System and method for camera calibration

By detecting feature points and generating depth maps through a stereo camera system, the complexity and inaccuracy of traditional calibration methods are solved, and fast and accurate camera calibration without a target is achieved, ensuring the reliability and accuracy of the camera during surgery.

CN120752910APending Publication Date: 2025-10-03DIGITAL SURGERY LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014362.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-13
Filing Date
2024-03-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In laparoscopic or endoscopic surgery, traditional camera calibration methods require the use of a checkerboard target, and the calibration parameters may change over time and use, especially after autoclaving and sterilization, making calibration inconvenient and inaccurate.

Method used

Using a stereo camera system, feature points in the image are detected and cross-matched to generate transformation parameters. These parameters are compared with pre-stored extrinsic calibration parameters to verify the camera calibration status. Depth maps and virtual volumes are used to confirm the success or failure of the calibration. Semi-automatic or automatic adjustment of calibration parameters is supported to achieve accurate calibration.

Benefits of technology

This enables quick and accurate verification and adjustment of camera calibration without the need for a checkerboard target, ensuring camera reliability and accuracy during surgery and reducing surgical interruptions and calibration complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752910A_ABST
    Figure CN120752910A_ABST
Patent Text Reader

Abstract

An imaging system includes a stereo camera having a first sensor configured to output a first image in a first channel and a second sensor configured to output a second image in a second channel. The system also includes an image processing device coupled to the stereo camera. The image processing apparatus includes a processor configured to perform a calibration process of the stereo camera. The process includes: generating a first map based on a first image of a first channel and generating a second map based on a second image of a second channel; generating a first surface based on the first map of the first channel; and generating a second surface based on the second map of the second channel; a virtual volume enclosed between the first surface and the second surface is generated. The process further includes determining whether the first image and the second image are displayed within the virtual volume; indicating that calibration of the stereo camera is successful based on the first image and the second image being displayed within the virtual volume; and starting an imaging operation of the stereo camera.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Laparoscopic or endoscopic visualization systems provide the surgeon with static images and / or video images during laparoscopic surgery. Traditional endoscopy involves the use of a camera to record the field of view and display the acquired images on a screen for the surgeon to review. In minimally invasive surgery (MIS), the laparoscope is typically a monocular endoscopic camera, while in robotic-assisted surgery (RAS), the endoscopic camera is typically a binocular stereo camera that provides depth perception to the surgeon during surgery. In recent years, with the advent of computational image processing, both monocular and stereo endoscopes have been increasingly used for depth mapping of the surgical scene. Many methods have been developed to reconstruct the 2.5D surgical scene from single and multiple monocular images or stereo image pairs, including stereo reconstruction, structure from motion (SfM), and simultaneous localization and mapping (SLAM).

[0002] Depth mapping of the surgical site has many applications in MIS and RAS, including avoiding critical structures and augmented reality (AR) overlays on preoperative imaging models. In order for these RAS applications to be clinically reliable, the camera's intrinsic and extrinsic parameters must be accurately calibrated to achieve robust stereo reconstruction. Traditionally, stereo camera calibration is performed using a checkerboard target placed at known locations in the scene. The camera is then used to acquire an image of the target, and the positions of the checkerboard corners in the image are used to estimate the camera's intrinsic and extrinsic parameters. However, in MIS or RAS applications in the operating room, the use of a checkerboard target and time-consuming calibration is less than ideal. Furthermore, camera calibration parameters may change over time and / or use, especially after autoclaving and sterilization processes, and may require regular recalibration. Therefore, there is a need for a system and method that verifies accurate calibration without a checkerboard target and calibrates the camera before and during each use. Summary of the Invention

[0003] The present disclosure provides a system and method for calibrating a laparoscope or endoscopic camera. The camera can be a monocular laparoscope commonly used in MIS or a binocular stereo endoscope commonly used in RAS. According to one embodiment of the present disclosure, an imaging system is disclosed. The imaging system includes a stereo camera having a first sensor configured to output a first image in a first channel, a second sensor configured to output a second image in a second channel, and an image processing device coupled to the stereo camera. The image processing device includes a processor configured to detect a first plurality of features in a first image of the first channel, detect a second plurality of features in a second image of the second channel, and generate cross-matches between the first plurality of features and the second plurality of features. The processor is further configured to transform the cross-matches to obtain a plurality of transformation parameters, and compare the plurality of transformation parameters with a plurality of extrinsic calibration parameters of the stereo camera. The processor is further configured to verify the calibration of the stereo camera based on the comparison of the transformation parameters with the extrinsic calibration parameters, and output a message indicating whether the calibration of the stereo camera is verified.

[0004] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the imaging system may include a monitor configured to display the first image and the second image. The monitor may be further configured to display a message indicating whether the calibration of the stereo camera is verified. The processor may be further configured to compare the difference between at least one of the multiple transformation parameters and at least one of the multiple extrinsic calibration parameters with a threshold value. The message may be an alarm indicating that the verification of the calibration has failed, and the processor may be configured to generate the alarm in response to the difference being greater than the threshold value. The message may also be a prompt indicating that the verification of the calibration has succeeded, and the processor may be configured to generate the prompt in response to the difference being less than the threshold value. The processor may be further configured to load the calibration parameters and the threshold value from the camera.

[0005] According to another embodiment of the present disclosure, an imaging system is disclosed. The imaging system includes a stereo camera having a first sensor configured to output a first image in a first channel and a second sensor configured to output a second image in a second channel. The system also includes an image processing device connected to the stereo camera. The image processing device includes a processor configured to generate a depth map based on a first image of the first channel and a second image of the second channel. Depth map generation may involve loading monocular or stereo camera calibration parameters, correcting the image to remove lens distortion, and calculating a dense disparity map. The processor is also configured to generate a first surface based on the depth map of the first channel and a second surface based on the depth map of the second channel. The processor is further configured to generate a virtual volume enclosed between the first surface and the second surface, and display the virtual volume on the first image and the second image.

[0006] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the imaging system may further include a monitor configured to display a virtual volume on the first image and the second image. The processor may be further configured to output a prompt on the monitor asking whether the first image and the second image are displayed within the virtual volume to confirm the calibration of the stereo camera. The processor may be further configured to receive a user response in response to the prompt. The processor may be further configured to output an alarm indicating that the verification of the calibration has failed in response to a negative response to the prompt. The processor may be further configured to output a message indicating that the verification of the calibration has succeeded in response to a positive response to the prompt. The processor may be further configured to generate a virtual volume based on the error value. The processor may be further configured to load the error value from the camera.

[0007] According to another embodiment of the present disclosure, a method for verifying the calibration of a stereo camera is disclosed. The method includes capturing a first image in a first channel at a first sensor of the stereo camera and capturing a second image in a second channel at a second sensor of the stereo camera. The method also includes generating a depth map at a processor based on the first image of the first channel and the second image of the second channel. The method also includes generating a first surface based on the depth map of the first channel, generating a second surface based on the depth map of the second channel, and generating a virtual volume enclosed between the first surface and the second surface at the processor. The method also includes displaying the virtual volume on the first image and the second image.

[0008] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include displaying a virtual volume on a first image and a second image on a monitor. In addition to or in lieu of depth maps and generation, key points or landmarks may be detected on known surfaces and / or structures (e.g., surgical instruments), where the distances between the key points and the landmarks are known. These maps may be used to refine the calibration. The key points and / or landmarks may be displayed and highlighted in the volume. The method may further include outputting a prompt on the monitor asking whether the first image and the second image are displayed within the virtual volume to confirm the calibration of the stereo camera. The method may further include receiving a user response to the prompt at the processor. The method may further include: outputting an alert on the monitor indicating that verification of the calibration failed in response to a negative response to the prompt; and outputting a message on the monitor indicating that verification of the calibration was successful in response to a positive response to the prompt.

[0009] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include, upon a validation test failure, displaying a rectified image and a virtual volume (e.g., a point cloud) to a user along with a set of slider controls for a plurality of calibration parameters. The method may rank the plurality of calibration parameters by the significance of their impact on the depth map and the morphology of the virtual volume (e.g., the point cloud). The method may also restrict the range of each slider for each ranked calibration parameter to a limited set of values ​​for user interaction. The value ranges may be generated based on an optimization mechanism. The method may further include receiving, at a processor, a user response in the form of adjustments to the plurality of sliders, each labeled with a name of a calibration parameter ranked by importance. The user response may be saved in the imaging system as a user preference. The method may generate, in real time, the rectified image, depth map, and point cloud based on each user input to the sliders. The method may further include, at the conclusion of the interactive calibration parameter optimization process, receiving, at the processor, a user response indicating that the generated rectified image is vertically aligned. The method may further include: at the end of the interactive calibration parameter optimization process, receiving a user response at the processor indicating that the first image and the second image are displayed within the virtual volume (point cloud) to confirm the calibration of the stereo camera. The method may further include receiving a user response to the prompt at the processor. The method may further include outputting an alert on the monitor indicating whether the semi-automatic calibration parameter optimization process passed or failed.

[0010] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include displaying the corrected left channel image and right channel image, and the polar lines of the binocular stereo endoscope on a monitor. The method may further include outputting a prompt on the monitor asking whether the polar lines displayed on the first image and the second image are aligned in the vertical direction. The method may additionally include receiving a user response to the prompt at the processor. The method may further include: outputting an alarm on the monitor indicating that the verification of the calibration has failed in response to a negative response to the prompt; and outputting a message on the monitor indicating that the verification of the calibration has succeeded in response to an affirmative response to the prompt.

[0011] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include a process for estimating intrinsic camera parameters and extrinsic camera parameters based on known or unknown objects in the operating environment of the robotic system. This feature is based on the observation that the system can load factory-calibrated camera parameters from a memory and use the acquired images to optimize the camera parameters if the verification fails. The method can present an interface to the user that displays a left and right channel stereo image pair and allows the user to select several matching points on both the left channel image and the right channel image. Deep learning or another machine learning image processing algorithm can automatically detect landmarks and / or key points. The method can further optimize the intrinsic camera parameters and extrinsic camera parameters to produce a final output in which similar features in both the left channel image and the right channel image are aligned on the same epipolar lines, thereby producing optimal correction and accurate stereo reconstruction.

[0012] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include performing a target-free self-calibration method based on intraoperative images at the beginning of a surgical procedure. The method may use a first stereo image pair to predict two dense depth maps from the perspectives of a left camera and a right camera using classical or deep learning methods for stereo reconstruction. The method may use the predicted depth maps and camera calibration parameters to generate a predicted point cloud from the first set of stereo image pairs. The method may further process a second stereo image pair from a previous timestamp to estimate the relative pose of the stereo endoscope in the surgical site based on the first and second stereo image pairs using a combination of visual SLAM and forward kinematics of a robotic arm holding the stereo endoscope. The method may use the deformed predicted point cloud and the actual first stereo image pair to synthetically generate a second pair of synthetic images. The view-synthesized left and right images may be subtracted from the actual second stereo image pair to generate a left visual loss image and a right visual loss image. The method may use the visual loss images to optimize intrinsic camera parameters and extrinsic camera parameters until the visual loss is below a threshold. Calibration parameter optimization methods can be based on classical or reinforcement learning algorithms, starting with the initial factory camera calibration parameters loaded from the camera or system memory. In addition to point clouds, keypoints can also be used. A keypoint detector can be used to find high-fidelity keypoints in the left and right images (or monocular images over time) to enhance rectification and calibration. Keypoints can also be displayed on a volume.

[0013] According to another embodiment of the present disclosure, a method for intraoperative calibration of a stereo camera is disclosed. The method includes: capturing a first stereo image pair from a stereo camera; capturing a second stereo image pair from the stereo camera; generating a predicted point cloud at a processor using a plurality of calibration parameters; and generating a deformed stereo image pair at the processor based on the second stereo image pair and the predicted point cloud. The method also includes: generating a visual loss image representing calibration parameter errors at the processor; optimizing the plurality of calibration parameters to minimize visual loss; and displaying a third stereo image pair on a monitor using the optimized plurality of camera calibration parameters.

[0014] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the method may include generating a depth map at a processor based on a plurality of calibration parameters. The method may also include generating a predicted point cloud at the processor based on the depth map. The method may further include generating a predicted pose of a stereo camera at the processor. The method may further include generating a deformed stereo image pair based on the predicted pose at the processor. The generation of the predicted pose may be based on kinematic data of a robotic arm that controls the movement of the stereo camera. The method may also include storing a plurality of optimized camera calibration parameters in a memory.

[0015] According to another embodiment of the present disclosure, an imaging system is disclosed. The imaging system includes a stereo camera, a monitor that displays images captured by the stereo camera, and an image processing device coupled to the stereo camera and the monitor. The image processing device includes a processor configured to: capture a first stereo image pair from the stereo camera; capture a second stereo image pair from the stereo camera; and generate a predicted point cloud using a plurality of camera calibration parameters. The processor is further configured to: generate a deformed stereo image pair based on the second stereo image pair and the predicted point cloud; generate a visual loss image representing calibration parameter errors; optimize the plurality of camera calibration parameters to minimize visual loss; and display a third stereo image pair on the monitor using the optimized plurality of camera calibration parameters.

[0016] Implementations of the above embodiments may include one or more of the following features. According to one aspect of the above embodiments, the system may further include a robotic arm that holds the stereo camera. The image processing device may be further configured to generate a predicted pose of the stereo camera based on kinematic data of the robotic arm.

[0017] According to another embodiment of the present disclosure, an imaging system is disclosed. The imaging system includes a stereo camera having a first sensor configured to output a first image in a first channel and a second sensor configured to output a second image in a second channel. The system also includes an image processing device connected to the stereo camera. The image processing device includes a processor configured to perform a calibration process for the stereo camera. The process includes: generating a first image based on a first image of the first channel and generating a second image based on a second image of the second channel; generating a first surface based on the first image of the first channel; and generating a second surface based on the second image of the second channel; generating a virtual volume enclosed between the first surface and the second surface. The process also includes: determining whether the first image and the second image are displayed within the virtual volume; indicating that the calibration of the stereo camera is successful based on the first image and the second image being displayed within the virtual volume; and enabling imaging operations of the stereo camera.

[0018] Implementations of the above embodiments may include one or more of the following features. Based on one aspect of the above embodiment, the calibration process may further include: indicating a calibration failure of the stereo camera based on at least one of the first image or the second image being displayed within the virtual volume; and disabling imaging operations of the stereo camera. The processor may be further configured to repeat the calibration process in response to a calibration failure. The first image and the second image may be at least one of a depth map or a surface normal map. The processor may be further configured to display the virtual volume over the first image and the second image. The imaging system may include a monitor configured to display the virtual volume over the first image and the second image. The monitor may be a heads-up display. The processor may be further configured to output a prompt on the monitor inquiring whether the first image and the second image are displayed within the virtual volume. The processor may be further configured to receive a user response in response to the prompt. The processor may be further configured to output an alert indicating a calibration verification failure in response to a negative user response. The processor may be further configured to output a message indicating a calibration verification success in response to a positive user response. The processor may be further configured to generate the virtual volume based on the error value. The processor may be further configured to load the error value from the stereo camera. The imaging processing apparatus may further include a memory, and the processor is further configured to load the error value from the memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] When considered in conjunction with the following detailed description, the present disclosure may be understood by reference to the accompanying drawings, in which:

[0020] Figure 1 is a schematic diagram of an imaging system according to an embodiment of the present disclosure;

[0021] Figure 2 yes Figure 1A stereoscopic image of a stereoscopic endoscope camera of an imaging system;

[0022] Figure 3A and Figure 3B is a schematic diagram of an image processing unit according to an embodiment of the present disclosure;

[0023] Figure 4 is a perspective view of a surgical robot system including an imaging system according to an embodiment of the present disclosure;

[0024] Figure 5 is a flow chart of a method for performing automatic calibration of an endoscopic camera of an imaging system according to an embodiment of the present disclosure;

[0025] Figure 6 is a schematic diagram of a stereo camera performing calibration according to an embodiment of the present disclosure;

[0026] Figure 7 is a flow chart of a method for performing automatic calibration of an endoscopic camera of an imaging system according to another embodiment of the present disclosure;

[0027] Figure 8 is a flow chart of a method for performing automatic calibration during white balance calibration according to an embodiment of the present disclosure;

[0028] Figure 9 is a flow chart of an intraoperative calibration method for correcting camera calibration parameters after verification failure according to an embodiment of the present disclosure;

[0029] Figure 10 shows a stereo image pair rectified using accurate calibration parameters;

[0030] Figure 11 shows a stereo image pair rectified using suboptimal camera calibration parameters;

[0031] Figure 12 A graphical user interface (GUI) showing an interactive semi-automatic parameter optimization method according to an embodiment of the present disclosure is shown; and

[0032] Figure 13 Shown by using Figure 12 The GUI adjusts the parameters of the generated point cloud to generate the virtual volume. DETAILED DESCRIPTION

[0033] Embodiments of the system disclosed herein are described in detail with reference to the accompanying drawings, in which like reference numerals denote like or corresponding elements in each of the several views. In the following description, well-known functions or configurations are not described in detail to avoid obscuring the present disclosure with unnecessary detail. Those skilled in the art will appreciate that the present disclosure can be adapted for use with any imaging system.

[0034] refer to Figure 1 , the imaging system 10 includes an image processing unit 20 that is configured to be coupled to one or more cameras, such as an endoscopic camera 12 or an open surgical camera 13 that is configured to be coupled to an endoscope 14. The system 10 also includes a light source 16 coupled to the cameras 12 and 13. The light source 16 can include any suitable light source including a light emitting diode, a lamp, a laser, etc., such as white light, near infrared light, etc.

[0035] The image processing unit 20 is configured to receive images from cameras 12 and 13 and process the raw image data signals to generate mixed white light and NIR images for recording and / or real-time display. The image processing unit 20 is also configured to use various AI image enhancements to mix the images.

[0036] refer to Figure 2 , the endoscope 14 is shown as a stereoscopic endoscope having a housing 15 and a shaft 17 extending distally from the housing. The endoscope 14 also includes a pair of objective lenses 18a and 18b (i.e., first and second optical channels for the left and right sides) disposed at the distal end of the shaft 17. The endoscope 14 also includes a fiber optic input adapter 15a extending radially from the housing 15. The housing 15 includes a proximal coupling interface 15b that is configured to couple the camera 12 to a pair of output optical elements (not shown). The endoscope 14 may include a plurality of lenses, prisms, reflectors, etc. to enable light transmission from the objective lenses 18a and 18b to the output elements. In another embodiment, the endoscope 14 may be a monocular endoscope, in which case a single image channel formed by a single objective lens and optical output element is provided to the camera 12.

[0037] refer to Figure 3A and Figure 3B, the image processing unit 20 is connected to the cameras 12 and 13 via a camera connector 22, which is in turn connected to a frame grabber 24, which is configured to capture individual digital still frames from the digital video stream. The frame grabber 24 is connected to a first processing unit 28 and a second processing unit 29 via a peripheral component interconnect express (PCI-E) bus 26. The first processing unit 28 can be configured to perform the operations, calculations and / or instruction sets described in this disclosure, and can be a hardware processor, a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), a microprocessor, and combinations thereof. Those skilled in the art will understand that the processor can be any logical processor (e.g., a control circuit) suitable for executing the algorithms, calculations and / or instruction sets described herein. The second processing unit 29 can be a graphics processing unit (GPU) or an FPGA, which is more suitable for processing images because it has a larger number of cores (e.g., thousands of compute unified device architecture (CUDA) cores) than a CPU (e.g., the first processing unit 28).

[0038] The image processing unit 20 also includes various other computer components, such as memory 70, storage device 73, peripheral ports 74, and input devices (e.g., a touch screen). In addition, the image processing unit 20 is coupled to one or more monitors 72 via output port 76. The image processing unit 20 is configured to output the processed image to any suitable video output port (e.g., DISPLAYPORT 76) capable of transmitting the processed image at any desired resolution, display rate, and / or bandwidth. TM 、 SDI, etc.) outputs the processed image.

[0039] Continue to refer Figure 3A and Figure 3B , cameras 12 and 13 include a pair of visible (VIS) image sensors 80a and 80b for stereoscopic white light (i.e., visible light) imaging (e.g., from about 380 nm to about 700 nm) and separate near-infrared (NIR) image sensors 81a and 81b for capturing NIR fluorescence (e.g., wavelengths from about 825 nm to about 850 nm). VIS image sensors 80a and 80b may include Bayer filters or any other filters suitable for single-chip color imaging. In embodiments, cameras 12 and 13 may also use multiple color sensors, for example, one sensor for each color (RGB) channel.

[0040] The imaging system 10 can also be used with Figure 4The surgical robotic system 11 is shown integrated. The control tower 21 is connected to all components of the surgical robotic system 11, including the surgeon's console 30 and one or more movable carts 60. Each of the movable carts 60 includes a robotic arm 40 with an attached device, such as an endoscopic camera 12. Each of the robotic arms 40 includes a plurality of links 42 that are movable relative to one another about joints 44. These joints can have any number of degrees of freedom (e.g., one or more), thereby providing the robotic arm 40 with multiple degrees of freedom. The robotic arm 40 includes actuators 45 (e.g., motors, transmissions, cables, drive shafts, etc.) and sensors 43 configured to provide feedback to control the movement of the robotic arm 40. The sensors can include electrical sensors, torque sensors, force sensors, strain sensors, temperature sensors, position sensors, etc. Each of the robotic arms 40 also includes an instrument drive unit (IDU) 52, which is configured to couple to the actuation mechanism of the attached device and is configured to move (e.g., rotate) and actuate the device. During endoscopic surgery, the endoscopic camera 12 may be inserted through an endoscope access port (not shown) held by the robotic arm 40 .

[0041] The surgeon's console 30 includes a first screen 32 that displays a video feed of the surgical site provided by the camera 12 and a second screen 34 that displays a user interface for controlling the surgical robotic system 11. The first screen 32 and the second screen 34 may be touch screens (e.g., monitors 72) that allow for the display of various graphical user inputs. In an embodiment, ultrasound images may also be displayed on the first screen 32 and the second screen 34. The surgeon's console 30 also includes a plurality of user interface devices, such as a foot pedal 36 and a pair of hand controllers 38a and 38b that are used by the user to remotely control the robotic arm 40 and the endoscopic camera 12.

[0042] The control tower 21 also serves as an interface between the surgeon's console 30 and one or more robotic arms 40. Specifically, the control tower 21 is configured to control the robotic arms 40, such as by moving the robotic arms 40 and attached devices based on a set of programmable instructions and / or input commands from the surgeon's console 30, such that the robotic arms 40 and attached devices execute a desired movement sequence in response to input from the foot pedals 36 and hand controls 38a and 38b. The foot pedals 36 can be used to enable and disable the hand controls 38a and 38b, thereby repositioning the endoscopic camera 12. Specifically, the foot pedals 36 can be used to clutch the hand controls 38a and 38b. By depressing one of the foot pedals 36, the clutch is activated, disconnecting the hand controls 38a and / or 38b from the robotic arms 40 and attached devices (i.e., preventing movement input). This allows the user to reposition the hand controls 38a and 38b without moving the robotic arm(s) 40 and endoscopic camera 12. This is useful when reaching the control boundaries of the surgical space.

[0043] The present disclosure provides a system and method for verifying the calibration of a stereo camera (i.e., camera 12) without the use of additional calibration targets and with minimal disruption to the surgical workflow. This rapid verification can be performed in the operating room before surgery or intraoperatively during surgery.

[0044] The camera can be a stereo camera that can be calibrated before use in a surgical environment. Calibration can be performed using a calibration pattern, which can be a checkerboard pattern of black and white squares. The calibration pattern can be set on one of the instruments 50, for example, on a shaft. In an embodiment, the calibration pattern can be set on one of the entry ports 55, inside or outside the cannula portion of the entry port. The calibration pattern can then be used when the camera 51 is inserted into the entry port 55. The insertion or any other movement of the camera 51 can be paused to allow the camera 51 to image the calibration pattern. In an additional embodiment, a custom calibration pattern can be etched or otherwise affixed to one of the robotic arms 40a-d, and the robotic arm 40 holding the camera 51 can be manually manipulated or automatically manipulated (e.g., programmed with path points) to capture the calibration pattern image for camera calibration. In addition, the current calibration parameters can be verified relative to the calibration pattern as described above based on the measured distance of the current calibration pattern.

[0045] In another embodiment, the calibration pattern can be projected by a projector disposed on the endoscope 14, which can include a light source (e.g., an LED) and a slide or lens on which the calibration pattern is disposed (e.g., etched, printed, etc.). In another embodiment, the calibration pattern can be disposed on a collapsible (e.g., foldable or rollable) sheet that is expanded or unfolded within the patient's body. The calibration pattern can be inserted in a collapsible form through one of the entry ports 55a-d, used for calibration, collapsed, and then removed. In an additional embodiment, the calibration pattern can be disposed in a trocar capable of calibration (e.g., on a foldable flap at the bottom of the trocar, wherein the calibration pattern is etched on the flap). The calibration process can be performed during insertion of the endoscope through the trocar capable of calibration. The process may require stopping the insertion of the endoscope through the trocar at a certain depth while the endoscope camera head is inside the trocar, capturing a few images of the calibration pattern on the calibration tabs of the trocar, performing the calibration, and instructing the user to push the endoscope further through the trocar for full insertion.

[0046] Calibration may include obtaining multiple images at different poses and orientations and providing these images as input to an image processing unit, which outputs calibration parameters for use during camera use. The calibration parameters may include one or more intrinsic and / or extrinsic parameters, including but not limited to principal point location, focal length, skew, sensor scale, distortion coefficients, rotation matrices, translation vectors, etc. The calibration parameters may be stored in a memory of camera 12 and loaded by image processing unit 20 when it is connected to camera 12.

[0047] refer to Figure 5 The method for calibrating the camera 12 and the endoscope 14 includes calibrating the sensors 80a and 80b by detecting features from each channel (i.e., from each of the sensors 80a and 80b of the camera 12) and then performing feature matching on the features detected in one channel with the features detected in the other channel. The feature matching is performed in the image processing unit 20 using a keypoint detector in combination with a descriptor, including but not limited to oriented FAST and rotated BRIEF (ORB) features.

[0048] After performing feature matching using robust keypoint descriptors, the relative transformation between the two cameras of the stereo endoscope is estimated. The system then estimates intrinsic and extrinsic camera parameters based on the initial factory calibration parameters and the estimated transformation based on feature matching. After performing correction using the existing intrinsic and extrinsic camera parameters, the features detected from one channel are checked to see if they match on the epipolar lines of the other channel, and vice versa. An affine transformation matrix is ​​then formed to transform the image from one channel to the other. The transformation matrix is ​​compared to a preset threshold (e.g., + / - 1% deviation) to determine whether it is within the tolerance range of the existing parameters. If the matrix is ​​outside the tolerance range, the test fails and the camera 12 and endoscope 14 are deemed unusable for surgery. The image processing unit 20 outputs an alarm on one of the monitors 72 and prevents the camera 12 and endoscope 14 from being further used with the imaging system 10.

[0049] refer to Figure 6 , which schematically shows Figure 5 In the calibration process, at step 100, the image processing unit 20 detects features from the first (e.g., right) channel to obtain a first plurality of features (FEATS_C1). Features can be detected using an Oriented FAST and Rotated BRIEF (ORB) feature detector, which is a computer vision algorithm used for object recognition or 3D reconstruction. The ORB feature detector itself is based on the Features from Accelerated Segmentation Test (FAST) keypoint detector and the Binary Robust Independent Elementary Features (BRIEF) visual descriptor. Any suitable computer vision feature detector, descriptor, and matching algorithm can be used.

[0050] In step 102, the image processing unit 20 detects features from the second (e.g., left) channel to obtain a second plurality of features (FEATS_C2). The features of the second channel are detected in the same manner as the features of the first channel. In step 104, the image processing unit 20 cross-matches the features detected from the first channel to the second channel image to obtain a cross-match of the features (FEATS_C1xC2). Any suitable image matching algorithm can be used, such as the Fast Nearest Neighbor Search Library (FLANN), which is an image matching algorithm for fast approximate nearest neighbor search in high-dimensional space. FLANN projects high-dimensional features into a low-dimensional space and then generates a compact binary code.

[0051] At step 106, the image processing unit 20 estimates a transformation from the first channel to the second channel based on the cross-matching of the features FEATS_C1xC2 to obtain a cross-matched transformation (TRANS_C1xC2) using an outlier-tolerant transformation estimation algorithm. In an embodiment, a random sample consensus (RANSAC) algorithm can be used to estimate a mathematical model from a data set containing outliers by identifying outliers in the data set and estimating an expected model using data that does not contain outliers.

[0052] At step 108 , the image processing unit 20 generates a combined transform that describes extrinsic camera parameters to transform salient points of the image from one channel (eg, the first channel) to another channel (eg, the second channel) to obtain generated combined transform parameters (TRANS).

[0053] In steps 110 and 111, the image processing unit 20 compares the generated transformation parameters (TRANS) with predetermined extrinsic camera calibration parameters. The difference between the corresponding transformation parameters and the calibration parameters is calculated, and then the difference is compared with a threshold value (DEL). If the difference between the tested extrinsic camera calibration parameters and the predetermined extrinsic camera calibration parameters is greater than the threshold value (i.e., the difference is greater than the threshold value), the image processing unit 20 issues an alarm indicating a verification failure in step 112, thereby advising the user to replace the laparoscopic camera 12 and / or endoscope 14. If the test is successful, the image processing unit 20 issues a success notification in step 114, indicating that the camera 12 and endoscope 14 can be used.

[0054] In addition to the calibration pattern, additional objects, such as one or more instruments 50, can be used for calibration. Instruments 50 include various landmarks with unique shapes and known dimensions, such as fasteners, edges, and components, that can be used as calibration patterns. Instruments 50 can be detected during use of the camera 12, and the detected instrument 50 landmarks can be used to optimize calibration parameters.

[0055] Any image processing algorithm, which may be an artificial intelligence or machine learning (AI / ML) algorithm, may be used to detect the instrument 50 in the video feed captured by the camera 12. The images or video feed of the instrument and its specific landmarks may be used as a dataset for training the AI / ML detection algorithm.

[0056] Once the instrument is detected, the AI / ML algorithm also detects the instrument's landmarks. After detecting the landmarks, a depth mapping algorithm can then be used to measure the distances between the instrument's landmarks in 3D using the existing stereo calibration. The distances can then be used to optimize the calibration parameters to minimize the difference between the measured instrument shape and size and the actual instrument shape and size.

[0057] Another embodiment of the present disclosure describes a verification method based on generating a stereo reconstruction using existing intrinsic and extrinsic camera parameters. Figure 7 The flowchart of FIG. 1 describes a verification algorithm based on stereo reconstruction, which can be implemented as software instructions executable by the image processing unit 20. At step 200, the image processing unit 20 receives one or more images (SCENE) from the camera 12. The images can be obtained when the camera 12 remains stationary for a short period of time (e.g., 2 to 5 seconds) to capture at least one or more frames. The scene can be a surgical scene or any other scene.

[0058] In step 202, the image processing unit 20 creates a depth map (DEPTH) of the observed scene using predetermined calibration parameters. The depth map may also include a surface normal map that uses RGB information corresponding to the X-axis, Y-axis, and Z-axis in 3D space. Any suitable depth map generation algorithm may be used, such as a depth map automatic generator (DMAG), etc. In step 204, the image processing unit 20 loads an acceptable error (ERR) threshold when reconstructing the surface. The error threshold may be stored in the memory of the camera 12 and represents the difference between the depth map and the image, and may be stored on the storage device 73. In step 206, the image processing unit 20 creates two surfaces from the reference system of each of the first channel and the second channel based on the loaded error threshold (ERR), and the distance between the two surfaces and the reconstructed surface (depth map) is + / - the threshold distance.

[0059] In an embodiment, the threshold value may be part of software being executed by the image processing unit 20. The threshold value may be stored in a memory that may store a plurality of threshold values, and an appropriate threshold value may be selected based on the type (e.g., model) of camera 12 being used. In another embodiment, the user may adjust the threshold value via a graphical user interface of the image processing unit 20. The user-adjusted threshold value may be saved as part of the surgeon's preferences.

[0060] In step 208, the image processing unit 20 creates a virtual error volume (ERR VOL) ​​that is enclosed between the two surfaces created in step 206. The virtual error volume represents an acceptable error. In step 210, the image processing unit 20 displays a semi-transparent image of the virtual error volume ERR VOL. The image processing unit 20 superimposes the semi-transparent image of the virtual error on the channel image used to create the depth map. The image processing unit 20 displays the superimposed image (ERR VOL OVER SCENE) on the monitor 72 and / or one of the screens 32 and 34. In an embodiment, the superimposed image can be displayed on a head-mounted device, which can be a virtual reality or augmented reality head-mounted device, such as a Microsoft Corporation of Redmond, Washington. METAQUEST, available from Meta, Inc., Menlo Park, California

[0061] At steps 212 and 213, the image processing unit 20 outputs a prompt asking the user to verify whether the image (SCENE) from the camera 12 is present in the ERR VOL. If the response is positive, the image processing unit 20 indicates that the calibration has passed and continues to use the camera 12 at step 214. The image processing unit 20 indicates via the success prompt that the camera 12 and the endoscope 14 can be used, and the image processing unit 20 enables the camera 12 for imaging, and the camera 12 can be used to image the surgical site.

[0062] If the response is negative, then at step 216, the image processing unit 20 outputs an alert indicating that the calibration verification failed and recommends that the user replace the laparoscopic camera 12 and / or endoscope 14. In an embodiment, after a verification failure, the verification algorithm may be repeated before recommending replacement of the camera 12 and / or endoscope 14.

[0063] In another embodiment, rather than relying on user verification, a computer vision algorithm can be used to perform the comparison of the volume with the image from the camera 12. The image processing device 20 is configured to automatically determine whether the first image and the second image are displayed within the virtual volume based on a threshold. An AI / ML image processing algorithm can be used to determine the position of the volume relative to the threshold.

[0064] Figure 7 The method can also be implemented during white balance calibration of the camera 12 and endoscope 14 and includes calibrating the white balance of the sensors 80a and 80b of the camera 12. The sensors 80a and 80b provide RGB color imaging and provide color and texture information to the surgeon during surgery. The camera color and texture information may drift over time, requiring white balance adjustment of the camera 12. Figure 8 Shown in Figure 7 The method used for white balance calibration during stereo calibration. Figure 8At step 300, the image processing unit 20 receives one or more images of a white balance control object (e.g., a white piece of paper) from the camera 12. At step 302, the image processing unit 20 verifies whether the extrinsic camera parameters are within an operational range by imaging the white balance control object. At step 302, the image processing unit 20 fits the plane of the white balance control object to the image in each channel and measures depth using the stored existing intrinsic and extrinsic camera parameters. Next, at step 304, the image processing unit 20 determines the size of the white balance control object based on the depth estimate and plane fitting. The image processing unit 20 then compares this size to a tolerance range, which may be + / - 1%, to pass the test step 306. If the size is within the tolerance range, the camera 12 is calibrated and can be used, and at step 308, the image processing unit 20 outputs a corresponding success message. If the size is outside the tolerance range, the camera 12 is not properly calibrated, and at step 310, the image processing unit 20 outputs a corresponding failure message.

[0065] Figure 9 It shows that no checkerboard pattern is required, and the use of Figure 4 The intraoperative calibration workflow for calibrating the surgical robotic system 10 is described. Alternatively, if necessary, the workflow uses a few frames from the initial exploratory phase of the procedure to optimize the camera calibration parameters. At step 400, the camera 12 and endoscope 14 are calibrated. These can be initial factory calibration parameters or newly optimized camera calibration parameters, which are stored in the memory of the endoscope 14 or stored individually on the system 10 for each endoscope 14 based on its unique serial number.

[0066] Figure 10 and Figure 11 A pair of stereoscopic images from camera 12 obtained during RAS surgery is shown. Figure 10 shows the left and right stereo image pairs rectified using accurate camera calibration parameters, and Figure 11 The same stereo image pair is shown rectified using suboptimal camera calibration parameters. The mismatch in the same orientation between the left and right stereo image pairs results in noisy dense depth maps and erroneous stereo reconstruction. Figure 10 In , rectangles 500 and 502 show similar content in the same (eg, vertical) orientation, demonstrating a highly accurate camera calibration. Figure 11 In FIG, rectangles 510 and 512 show different content in the same direction due to inaccurate camera calibration.

[0067] At step 402, the system 10 loads the most recently optimized camera calibration parameters, and at step 404, using the current stereo image pair 406, the system 10 first uses stereo reconstruction through a depth prediction network 410 to predict a dense depth map 408. The system 10 predicts the dense depth map 408 from both the left camera image perspective and the right camera image perspective. At step 412, the dense depth map 408 is then used together with the camera calibration parameters to generate a predicted point cloud 414 relative to the left channel image and the right channel image.

[0068] The system 10 is also configured to process the current image and the previous image to predict the camera pose via a pose prediction network 416 in the form of visual SLAM. The pose network 416 uses the robotic arm kinematics data 418 and the previous and current image sets to estimate the position and pose 420 of the endoscope 14. The system 10 further generates a deformed point cloud 422 based on the predicted pose corresponding to the previous stereo image pair 426 provided to the pose network 416. Then, at step 424, the system 10 uses the deformed point cloud to synthesize a deformed stereo image pair 430 representing the first stereo pair to generate a visual loss function image 432. The visual loss image 432 represents the calibration parameter error and is then passed to the calibration parameter optimization network 434, which optimizes the camera calibration parameters. The depth prediction network 410, the pose prediction network 416, and the optimization network 434 can be any suitable deep learning network (e.g., a convolutional neural network, a recurrent neural network, a deep reinforcement network, a deep belief network, a transformer network, etc.) that can be trained on image and video data and corresponding camera calibration parameters.

[0069] refer to Figure 12 , shows a GUI 600 that can be displayed on one of the screens 32, 34. The GUI 600 outputs the optimized calibration parameters and allows the user to modify these parameters in real time while viewing the image. The effect of the correction on the image is also shown in FIG. Figure 10 and Figure 11 , where parallel (e.g., vertical) lines represent how the left and right images are shifted. GUI 600 includes a plurality of inputs 602a-d, each of which allows control of a single calibration parameter. The calibration parameters may include one or more intrinsic and / or extrinsic parameters, including but not limited to principal point location, focal length, skew, sensor scale, distortion coefficients, rotation matrix, translation vector, etc. Each of inputs 602a-d may be a slider, a drop-down menu, or any other GUI adjustment selector suitable for changing a numerical value. As the calibration parameters are changed through GUI 600, Figure 10 and Figure 11 Images and Figure 13The point cloud 508 shown is adjusted in real time. In one embodiment, these interactively optimized endoscope-specific calibration parameters can be saved in the system along with the endoscope camera serial number each time the calibration parameters are updated. In another embodiment, the calibration parameters can be saved in the endoscope camera programmable memory each time the calibration parameters are updated. In yet another embodiment, the calibration parameters can be saved as part of a particular user's preferences.

[0070] Although several embodiments of the present disclosure have been shown in the drawings and / or described herein, it is not intended that the disclosure be limited thereto, as the intention is that the scope of the disclosure be as broad as the art will allow, and this specification should be read in the same manner. Therefore, the above description should not be construed as limiting, but rather as merely exemplifying particular embodiments. Those skilled in the art will envision other modifications within the scope of the appended claims.

Claims

1. An imaging system comprising: a stereo camera comprising a first sensor configured to output a first image in a first channel and a second sensor configured to output a second image in a second channel; as well as An image processing device coupled to the stereo camera, the image processing device comprising a processor configured to perform a calibration process for the stereo camera, the process comprising: generating a first image based on the first image of the first channel and generating a second image based on the second image of the second channel; generating a first surface based on the first map of the first channel; generating a second surface based on the second map of the second channel; generating a virtual volume enclosed between the first surface and the second surface; determining whether the first image and the second image are displayed within the virtual volume; indicating a successful calibration of the stereo camera based on the first image and the second image being displayed within the virtual volume; and Enables imaging operations for this stereo camera.

2. The imaging system according to claim 1, wherein The calibration process further includes indicating a calibration failure of the stereo camera based on at least one of the first image or the second image being displayed within the virtual volume; as well as Disables imaging operations for this stereo camera.

3. The imaging system according to claim 2, wherein: The processor is further configured to repeat the calibration process in response to a calibration failure.

4. The imaging system according to claim 1 or claim 2, wherein: The first map and the second map are at least one of a depth map, a surface normal map, a key point map, or a landmark point map of a known structure.

5. An imaging system according to any preceding claim, wherein The processor is further configured to display the virtual volume on the first image and the second image. 6 . The imaging system of claim 5 , further comprising a monitor configured to display the virtual volume on the first image and the second image.

7. The imaging system according to claim 6, wherein: The monitor is a head-up display.

8. The imaging system according to claim 6 or claim 7, wherein: The processor is further configured to output a prompt on the monitor asking whether the first image and the second image are displayed within the virtual volume.

9. The imaging system according to any one of claims 6 to 8, wherein: The processor is further configured to automatically determine whether the first image and the second image are displayed within the virtual volume based on a threshold value that is adjustable and saved in user settings.

10. The imaging system according to claim 8 or claim 9, wherein: The processor is further configured to receive a user response in response to the prompt.

11. The imaging system according to claim 10, wherein: The processor is further configured to output an alert indicating a calibration verification failure in response to a negative user response.

12. The imaging system of claim 10 or claim 11, wherein: The processor is further configured to output a message indicating that the calibration verification was successful in response to the affirmative user response.

13. An imaging system according to any preceding claim, wherein: The processor is further configured to generate the virtual volume based on an error value, and wherein the error value is adjustable and saved in user settings.

14. The imaging system of claim 13, wherein: The processor is further configured to load the error value from the stereo camera.

15. An imaging system according to claim 13 or claim 14, wherein: The imaging processing device includes a memory, and the processor is further configured to load the error value from the memory.