Robotic endoluminal intervention remote presentation method and system based on augmented reality
Patent Information
- Application Number
- CN202311266997.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-27
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2043-09-27
AI Technical Summary
[0004]现有技术中通过优化医生端的远程控制模块,引入远程通信模块,实现血管介入手术信息的全面感知和加密传输,但是无法为临床医生提供直观的视觉反馈,存在待改进之处
[0064]1、本发明通过将一个单目相机固定于多自由度机械臂末端,并将单目相机和机械臂运动特性相结合以实现全向增强现实的功能,通过多目标手眼标定、迭代最近点和图像叠加算法将虚拟关键解剖结构和介入器械叠加在相机视频图像上,单目相机的视野方向可以通过机械臂的运动来改变,允许从不同方向观察并透视关键解剖结构和介入器械。此外,重建了患者端的场景,从而为机械臂的安全远程遥操作提供虚拟现实交互界面。将全向增强现实和虚拟现实与机器人远程呈现相结合,可以为临床医生提供直观的视觉反馈,从而促进机器人远程医疗的安全性。
Smart Images

Figure CN117064545B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote surgery technology, and more specifically, to a method and system for remote presentation of robotic intracavitary intervention based on augmented reality. Background Technology
[0002] Robotic telemedicine can provide timely treatment for critically ill patients in geographically remote areas. However, due to issues such as lack of depth information, occlusion, and limited field of vision, the visual feedback from the remote patient to the local clinician is often not intuitive, which can affect surgical safety.
[0003] Chinese patent application CN116370092A discloses a remote surgical system and control method for interventional surgery. The remote surgical system includes a surgical execution module, a remote control module, and a remote communication module. The surgical execution module includes an interventional surgical robot, a driver and motion control system for the interventional surgical robot, and sensors mounted on the interventional surgical robot. The remote control module senses the doctor's action commands as input and outputs control signals to the surgical execution module based on the input, thereby controlling and operating the surgical execution module. The remote communication module uses remote communication software, a remote transmission system, and an encrypted transmission protocol to achieve encrypted transmission between the driver and sensor signals of the surgical execution module and the commands of the remote control module.
[0004] Existing technologies optimize the remote control module on the doctor's end and introduce a remote communication module to achieve comprehensive perception and encrypted transmission of vascular interventional surgery information. However, they cannot provide clinicians with intuitive visual feedback and therefore require improvement. Summary of the Invention
[0005] In view of the deficiencies in the prior art, the purpose of this invention is to provide a method and system for remote presentation of robotic intracavitary intervention based on augmented reality.
[0006] According to the present invention, a robotic intracavitary intervention telepresence method based on augmented reality is provided, the presentation method comprising the following steps:
[0007] Step S1: Calibrate the hardware system on the patient's end;
[0008] Step S2: Collect multiple images from a monocular camera in different fields of view, perform 3D reconstruction of the actual scene at the patient's end, and build a VR interactive interface;
[0009] Step S3: Move the monocular camera to obtain multiple viewing directions and record the corresponding robotic arm posture information;
[0010] Step S4: On the monocular camera image, the virtual anatomical structure and virtual interventional instruments are superimposed onto the real scene using calibration results and image processing to construct an AR feedback interface;
[0011] Step S5: The doctor's visualization device remotely acquires the VR and AR interfaces of the patient's device and changes the field of view through remote operation or teaching.
[0012] Preferably, step S1 includes: calibration between the visualization robotic arm and the electromagnetic tracking system, calibration between the visualization robotic arm and the rigid instrument, hand-eye calibration between the monocular camera and the visualization robotic arm, and registration between the electromagnetic tracking system and the anatomical structure.
[0013] Preferably, the calibration between the visualization robotic arm and the electromagnetic tracking system includes: connecting a six-degree-of-freedom electromagnetic sensor to a rigid flange fixed to the end of the visualization robotic arm, moving the robotic arm, causing the six-degree-of-freedom electromagnetic sensor to move in the magnetic field of the electromagnetic tracking system, recording the poses of the robotic arm and the six-degree-of-freedom electromagnetic sensor, and constructing the following equations:
[0014]
[0015]
[0016] in, It visualizes the pose of the robotic arm's end effector. It is the transformation matrix from the end effector of the visualized robotic arm to the six-degree-of-freedom electromagnetic sensor. It is the transformation matrix from a visualized robotic arm to an electromagnetic tracking system. It refers to the pose of a six-degree-of-freedom electromagnetic sensor. and In this equation, i represents the dataset index for different visualized robotic arm postures. The above equation can be expressed as:
[0017]
[0018] make and The above equation can then be rewritten as AX = XB, and finally, X can be obtained.
[0019] Preferably, the calibration between the visualization robotic arm and the rigid device includes: fixing a monocular camera to the end of the visualization robotic arm, fixing an ArUco code to the end of the operating robotic arm, moving the visualization robotic arm and the operating robotic arm to move the monocular camera and the ArUco code to different positions, and ensuring that the ArUco code is always within the field of view of the monocular camera. The pose of the ArUco code is estimated by the monocular camera, and the following equation is constructed:
[0020]
[0021]
[0022] in, It is the transformation matrix from the visualization robotic arm to the monocular camera; It is the transformation matrix from the end effector of the visualization robotic arm to the monocular camera, which can be obtained in advance through basic hand-eye calibration. It is the transformation matrix from monocular camera to ArUco code. It is the transformation matrix from visualizing the robotic arm to operating the robotic arm; It is the transformation matrix from the end effector of the robotic arm to the ArUco code; therefore, the above can be expressed as:
[0023]
[0024] make and The transformation matrix between the visualized robotic arm and the operating robotic arm is then transformed into solving for AX = XB;
[0025] The registration of the electromagnetic tracking system and the anatomical structure includes the following steps:
[0026] Step S101: Perform a CT scan on the anatomical structure;
[0027] Step S102: Reconstruct a three-dimensional virtual model of the anatomical structure using multiple CT slices, and use this model as the target point cloud P;
[0028] Step S103: Collect the point cloud of the actual anatomical structure as the source point cloud G;
[0029] Step S104: Use the ICP algorithm to register the target point cloud P and the source point cloud G to obtain the transformation matrix between the electromagnetic tracking system and the anatomical structure.
[0030] Preferably, for step 2, the actual scene at the patient end is reconstructed in three dimensions, including sparse reconstruction and dense reconstruction. The incremental SfM algorithm is used for sparse reconstruction, and the MVS algorithm is used for dense reconstruction. The open source library OpenMVS is used to restore the complete surface of the reconstructed scene through surface reconstruction, optimization and texture mapping.
[0031] In sparse reconstruction, the monocular camera pose estimated by the SfM algorithm is represented as: The pose of the visualized robotic arm is The true scale of the monocular camera pose is recovered through a similarity transformation, as shown below:
[0032]
[0033] in, It is a similarity transformation matrix; and They are Translation vectors and rotation matrices; It is the scale of the isotropic scaling transformation, obtained using Umeyama's method. As shown below:
[0034]
[0035] Preferably, for step S4, the change matrix of anatomical structure and monocular camera:
[0036]
[0037] in, This refers to the pose of the anatomical structure within the space of a monocular camera. Since the pose of the visualization robotic arm's end effector is updated in real time, the virtual anatomical structure is superimposed onto the real-world scene using the following four steps:
[0038] Step S401: Use an open-source library to create a virtual 3D space containing a virtual camera and a 3D anatomical model;
[0039] Step S402: Set the intrinsic parameters of the virtual camera to be the same as those of the actual monocular camera, and make the reference frame of the virtual camera coincide with the reference frame of the virtual space.
[0040] Step S403, use Update the pose of the 3D anatomical model in the virtual space and use a virtual camera to capture the scene of the virtual space in real time;
[0041] Step S404: Use image processing algorithms to overlay the images captured by the virtual camera and the real monocular camera for enhanced display of anatomical structures;
[0042] For rigid instruments, the transformation matrix from a monocular camera to a rigid instrument is as follows:
[0043]
[0044] in, It is the pose of the rigid instrument relative to the monocular camera, the pose of the visualization robotic arm end effector, and the real-time update of the pose of the manipulating robotic arm end effector.
[0045] For flexible instruments, after shape reconstruction, the shape of the flexible instrument relative to the monocular camera can be obtained as follows:
[0046]
[0047] Where C(s) is the position vector relative to {C} at the arc length s of the flexible instrument, and the pose of the visualized end of the robotic arm, the pose of the operated end of the robotic arm, and p′(s) are updated in real time.
[0048] Preferably, a cube with an ArUco code attached to each face is used as the actual object, and a three-dimensional virtual cube model with an ArUco code image on each face is used as the virtual object. The actual cube and the electromagnetic tracking system are registered to obtain the pose of the cube under the electromagnetic tracking system.
[0049] The poses of real and virtual objects are transformed from an electromagnetic tracking system to a monocular camera, and an instrumental augmentation display method is used to enhance the display of the cube.
[0050] Move the monocular camera to different positions while keeping the actual object within the camera's field of view, and capture camera images before and after the virtual object is superimposed. Using the acquired images, estimate the 3D poses of the virtual and actual objects respectively. The superposition error can be calculated using the following formula:
[0051]
[0052] Where E overlay This is the average 3D error of the virtual and real images superimposed, where n is the number of image groups acquired. Each group includes the position vectors of the eight corner points of the virtual object estimated by the algorithm and the position vectors of the eight corner points of the actual object. Additionally, m = 8 represents the number of corner points of the cube. Let {C} be the position vector data of the j-th corner point of the three-dimensional virtual cube in the i-th group of {C}. Let {C} be the position vector data of the j-th corner point of the three-dimensional actual cube in the i-th group of {C}.
[0053] Preferably, for step S5, the feedback visuals provided from the patient to the clinician include AR view, VR view, and endoscopic view;
[0054] In the AR view, key anatomical structures of the patient and instruments inserted into the patient's body can be seen through, and different AR views can be obtained by moving the robotic arm.
[0055] By interacting with the display interface, the direction of the VR feedback field of view can be changed, thereby determining the safe distance between the visualized robotic arm and surrounding objects.
[0056] The movement of the robotic arm can be remotely controlled by a doctor, or the robotic arm can be taught to move using the posture information recorded in step S3.
[0057] The present invention provides an augmented reality-based robotic intracavitary intervention telepresence system, employing the augmented reality-based robotic intracavitary intervention telepresence method according to any one of claims 1-9, comprising the following modules:
[0058] Module M1 is used to calibrate the hardware system at the patient end;
[0059] Module M2 is used to acquire multiple images from a monocular camera in different fields of view, perform 3D reconstruction of the actual scene at the patient's end, and build a VR interactive interface;
[0060] Module M3 is used to move the monocular camera to obtain multiple viewing directions and record the corresponding robotic arm posture information;
[0061] Module M4 is used to overlay virtual anatomical structures and virtual interventional instruments onto a real scene using calibration results and image processing on a monocular camera image, thereby constructing an AR feedback interface;
[0062] Module M5 is a visualization device used by doctors to remotely acquire the VR and AR interfaces of the patient's device and change the field of view through remote operation or teaching.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1. This invention achieves omnidirectional augmented reality functionality by fixing a monocular camera to the end effector of a multi-degree-of-freedom robotic arm and combining the motion characteristics of the monocular camera and the robotic arm. Through multi-target hand-eye calibration, iterative nearest point, and image overlay algorithms, virtual key anatomical structures and interventional instruments are superimposed onto the camera's video images. The field of view of the monocular camera can be changed by the movement of the robotic arm, allowing for observation and perspective viewing of key anatomical structures and interventional instruments from different directions. Furthermore, the patient-side scene is reconstructed, providing a virtual reality interactive interface for safe remote teleoperation of the robotic arm. Combining omnidirectional augmented reality and virtual reality with robotic telepresence provides clinicians with intuitive visual feedback, thereby improving the safety of robotic telemedicine.
[0065] 2. This invention uses a method based on virtual ArUco codes to calculate the superposition error between virtual and reality, fully considering the errors that may be introduced by each implementation process, and providing a widely applicable reference for error assessment in augmented reality. Attached Figure Description
[0066] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0067] Figure 1 This is a flowchart illustrating the remote presentation method of the present invention;
[0068] Figure 2 This invention primarily embodies the components of a robot telepresence framework and their relative transformation relationships;
[0069] Figure 3 This is a schematic diagram illustrating the three-dimensional reconstruction geometry of the flexible device, which is the main feature of this invention.
[0070] Figure 4 This invention primarily embodies the flowchart of three-dimensional scene reconstruction;
[0071] Figure 5 This is a schematic diagram illustrating the remote presentation of robotic intracavitary intervention based on omnidirectional augmented reality, which is the main feature of this invention. Detailed Implementation
[0072] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0073] like Figure 1 , Figure 2 , Figure 3 , Figure 4 as well as Figure 5 As shown, a robotic intracavitary intervention telepresence method based on augmented reality according to the present invention includes the following steps:
[0074] Step S1: Calibrate the hardware system on the patient side. Specifically, the system is calibrated using multi-target hand-eye calibration and the Iterative Closest Point (ICP) algorithm. This mainly includes calibration between the visualization robotic arm and the electromagnetic tracking system, calibration between the visualization robotic arm and the rigid interventional device, hand-eye calibration, and registration between the electromagnetic tracking system and key anatomical structures.
[0075] It should be noted that the patient-side component of this application mainly includes a robotic arm for manipulation, a robotic arm for visualization, an electromagnetic tracking system, a monocular camera, and interventional devices. The interventional devices include rigid devices and flexible devices. The monocular camera is fixed to the end of the robotic arm for visualization, and the interventional devices are fixed to the end of the robotic arm for manipulation.
[0076] like Figure 1 and Figure 2As shown, to achieve omnidirectional AR functionality, it is necessary to combine the positioning of the robotic arm with the motion of the camera, and to overlay the virtual model and instruments onto the real-world scene. Therefore, it is necessary to transform the reference frame of each part of the system to the camera coordinate system. This mainly involves four steps of calibration or registration: calibration between the visual robotic arm and the electromagnetic tracking system, calibration between the visual robotic arm and the rigid instruments, hand-eye calibration between the monocular camera and the visual robotic arm, and registration between the electromagnetic tracking system and key anatomical structures. All calibration methods are based on AX = XB, and the registration method is the ICP algorithm.
[0077] Specifically, step S1 includes: calibration between the visualization robotic arm and the electromagnetic tracking system, calibration between the visualization robotic arm and the rigid instrument, hand-eye calibration between the monocular camera and the visualization robotic arm, and registration between the electromagnetic tracking system and the anatomical structure.
[0078] The calibration between the visualization robotic arm and the electromagnetic tracking system includes: connecting a six-degree-of-freedom electromagnetic sensor to a rigid flange fixed to the end of the visualization robotic arm; moving the robotic arm; the six-degree-of-freedom electromagnetic sensor moving in the magnetic field of the electromagnetic tracking system; recording the poses of the robotic arm and the six-degree-of-freedom electromagnetic sensor; and constructing the following equations:
[0079]
[0080]
[0081] in, It visualizes the pose of the robotic arm's end effector. It is the transformation matrix from the end effector of the visualized robotic arm to the six-degree-of-freedom electromagnetic sensor. It is the transformation matrix from a visualized robotic arm to an electromagnetic tracking system. It refers to the pose of a six-degree-of-freedom electromagnetic sensor. and In this equation, i represents the dataset index for different visualized robotic arm postures. The above equation can be expressed as:
[0082]
[0083] make and The above equation can then be rewritten as AX = XB, and finally, X can be obtained.
[0084] The calibration between the visualization robot arm and the rigid device includes: fixing a monocular camera to the end effector of the visualization robot arm and fixing an ArUco code to the end effector of the manipulator arm; moving the visualization robot arm and the manipulator arm to move the monocular camera and the ArUco code to different positions, with the ArUco code always within the field of view of the monocular camera; the pose of the ArUco code is estimated by the monocular camera; and the following equation is constructed:
[0085]
[0086]
[0087] in, It is the transformation matrix from the visualization robotic arm to the monocular camera; It is the transformation matrix from the end effector of the visualization robotic arm to the monocular camera, which can be obtained in advance through basic hand-eye calibration. It is the transformation matrix from monocular camera to ArUco code. It is the transformation matrix from visualizing the robotic arm to operating the robotic arm; It is the transformation matrix from the end effector of the robotic arm to the ArUco code; therefore, the above can be expressed as:
[0088]
[0089] make and The transformation matrix between the visualized robotic arm and the operating robotic arm is then transformed into solving for AX = XB.
[0090] The registration of the electromagnetic tracking system and the anatomical structure includes the following steps:
[0091] Step S101: Perform a CT scan on the anatomical structure;
[0092] Step S102: Reconstruct a three-dimensional virtual model of the anatomical structure using multiple CT slices, and use this model as the target point cloud P;
[0093] Step S103: Collect the point cloud of the actual anatomical structure as the source point cloud G;
[0094] Step S104: Use the ICP algorithm to register the target point cloud P and the source point cloud G to obtain the transformation matrix between the electromagnetic tracking system and the anatomical structure.
[0095] More specifically, Figure 2 To remotely present the components of the framework for the robot and their relative transformation relationships. Figure 2 In the diagram, E1 represents a six-degree-of-freedom electromagnetic sensor, RB1 represents a visualization robotic arm, End1 represents the end effector of the visualization robotic arm, C represents a monocular camera, M represents ArUco code, A represents anatomical structure, E represents a six-degree-of-freedom sensor, E3 represents a six-degree-of-freedom sensor, EB represents an electromagnetic tracking system, E2 represents a six-degree-of-freedom sensor, In represents a rigid instrument, FIn represents a flexible instrument, RB2 represents a manipulating robotic arm, and End2 represents the end effector of the manipulating robotic arm.
[0096] During the calibration process between the visualization robotic arm and the electromagnetic tracking system, E1 is fixed to the end effector of the visualization robotic arm via a rigid flange. After calibration, E1 is removed, and the monocular camera is fixed to the end effector of the visualization robotic arm. Similarly, during the calibration process between the visualization robotic arm and the rigid instrument, ArUco codes are fixed to the end effector of the operating robotic arm. After calibration, the ArUco codes are removed, and then the rigid or flexible instrument is fixed to the operating robotic arm according to the surgical requirements. All calibrations and registrations are performed preoperatively, while the shape reconstruction of the flexible robot is performed in real time during the operation.
[0097] The calibration between the visual robotic arm and the electromagnetic tracking system is transformed into a hand-eye calibration problem, primarily solving AX = XB. First, a six-DOF electromagnetic sensor is connected to a rigid flange fixed to the end effector of the visual robotic arm. Then, through telemanipulation of the robotic arm, E1 moves within the magnetic field of the electromagnetic tracking system, and the poses of the visual robotic arm and E1 are recorded. Finally, these poses are used to construct the following equations:
[0098]
[0099] in, It visualizes the pose of the robotic arm's end effector. It is the transformation matrix from {End1} to {E1}. It is the transformation matrix from {RB1} to {EB}. It is the pose of E1. and In this context, i represents the dataset index for different robotic arm postures. Equation 1 can be expressed as:
[0100]
[0101] If let and The above equation can then be rewritten as AX = XB, and finally, X can be obtained.
[0102] During the calibration process between the manipulating arm and the rigid device, since the rigid device is fixed at the end of the manipulating arm, the transformation matrix between {In} and {RB2} can be calculated based on the kinematics of the manipulating arm. Therefore, the calibration between the visual manipulating arm and the rigid device is equivalent to calibrating two manipulating arms. The calibration of two manipulating arms can also be transformed into a hand-eye calibration problem. First, a monocular camera (C) and an ArUco code (M) are fixed at the end of the visual and manipulating manipulating arms, respectively. Then, the camera and the ArUco code are moved to different positions through teleoperation of the two manipulating arms, with the ArUco code always within the camera's field of view. The pose of the ArUco code is estimated by the camera using the Perspective-n-Point (PnP) algorithm. Finally, the poses of the two manipulating arms and the ArUco code can be used to construct the following equations:
[0103]
[0104] in, It is the transformation matrix from {RB1} to {C}; The transformation matrix from {End1} to {C} can be obtained in advance through basic hand-eye calibration. It is the transformation matrix from {C} to {M}. It is the transformation matrix from {RB1} to {RB2}; It is the transformation matrix from {End2} to {M}. Therefore, equation 3 can be expressed as:
[0105]
[0106] If let and The transformation matrix between the visualized robotic arm and the operating robotic arm is then transformed into solving for AX = XB.
[0107] To obtain the 3D pose of the anatomical structure relative to {C}, the electromagnetic tracking system and the anatomical structure need to be registered. The detailed registration process is as follows:
[0108] Perform CT scans on the anatomical structures.
[0109] A three-dimensional virtual model of the anatomical structure was reconstructed using multiple CT slices, and this model was used as the target point cloud P.
[0110] Point clouds of actual anatomical structures were acquired using EM sensors and used as the source point cloud G.
[0111] The ICP algorithm was used to register point clouds P and G to obtain the transformation matrix between the electromagnetic tracking system and the anatomical structure.
[0112] Step S2: Acquire multiple images from monocular cameras in different field-of-view directions to perform 3D reconstruction of the actual scene at the patient end and construct a VR interactive interface. Specifically, twenty images from monocular cameras are acquired from different field-of-view directions through teleoperation of the visual robotic arm, and the SfM and MVS algorithms are used to perform 3D reconstruction of the actual scene at the patient end. The virtual models of the visual robotic arm and the operating robotic arm are imported into the 3D reconstruction results to construct a VR scene, and the joint values of the robotic arm virtual model are updated in real time according to the actual state of the two robotic arms.
[0113] like Figure 2 and Figure 3 As shown, for step 2, which includes the 3D shape reconstruction of the flexible device, two six-degree-of-freedom (DOF) sensors are respectively installed at the root and head of the active bending segment of the flexible device to reconstruct its shape. Specifically, in order to overlay the virtual flexible device model onto the real-world scene, it is necessary to reconstruct the shape of the active bending segment of the flexible device to achieve intuitive remote operation. Two six-DOF EM sensors, namely E2 and E3, are installed at the root and head of the active bending segment, respectively. Figure 2 As shown. The geometry for shape estimation is as follows. Figure 3 As shown. S2 and S3 are the ends of E2 and E3 respectively, p2 and p3 are the position vectors of S2 and S3 respectively, and r2 and r3 are the direction vectors of E2 and E3 respectively. Line segments l2 and l3 coincide with r2 and r3 respectively, which can be represented as:
[0114] l i =p i +r i t i (5)
[0115] Where i = 2 and 3, t i Represents line segment l i From any point on S i The distance between line segments l0 and l2 is given. Line segment l0 is defined as the common perpendicular to l2 and l3, and M2 and M3 are the intersection points of l0 with l2 and l3, respectively. M0 is the midpoint of M2 and M3. and These are the distance parameters of M2 and M3 in Equation 5, respectively.
[0116] Since l2 and l3 may not intersect, the shape reconstruction of the flexible device is based on two assumptions. First, it is assumed that the active bending segment has a constant curvature. Second, it is assumed that the bending plane of the active bending segment lies within the plane defined by points S2, S3, and M0. and The following equation can be obtained:
[0117]
[0118] in, and The coordinates of the intermediate point M0 can be obtained by calculating from Equation 6. Figure 3 In this diagram, S0 is the midpoint of S2S3, O is the center of the arc of the curved segment, and three mutually perpendicular unit vectors (s, n, m) are defined. s is... The unit vector, n is perpendicular to the bending plane, and m is... The unit vectors are calculated as follows:
[0119]
[0120] Based on the above geometric relationships, the arc length can be obtained. The central angles are as follows:
[0121]
[0122] Where R is The radius, as expressed in equation 8, is:
[0123]
[0124] In equation 9, It can be calculated in real time, and can be obtained in advance based on the physical relationship between E2 and E3. To speed up the solution of θ, θ j (j = 1, 2, 3…) Select values between 0 and π at equal intervals and calculate the corresponding values. Finally, calculate the minimum error. When E is at its minimum, θ j This is the solution required for θ.
[0125] Once θ is obtained, the radius And the position vector of the arc center O can be calculated as:
[0126] p0=(p2+p3) / 2+Rcos(θ / 2)m (10)
[0127] Establish a local coordinate system {B} at point O, with unit vectors in the x, y, and z directions, respectively. B y B z B Among them, x B yes The unit vector, z B =n,y B =z B ×x B . Any point in the equation can be considered as S3 orbiting around z in the S2S3O plane. BWe obtain it by rotating γ. Where γ∈[0,θ]. If θ>π, then γ∈[-θ,0]. The mathematical expression for any point in is:
[0128]
[0129] Where p(s) is the position vector under {EB}, and s is the arc parameter.
[0130] Due to modeling and calculation errors, S2 and S3 may not lie on the arc with radius R and center p0. Therefore, according to Equation 11, three points are selected at equal intervals from the arc. Finally, a cubic spline curve is used to fit these three points on the arc, as well as S2 and S3, to obtain the shape p′(s) of the active bending segment of the flexible device.
[0131] For step 2, the actual scene at the patient end is reconstructed in 3D, including sparse reconstruction and dense reconstruction. The incremental SfM algorithm is used for sparse reconstruction, and the MVS algorithm is used for dense reconstruction. The open source library OpenMVS is used to restore the complete surface of the reconstructed scene through surface reconstruction, optimization and texture mapping.
[0132] In sparse reconstruction, the monocular camera pose estimated by the SfM algorithm is represented as: The pose of the visualized robotic arm is The true scale of the monocular camera pose is recovered through a similarity transformation, as shown below:
[0133]
[0134] in, It is a similarity transformation matrix; and They are Translation vectors and rotation matrices; It is the scale of the isotropic scaling transformation, obtained using Umeyama's method. As shown below:
[0135]
[0136] More specifically, to more intuitively teleoperate the visualized robotic arm and determine whether it collides with surrounding objects, a real-world scenario is reconstructed preoperatively at the patient's end, providing a VR view for the operator at the physician's end. The VR view displays the reconstructed scene, the robotic arm, and rigid instruments, with the poses of the robotic arm and instruments updated in real time. Operators can interact with the VR display interface to change their viewing direction, thereby enhancing depth perception and avoiding visual obstruction, thus enabling safe teleoperation of the robotic arm.
[0137] 3D reconstruction comprises two steps: sparse reconstruction and dense reconstruction. Based on the open-source library COLMAP, incremental SfM algorithm is used for sparse reconstruction, and MVS algorithm is used for dense reconstruction. Additionally, the open-source library OpenMVS is used to restore the complete surface of the reconstructed scene through surface reconstruction, optimization, and texture mapping. The 3D scene reconstruction process is as follows: Figure 4 As shown.
[0138] In sparse reconstruction, the estimated camera pose scale differs from the actual pose scale, and the robotic arm itself boasts high positioning accuracy. Therefore, the robotic arm's pose scale is used to recover the camera pose scale estimated by the SfM algorithm. The camera pose estimated by the SfM algorithm is represented as follows: The pose of the visualized robotic arm is The true scale of the camera pose can be recovered through a similarity transformation, as shown below:
[0139]
[0140] in, It is a similarity transformation matrix; and They are Translation vectors and rotation matrices; It is the scale of the isotropic scaling transformation. It is obtained using Umeyama's method. As shown below:
[0141]
[0142] Step S3: Obtain images from different fields of view of the monocular camera and record the status information of multiple robotic arms. Specifically, images from different fields of view of the monocular camera are obtained through teleoperation of the visualized robotic arms, and the status information of the robotic arms under multiple good fields of view is recorded according to the experience and needs of clinicians.
[0143] Step S4: On the monocular camera image, the virtual anatomical structure and virtual interventional instruments are superimposed onto the real scene using calibration results and image processing.
[0144] For step S4, the change matrix for anatomical structure and monocular camera:
[0145]
[0146] in, This refers to the pose of the anatomical structure within the space of a monocular camera. Since the pose of the visualization robotic arm's end effector is updated in real time, the virtual anatomical structure is superimposed onto the real-world scene using the following four steps:
[0147] Step S401: Use an open-source library to create a virtual 3D space containing a virtual camera and a 3D anatomical model;
[0148] Step S402: Set the intrinsic parameters of the virtual camera to be the same as those of the actual monocular camera, and make the reference frame of the virtual camera coincide with the reference frame of the virtual space.
[0149] Step S403, use Update the pose of the 3D anatomical model in the virtual space and use a virtual camera to capture the scene of the virtual space in real time;
[0150] Step S404: Use image processing algorithms to overlay the images captured by the virtual camera and the real monocular camera for enhanced display of anatomical structures;
[0151] For rigid instruments, the transformation matrix from a monocular camera to a rigid instrument is as follows:
[0152]
[0153] in, It is the pose of the rigid instrument relative to the monocular camera, the pose of the visualization robotic arm end effector, and the real-time update of the pose of the manipulating robotic arm end effector.
[0154] For flexible instruments, after shape reconstruction, the shape of the flexible instrument relative to the monocular camera can be obtained as follows:
[0155]
[0156] Where C(s) is the position vector relative to {C} at the arc length s of the flexible instrument, and the pose of the visualized end of the robotic arm, the pose of the operated end of the robotic arm, and p′(s) are updated in real time.
[0157] More specifically, to achieve omnidirectional augmented reality, the monocular camera must be moved to observe objects from different directions. Because the visualization robotic arm has high positioning accuracy, it can effectively compensate for the movement of the monocular camera. After calibration, the transformation matrix from {A} to {C} can be obtained:
[0158]
[0159] This refers to the pose of the anatomical structure in {C} space, updated in real-time by the pose of the visualized robotic arm's end effector. The virtual anatomical structure is overlaid onto the real-world scene using the following four steps:
[0160] Use the open-source library Visualization ToolKit (VTK) to create a virtual 3D space that includes a virtual camera and a 3D anatomical model.
[0161] Set the intrinsic parameters of the virtual camera to be the same as those of the actual monocular camera, and make the reference frame of the virtual camera coincide with the reference frame of the virtual space.
[0162] use Update the pose of the 3D anatomical model in the virtual space and use a virtual camera to capture the scene in the virtual space in real time.
[0163] Images captured by a virtual camera and a real monocular camera are superimposed using image processing algorithms to enhance the display of anatomical structures.
[0164] The enhanced display of the instrument in this invention is applicable to both rigid and flexible instruments. For rigid instruments, the transformation matrix from {C} to {In} is as follows:
[0165]
[0166] in, The pose of the rigid instrument relative to {C} is updated in real time by the poses of the two robotic arms' end effectors. For the flexible instrument, after shape reconstruction, the shape of the flexible instrument relative to {C} is obtained as follows:
[0167]
[0168] Where C(s) is the position vector relative to {C} at the arc length s of the flexible instrument, which is updated in real time by the pose of the end effector and p′(s).
[0169] If the virtual 3D space of the anatomical structure and the virtual 3D space of the instruments are the same, the anatomical structure and the instruments may occlude each other during virtual camera imaging. Therefore, another virtual 3D space is created for the enhanced display of the instruments. Finally, the virtual instruments are overlaid onto the real scene using the methods and steps for enhancing the display of anatomical structures.
[0170] More specifically, this application proposes a method for evaluating the accuracy of virtual and real image overlay. The sources of overlay error mainly include calibration, registration, and image fusion. Since the final overlay effect is displayed on a two-dimensional image plane, it is difficult to estimate the three-dimensional overlay error. Therefore, we have developed a method based on virtual ArUco codes to evaluate the overlay accuracy.
[0171] A cube with an ArUco code attached to each face is used as the real object, and a 3D virtual cube model with an ArUco code image on each face is used as the virtual object. The real cube and the electromagnetic tracking system are registered using the registration method described above to obtain the cube's pose under the electromagnetic tracking system. Then, Equation 14 is used to transform the cube's pose from {EB} to {C} to obtain... Finally, the cube is enhanced using the instrument-enhanced display method described above.
[0172] Move the monocular camera to different positions, keeping the actual cube within the camera's field of view, and capture images of the virtual cube before and after it is superimposed. Using the acquired images, estimate the 3D poses of the virtual and actual cubes using the PnP algorithm. The superposition error can be calculated using the following formula:
[0173]
[0174] Where E overlay This is the 3D average error of the virtual and real images superimposed, where n is the number of image groups acquired. Each group includes the position vectors of the eight corner points of the virtual cube estimated using the PnP algorithm and the position vectors of the eight corner points of the actual cube. Simultaneously, m = 8 represents the number of corner points of the cube. Let {C} be the position vector data of the j-th corner point of the three-dimensional virtual cube in the i-th group of {C}. Let {C} be the position vector data of the j-th corner point of the 3D real cube in the i-th group. Since the pose of the virtual cube superimposed on the real scene is obtained after calibration and registration, and the captured images are acquired from different directions, the developed superposition accuracy estimation method fully considers the errors caused by calibration, registration, and image fusion in the omnidirectional augmented reality implementation process.
[0175] Step S5: The doctor's visualization device remotely acquires the VR and AR interfaces of the patient's device and changes the field of view direction through remote operation or teaching. Specifically, the visual feedback provided from the patient's end to the clinician's end includes AR view, VR view, and endoscopic view. In the AR view, key anatomical structures of the patient and instruments inserted into the patient's body can be visualized, and different AR views with different field of view directions can be obtained through the movement of the robotic arm. By interacting with the display interface, the feedback field of view direction of VR can be changed, thereby determining the safe distance between the visualization robotic arm and surrounding objects, ensuring the safety of robotic arm operation; the movement of the robotic arm can be completed by remote operation by the doctor or by teaching the robotic arm movement using the posture information recorded in step S3.
[0176] More specifically, the TeamView software allows doctors to log into remote patient computers to view AR, VR, and endoscopic views. The visual feedback provided from the patient to the clinician includes AR, VR, and endoscopic views. In the AR view, key anatomical structures and instruments inserted into the patient's body are visible, and different viewing directions can be obtained by observing the movement of the robotic arm. Interaction with the display interface allows the VR feedback view to be changed, determining the safe distance between the robotic arm and surrounding objects, ensuring the safety of teleoperation. The AR view is generated by rendering images from a monocular camera, while the endoscopic view is an image from a camera fixed to the end of the interventional instrument. The movement of the robotic arm and interventional instruments is controlled remotely by the clinician via a handle, while the movement of the robotic arm is controlled remotely by the clinician via a handle or according to a pre-planned trajectory. Movement commands from the doctor's end are sent to the remote patient via TCP.
[0177] It is important to note that the physician's end can be configured with one surgical operator and one assistant. The operator remotely manipulates the interventional instruments with feedback from three views, while the assistant adjusts the AR and VR field of view according to the physician's requirements. Furthermore, the AR field of view can also be programmed based on pre-operatively recorded robotic arm status information.
[0178] This invention also provides an augmented reality-based robotic intracavitary intervention telepresence system, comprising the following modules:
[0179] Module M1 is used to calibrate the hardware system at the patient end.
[0180] Module M2 is used to acquire multiple images from a monocular camera in different fields of view, perform 3D reconstruction of the actual scene at the patient's end, and build a VR interactive interface.
[0181] Module M3 is used to move the monocular camera to obtain multiple viewing directions and record the corresponding robotic arm posture information.
[0182] Module M4 is used to overlay virtual anatomical structures and virtual interventional instruments onto a real scene using calibration results and image processing on a monocular camera image, thereby constructing an AR feedback interface.
[0183] Module M5 is a visualization device used by doctors to remotely acquire the VR and AR interfaces of the patient's device and change the field of view through remote operation or teaching.
[0184] This application fixes a monocular camera to the end effector of a multi-degree-of-freedom robotic arm and combines the motion characteristics of the monocular camera and the robotic arm to achieve omnidirectional augmented reality functionality. Through multi-target hand-eye calibration, iterative nearest point, and image overlay algorithms, virtual key anatomical structures and interventional instruments are superimposed onto the camera's video images. The field of view of the monocular camera can be changed by the movement of the robotic arm, allowing observation and perspective of key anatomical structures and interventional instruments from different directions, reconstructing the scene at the patient's end. This provides a virtual reality interactive interface for the safe remote teleoperation of the robotic arm. Furthermore, a method based on virtual ArUco codes is proposed to calculate the superposition error between virtual and reality. The method proposed in this application combines omnidirectional augmented reality and virtual reality with robotic telepresence, providing clinicians with intuitive visual feedback and thus promoting the safety of robotic telemedicine.
[0185] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0186] In the description of this application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0187] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A method for remote presentation of robotic intracavitary intervention based on augmented reality, characterized in that, The presentation method includes the following steps: Step S1: Calibrate the hardware system on the patient's end; Step S2: Collect multiple images from a monocular camera in different fields of view, perform 3D reconstruction of the actual scene at the patient's end, and build a VR interactive interface; Step S3: Move the monocular camera to obtain multiple viewing directions and record the corresponding robotic arm posture information; Step S4: On the monocular camera image, the virtual anatomical structure and virtual interventional instruments are superimposed onto the real scene using calibration results and image processing to construct an AR feedback interface; Step S5: The doctor's visualization device remotely acquires the VR and AR interfaces of the patient's device and changes the field of view through remote operation or teaching. For step S2, the actual scene at the patient end is reconstructed in three dimensions, including sparse reconstruction and dense reconstruction. The incremental SfM algorithm is used for sparse reconstruction, and the MVS algorithm is used for dense reconstruction. The open source library OpenMVS is used to restore the complete surface of the reconstructed scene through surface reconstruction, optimization and texture mapping. In sparse reconstruction, the monocular camera pose estimated by the SfM algorithm is represented as: The pose of the visualized robotic arm is The true scale of the monocular camera pose is recovered through a similarity transformation, as shown below: in, It is a similarity transformation matrix; and They are Translation vectors and rotation matrices; It is the scale of the isotropic scaling transformation, obtained using Umeyama's method. As shown below: 。 2. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 1, characterized in that, Step S1 includes: calibration between the visualization robotic arm and the electromagnetic tracking system, calibration between the visualization robotic arm and the rigid instrument, hand-eye calibration between the monocular camera and the visualization robotic arm, and registration between the electromagnetic tracking system and the anatomical structure.
3. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 2, characterized in that, The calibration between the visualization robotic arm and the electromagnetic tracking system includes: connecting a six-degree-of-freedom electromagnetic sensor to a rigid flange fixed to the end of the visualization robotic arm; moving the robotic arm; the six-degree-of-freedom electromagnetic sensor moving in the magnetic field of the electromagnetic tracking system; recording the poses of the robotic arm and the six-degree-of-freedom electromagnetic sensor; and constructing the following equations: in, It visualizes the pose of the robotic arm's end effector. It is the transformation matrix from the end effector of the visualized robotic arm to the six-degree-of-freedom electromagnetic sensor. It is the transformation matrix from a visualized robotic arm to an electromagnetic tracking system. It is the pose of a six-degree-of-freedom electromagnetic sensor; in and Among them , i These are the dataset indices for different visualized robotic arm poses, and the above equation is expressed as: make , ,and Then the above equation is rewritten as AX=XB, and finally, X is obtained.
4. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 2, characterized in that, The calibration between the visualization robot arm and the rigid device includes: fixing a monocular camera to the end effector of the visualization robot arm and fixing an ArUco code to the end effector of the manipulator arm; moving the visualization robot arm and the manipulator arm to move the monocular camera and the ArUco code to different positions, with the ArUco code always within the field of view of the monocular camera; the pose of the ArUco code is estimated by the monocular camera; and the following equation is constructed: in, It is the transformation matrix from the visualization robotic arm to the monocular camera; It is the transformation matrix from the end effector of the visualization robotic arm to the monocular camera, which can be obtained in advance through basic hand-eye calibration. It is the transformation matrix from monocular camera to ArUco code. It is the transformation matrix from visualizing the robotic arm to operating the robotic arm; It is the transformation matrix from the end effector of the robotic arm to the ArUco code; therefore, the above can be expressed as: make , ,and Then the transformation matrix of the visualized robotic arm and the operating robotic arm is transformed into solving for AX=XB; The registration of the electromagnetic tracking system and the anatomical structure includes the following steps: Step S101: Perform a CT scan on the anatomical structure; Step S102: Reconstruct a three-dimensional virtual model of the anatomical structure using multiple CT slices, and use this model as the target point cloud P; Step S103: Collect the point cloud of the actual anatomical structure as the source point cloud G; Step S104: Use the ICP algorithm to register the target point cloud P and the source point cloud G to obtain the transformation matrix between the electromagnetic tracking system and the anatomical structure. .
5. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 1, for step S2, includes three-dimensional shape reconstruction of the flexible instrument, by installing two six-degree-of-freedom sensors at the root and head of the active bending segment of the flexible instrument, respectively, to reconstruct the shape of the active bending segment of the flexible instrument.
6. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 1, characterized in that, For step S4, the change matrix for anatomical structure and monocular camera: in, This refers to the pose of the anatomical structure within the space of a monocular camera. Since the pose of the visualization robotic arm's end effector is updated in real time, the virtual anatomical structure is superimposed onto the real-world scene using the following four steps: Step S401: Use an open-source library to create a virtual 3D space containing a virtual camera and a 3D anatomical model; Step S402: Set the intrinsic parameters of the virtual camera to be the same as those of the actual monocular camera, and make the reference frame of the virtual camera coincide with the reference frame of the virtual space. Step S403, use Update the pose of the 3D anatomical model in the virtual space and use a virtual camera to capture the scene of the virtual space in real time; Step S404: Use image processing algorithms to overlay the images captured by the virtual camera and the real monocular camera for enhanced display of anatomical structures; For rigid instruments, the transformation matrix from a monocular camera to a rigid instrument is as follows: in, It is the pose of the rigid instrument relative to the monocular camera, the pose of the visualization robotic arm end effector, and the real-time update of the pose of the manipulating robotic arm end effector. For flexible instruments, after shape reconstruction, the shape of the flexible instrument relative to the monocular camera can be obtained as follows: in, In the arc length of flexible instruments s The position vector relative to {C} is used to visualize the pose of the robotic arm's end effector, manipulate the pose of the robotic arm's end effector, and... Updated in real time.
7. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 1, characterized in that, A cube with an ArUco code attached to each face is used as the actual object, and a 3D virtual cube model with an ArUco code image on each face is used as the virtual object. The actual cube and the electromagnetic tracking system are registered to obtain the pose of the cube under the electromagnetic tracking system. The poses of real and virtual objects are transformed from an electromagnetic tracking system to a monocular camera, and an instrumental augmentation display method is used to enhance the display of the cube. Move the monocular camera to different positions while keeping the actual object within the camera's field of view, and capture camera images before and after the virtual object is superimposed. Using the acquired images, estimate the 3D poses of the virtual and actual objects respectively. The superposition error can be calculated using the following formula: in It is the average error of three dimensions, which is a superposition of virtual and reality. n This refers to the number of image groups acquired. Each group includes the position vectors of the eight corner points of the virtual object estimated using an algorithm, and the position vectors of the eight corner points of the actual object. m =8 represents the number of corner points of the cube. For the third three-dimensional virtual cube j The corner point at {C} is the... i Group location vector data, For the three-dimensional real cube j The corner point at {C} is the... i Group location vector data.
8. The augmented reality-based robotic intracavitary intervention telepresence method as described in claim 1, characterized in that, For step S5, the visual feedback provided from the patient to the clinician includes AR view, VR view and endoscopic view; In the AR view, key anatomical structures of the patient and instruments inserted into the patient's body can be seen through, and different AR views can be obtained by moving the robotic arm. By interacting with the display interface, the direction of the VR feedback field of view can be changed, thereby determining the safe distance between the visualized robotic arm and surrounding objects. The movement of the robotic arm can be remotely controlled by a doctor, or the robotic arm can be taught to move using the posture information recorded in step S3.
9. A robotic intracavitary intervention telepresence system based on augmented reality, characterized in that, The augmented reality-based robotic intracavitary intervention telepresence method according to any one of claims 1-8 includes the following modules: Module M1 is used to calibrate the hardware system at the patient end; Module M2 is used to acquire multiple images from a monocular camera in different fields of view, perform 3D reconstruction of the actual scene at the patient's end, and build a VR interactive interface; Module M3 is used to move the monocular camera to obtain multiple viewing directions and record the corresponding robotic arm posture information; Module M4 is used to overlay virtual anatomical structures and virtual interventional instruments onto a real scene using calibration results and image processing on a monocular camera image, thereby constructing an AR feedback interface; Module M5 is a visualization device used by doctors to remotely acquire the VR and AR interfaces of the patient's device and change the field of view through remote operation or teaching.
Citation Information
Patent Citations
Remote operation system for vascular interventional operation and control method
CN116370092A
Orthopedic surgery robot system based on optical positioning navigation and control method thereof
CN115363773A
System for facilitating navigation of anatomical lumen network and computer readable storage medium
CN116725669A