A surgical robot high-presence visual perception method, system and storage medium

By collecting preoperative three-dimensional images and combining them with binocular cameras and semi-supervised depth estimation models, accurate three-dimensional reconstruction and holographic fusion superposition of organ tissues during surgery are achieved, solving the problem of obstructed spatial vision during laparoscopic surgery and improving the sense of presence and effect of the surgery.

CN120093440BActive Publication Date: 2025-09-16XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510576186.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-16
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In the existing technology, it is difficult for doctors to intuitively grasp the tissue conditions in the obstructed space during laparoscopic surgery, resulting in a lack of high sense of presence during the operation.

Method used

Preoperative 3D images are collected for 3D reconstruction, and the binocular camera is used to capture the real-time scene during the operation. The semi-supervised binocular image depth estimation model is used for depth estimation, and the intraoperative field point cloud is generated and 3D reconstruction is performed. The preoperative 3D model is superimposed with the real-time scene during the operation through holographic projection to achieve highly immersive visual perception.

Benefits of technology

It provides complete information about organ-occluded areas and soft tissue interiors, enhancing the doctor's visual perception and improving the sense of presence and efficiency of surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120093440B_ABST
    Figure CN120093440B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical visual perception technology, and specifically to a highly immersive visual perception method, system, and storage medium for a surgical robot, comprising the following steps: collecting preoperative three-dimensional images, and reconstructing them to obtain a three-dimensional model of the surgical scene; collecting the real-time scene during the operation through a binocular camera to obtain a binocular image of the intraoperative scene, and using a semi-supervised binocular image depth estimation model to perform binocular image depth estimation to obtain an intraoperative scene depth map; generating an intraoperative field point cloud from the intraoperative scene depth map, and reconstructing the three-dimensional model of the intraoperative scene; aligning the three-dimensional model of the intraoperative scene with a three-dimensional model of the surgical scene, and superimposing the aligned three-dimensional model of the surgical scene onto the real-time intraoperative scene through holographic projection to obtain a highly immersive visual scene during the operation. The present invention collects real-life intraoperative scene information and integrates it with preoperative image information to enhance visual perception of the surgical scene, providing doctors with advanced visual perception of organ occlusion areas and complete information within soft tissues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical visual perception technology, and in particular to a high-presence visual perception method, system and storage medium for a surgical robot. Background Art

[0002] During laparoscopic surgery, doctors operating surgical robots use live-action images transmitted to a monitor to understand the real-time surgical scene and guide the use of surgical instruments. Although the live-action images returned by the monitor can reflect the actual surgical scene, the limited field of view of the endoscope and the occlusion caused by stacked tissues make it difficult to understand the internal conditions of the tissues in the obscured space. Therefore, the doctor's intraoperative judgment must rely on his or her experience. Therefore, the effectiveness of the surgery requires extremely high technical skills of the doctor. Even so, subjective judgment cannot fully guarantee the effectiveness and efficiency of the surgery.

[0003] To achieve real-time surgical scene perception during laparoscopic surgery, existing technologies combine laparoscopic visual perception with mechanical perception to achieve a multi-dimensional perception of the real-time scene and enhance the surgeon's perception of the surgical environment. However, mechanical perception of tissue occlusion still makes it difficult to intuitively grasp the tissue conditions within the occluded space, making it difficult to achieve a high sense of presence during the operation. Summary of the Invention

[0004] The purpose of the present invention is to provide a highly immersive visual perception method, system and storage medium for a surgical robot, so as to solve the technical problems in the prior art that it is difficult to intuitively grasp the tissue conditions in an obstructed space and to achieve a highly immersive sense of presence during the operation.

[0005] In order to solve the above technical problems, the present invention specifically provides the following technical solutions:

[0006] A method for highly telepresence visual perception of a surgical robot comprises the following steps:

[0007] Collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene;

[0008] The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene. Based on the binocular image of the intraoperative scene, a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation to obtain a depth map of the intraoperative scene.

[0009] The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene;

[0010] The intraoperative three-dimensional model of the scene is registered with the three-dimensional model of the surgical scene, and the registered three-dimensional model of the surgical scene is superimposed on the real-time intraoperative scene through holographic projection to obtain a highly immersive visual scene during the operation.

[0011] As a preferred solution of the present invention, the method for constructing the semi-supervised binocular image depth estimation model includes:

[0012] The left and right eye images in the binocular image of the intraoperative scene are used to perform bidirectional disparity estimation using a two-branch convolutional neural network to obtain the left and right bidirectional disparity;

[0013] Establishing a reconstruction loss and a consistency loss for left-right bidirectional disparity, and performing bidirectional self-supervised training on the two-branch convolutional neural network based on the reconstruction loss and the consistency loss for left-right bidirectional disparity to obtain a semi-supervised binocular image depth estimation model;

[0014] The semi-supervised binocular image depth estimation model is:

[0015] ;

[0016] ;

[0017] Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network;

[0018] The reconstruction loss is:

[0019] ;

[0020] Where, To rebuild the losses, is the left disparity of the i-th binocular image in the dataset used to train the semi-supervised binocular image depth estimation model, is the left eye image of the i-th binocular image in the dataset, is the right eye image of the i-th binocular image in the dataset, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right eye image is for pass The reconstructed left eye image, n is the total number of binocular images in the dataset;

[0021] The consistency loss is:

[0022] ;

[0023] Where, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for parallax gradient.

[0024] As a preferred solution of the present invention, the method for constructing the intraoperative scene depth map includes:

[0025] Input the binocular image of the intraoperative scene into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the binocular image of the intraoperative scene;

[0026] Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map;

[0027] The first intraoperative scene depth map is: ;

[0028] Where h1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax;

[0029] Obtain the baseline distance and focal length of the binocular camera, and convert the right parallax in the left and right bidirectional parallax to obtain the second intraoperative scene depth map;

[0030] The second intraoperative scene depth map is: ;

[0031] Where h2 is the depth map of the second intraoperative scene, b is the baseline distance, f is the focal length, is the right parallax.

[0032] As a preferred solution of the present invention, the method for generating the intraoperative field view cloud includes:

[0033] Get the camera intrinsic parameters of the binocular camera;

[0034] Perform point cloud conversion using the camera intrinsic parameters and the first intraoperative scene depth map to obtain the first intraoperative field point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud;

[0035] in, , , ;

[0036] Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth value at

[0037] The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative field point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud;

[0038] in, , , ;

[0039] Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth value at .

[0040] As a preferred embodiment of the present invention, the method for reconstructing the three-dimensional model of the intraoperative scene includes:

[0041] The ICP point cloud registration method is used to align the first and second intraoperative scene cloud to obtain the intraoperative scene registration point cloud.

[0042] Performing three-dimensional reconstruction on the intraoperative scene registration point cloud using a Poisson reconstruction method to obtain a three-dimensional model of the intraoperative scene;

[0043] The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

[0044] As a preferred solution of the present invention, the network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same.

[0045] As a preferred embodiment of the present invention, the present invention provides a surgical robot high telepresence visual perception system, which is applied to a surgical robot high telepresence visual perception method. The system includes:

[0046] A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene;

[0047] A depth estimation unit is used to capture the real-time scene during surgery through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene;

[0048] An intraoperative reconstruction unit is used to generate an intraoperative field point cloud from the intraoperative scene depth map, and to perform 3D reconstruction based on the intraoperative field point cloud to obtain a 3D model of the intraoperative scene;

[0049] The preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model;

[0050] The holographic projection unit is used to superimpose the registered three-dimensional model of the surgical scene onto the real-time scene during the operation through holographic projection, thereby obtaining a highly immersive visual scene during the operation.

[0051] As a preferred solution of the present invention, the semi-supervised binocular image depth estimation model is:

[0052] ;

[0053] ;

[0054] Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network.

[0055] As a preferred solution of the present invention, the three-dimensional model of the surgical scene is obtained by three-dimensional reconstruction using 3D Slicer software.

[0056] As a preferred embodiment of the present invention, the present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, a high-presence visual perception method for a surgical robot is implemented.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] The present invention collects intraoperative real-scene information and integrates it with preoperative image information, enhances the visual perception of the surgical scene, develops a semi-supervised binocular image depth estimation model, realizes accurate three-dimensional reconstruction of the organ surface during surgery, and utilizes preoperative and intraoperative organ surface registration to realize holographic fusion and superposition of preoperative three-dimensional images and intraoperative endoscopic images, providing augmented reality visual perception and providing doctors with advanced visual perception of organ occlusion areas and complete information inside soft tissues. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other implementation drawings based on the provided drawings without inventive effort.

[0060] Figure 1 A flow chart of a highly telepresence visual perception method for a surgical robot provided by an embodiment of the present invention;

[0061] Figure 2 A block diagram of a highly telepresence visual perception system for a surgical robot according to an embodiment of the present invention;

[0062] Figure 3 Schematic diagram of the structure of a semi-supervised binocular image depth estimation model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0064] like Figure 1 As shown, the present invention provides a method for highly telepresence visual perception of a surgical robot, comprising the following steps:

[0065] Collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene;

[0066] The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene. Based on the binocular image of the intraoperative scene, a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation to obtain a depth map of the intraoperative scene.

[0067] The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene;

[0068] The intraoperative three-dimensional model of the scene is registered with the three-dimensional model of the surgical scene, and the registered three-dimensional model of the surgical scene is superimposed on the real-time intraoperative scene through holographic projection to obtain a highly immersive visual scene during the operation.

[0069] The present invention first obtains three-dimensional images of organ tissues before surgery through CT imaging technology, and then reconstructs the preoperative organ tissues in three dimensions through three-dimensional reconstruction technology to obtain a three-dimensional model of the organ tissues. The surgical operation object is the organ tissue, and the three-dimensional model of the organ tissue is also the three-dimensional model of the surgical scene. Therefore, the three-dimensional model of the surgical scene panoramically displays the organizational structure and morphology of the organ tissue, so that the organizational structure of the organ tissue can be grasped through the three-dimensional model of the surgical scene.

[0070] The present invention uses a binocular camera loaded in a laparoscope to capture the real scene during surgery, and then performs stereoscopic image depth estimation based on the binocular image of the real scene during surgery obtained by the binocular camera to obtain a real scene depth map during surgery, which is then converted into real scene point cloud data, and finally a three-dimensional model of the organ tissue during surgery is obtained through point cloud three-dimensional reconstruction.

[0071] The present invention aligns the three-dimensional model of intraoperative organ tissue with the three-dimensional model of preoperative organ tissue (the three-dimensional model of the surgical scene), and converts the three-dimensional model of preoperative organ tissue into the three-dimensional model space of intraoperative organ tissue, providing a guarantee for holographic projection.

[0072] The present invention uses holographic projection to project the three-dimensional model of preoperative organ tissue into the real-time scene during surgery, realizing the holographic fusion and superposition of the preoperative three-dimensional image and the intraoperative endoscopic image, providing augmented reality visual perception, and providing doctors with advanced visual perception of organ occlusion areas and complete information inside soft tissues.

[0073] When performing stereo image depth estimation on real-scene binocular images during surgery, the present invention adopts a semi-supervised binocular image depth estimation model to achieve it. The semi-supervised binocular image depth estimation model includes a dual-branch convolutional neural network and bidirectional adaptive supervision to ensure that accurate left and right disparity can be obtained based on the binocular images.

[0074] Specifically, the dual-branch convolutional neural network includes a first-branch convolutional neural network and a second-branch convolutional neural network. The first-branch convolutional neural network is used to predict the left disparity of the left-eye image relative to the right-eye image, and the second-branch convolutional neural network is used to predict the right disparity of the right-eye image relative to the left-eye image.

[0075] Bidirectional adaptive supervision includes bidirectional supervision of reconstruction loss and bidirectional supervision of consistency loss. In the bidirectional supervision of reconstruction loss, one direction of supervision is the loss between the original left-eye image and the right-eye image obtained by reconstructing the left disparity of the left-eye image relative to the right-eye image (reconstruction loss of left disparity). The other direction of supervision is the loss between the original right-eye image and the original left-eye image obtained by reconstructing the right disparity of the right-eye image relative to the left-eye image (reconstruction loss of right disparity).

[0076] The bidirectional supervision of the reconstruction loss is quantified by the reconstruction loss between the original left image and the original right image. The original left image and the original right image are the label values ​​of the reconstruction loss of the two-branch convolutional neural network, and the label value is the true value. Therefore, the bidirectional supervision of the reconstruction loss belongs to supervised training with accurate supervision truth value for the two-branch convolutional neural network.

[0077] One direction of supervision in the two-way supervision of disparity consistency loss is the loss between the left disparity predicted by the first branch neural network, the right disparity reconstructed by the left disparity predicted by the first branch neural network, and the right disparity predicted by the second branch neural network (left-right disparity consistency loss). The other direction of supervision is the loss between the right disparity predicted by the second branch neural network, the left disparity reconstructed by the right disparity predicted by the second branch neural network, and the left disparity predicted by the first branch neural network (right-left disparity consistency loss).

[0078] The bidirectional supervision of disparity consistency loss is quantified by the reconstruction loss between the left disparity predicted by the first branch neural network and the right disparity predicted by the second branch neural network. The left disparity and the right disparity are the label values ​​of the consistency loss of the two-branch convolutional neural network, but the label value is a predicted value. Therefore, the bidirectional supervision of the reconstruction loss belongs to unsupervised training of the two-branch convolutional neural network without supervised truth value.

[0079] The combination of reconstruction loss and consistency loss enables the combination of supervised and unsupervised training, resulting in a semi-supervised two-branch convolutional neural network. This constructs a semi-supervised binocular image depth estimation model, which can obtain accurate left and right disparity estimates. The disparity can be directly converted into depth, thus enabling accurate stereo depth estimation.

[0080] The present invention also superimposes the disparity gradient in the disparity consistency loss, thereby ensuring the left and right disparity consistency with the highest disparity value obtained after training is completed, while also having the smallest disparity gradient, so that the obtained disparity map has the highest density and the disparity is kept smooth locally.

[0081] After obtaining the left and right bidirectional parallax, the present invention converts it into two corresponding real-world point cloud data, and obtains an accurate three-dimensional model of the intraoperative organ through point cloud registration and optimization to construct the surface point cloud of the intraoperative organ.

[0082] The present invention uses a semi-supervised binocular image depth estimation model to perform stereo depth estimation on real-scene binocular images during surgery. The semi-supervised binocular image depth estimation model includes a two-branch convolutional neural network and bidirectional adaptive supervision to ensure that accurate left and right disparity can be obtained based on the binocular images. The details are as follows:

[0083] The method for constructing a semi-supervised binocular image depth estimation model includes:

[0084] The left and right eye images in the binocular image of the intraoperative scene are used to perform bidirectional disparity estimation using a two-branch convolutional neural network to obtain the left and right bidirectional disparity;

[0085] A reconstruction loss and consistency loss for left-right bidirectional disparity are established, and a two-branch convolutional neural network is trained in a bidirectional self-supervised manner based on the reconstruction loss and consistency loss for left-right bidirectional disparity to obtain a semi-supervised binocular image depth estimation model.

[0086] like Figure 3 As shown in Figure 2, the semi-supervised binocular image depth estimation model is:

[0087] ;

[0088] ;

[0089] Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network;

[0090] The reconstruction loss is:

[0091] ;

[0092] Where, To rebuild the losses, is the left disparity of the i-th binocular image in the dataset used to train the semi-supervised binocular image depth estimation model, is the left eye image of the i-th binocular image in the dataset, is the right eye image of the i-th binocular image in the dataset, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right eye image is for pass The reconstructed left eye image, n is the total number of binocular images in the dataset;

[0093] Figure 3 middle, , refers to the original left eye image By left parallax Reconstructed right eye image , , refers to the original right eye image By right parallax Reconstructed left eye image .

[0094] The dual-branch convolutional neural network includes a first-branch convolutional neural network and a second-branch convolutional neural network. The first-branch convolutional neural network is used to predict the left disparity of the left-eye image relative to the right-eye image, and the second-branch convolutional neural network is used to predict the right disparity of the right-eye image relative to the left-eye image.

[0095] Bidirectional adaptive supervision includes bidirectional supervision of reconstruction loss and bidirectional supervision of consistency loss. In the bidirectional supervision of reconstruction loss, one direction of supervision is the loss between the original left-eye image and the right-eye image obtained by reconstructing the left disparity of the left-eye image relative to the right-eye image (reconstruction loss of left disparity). The other direction of supervision is the loss between the original right-eye image and the original left-eye image obtained by reconstructing the right disparity of the right-eye image relative to the left-eye image (reconstruction loss of right disparity).

[0096] The bidirectional supervision of the reconstruction loss is quantified by the reconstruction loss between the original left image and the original right image. The original left image and the original right image are the label values ​​of the reconstruction loss of the two-branch convolutional neural network, and the label value is the true value. Therefore, the bidirectional supervision of the reconstruction loss belongs to supervised training with accurate supervision truth value for the two-branch convolutional neural network.

[0097] The consistency loss is:

[0098] ;

[0099] Where, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for parallax gradient.

[0100] Figure 3 middle, , refers to the left parallax By left parallax Reconstructed right disparity , , refers to the right parallax By right parallax Reconstructed left disparity .

[0101] One direction of supervision in the two-way supervision of disparity consistency loss is the loss between the left disparity predicted by the first branch neural network, the right disparity reconstructed by the left disparity predicted by the first branch neural network, and the right disparity predicted by the second branch neural network (left-right disparity consistency loss). The other direction of supervision is the loss between the right disparity predicted by the second branch neural network, the left disparity reconstructed by the right disparity predicted by the second branch neural network, and the left disparity predicted by the first branch neural network (right-left disparity consistency loss).

[0102] The bidirectional supervision of disparity consistency loss is quantified by the reconstruction loss between the left disparity predicted by the first branch neural network and the right disparity predicted by the second branch neural network. The left disparity and the right disparity are the label values ​​of the consistency loss of the two-branch convolutional neural network, but the label value is a predicted value. Therefore, the bidirectional supervision of the reconstruction loss belongs to unsupervised training of the two-branch convolutional neural network without supervised truth value.

[0103] The present invention also superimposes the disparity gradient in the disparity consistency loss, thereby ensuring the left and right disparity consistency with the highest disparity value obtained after training is completed, while also having the smallest disparity gradient, so that the obtained disparity map has the highest density and the disparity is kept smooth locally.

[0104] The combination of reconstruction loss and consistency loss enables the combination of supervised and unsupervised training, resulting in a semi-supervised two-branch convolutional neural network. This constructs a semi-supervised binocular image depth estimation model, which can obtain accurate left and right disparity estimates. The disparity can be directly converted into depth, thus enabling accurate stereo depth estimation.

[0105] After obtaining the left and right bidirectional parallax, the present invention converts it into two corresponding real-world point cloud data. By registering and optimizing the surface point cloud of the intraoperative organ, an accurate three-dimensional model of the intraoperative organ is obtained, as follows:

[0106] The method for constructing the intraoperative scene depth map includes:

[0107] Input the binocular image of the intraoperative scene into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the binocular image of the intraoperative scene;

[0108] Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map;

[0109] The depth map of the first intraoperative scene is: ;

[0110] Where h1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax;

[0111] Obtain the baseline distance and focal length of the binocular camera, and convert the right parallax in the left and right bidirectional parallax to obtain the second intraoperative scene depth map;

[0112] The depth map of the second intraoperative scene is: ;

[0113] Where h2 is the depth map of the second intraoperative scene, b is the baseline distance, f is the focal length, is the right parallax.

[0114] The methods for generating the scenic spot cloud during the operation include:

[0115] Get the camera intrinsic parameters of the binocular camera;

[0116] Perform point cloud conversion using the camera intrinsic parameters and the first intraoperative scene depth map to obtain the first intraoperative field point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud;

[0117] in, , , ;

[0118] Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth value at

[0119] The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative field point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud;

[0120] in, , , ;

[0121] Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth of the .

[0122] Methods for reconstructing the 3D model of the intraoperative scene include:

[0123] The ICP point cloud registration method is used to align the first and second intraoperative scene cloud to obtain the intraoperative scene registration point cloud.

[0124] The Poisson reconstruction method is used to reconstruct the intraoperative scene registration point cloud into three dimensions to obtain a three-dimensional model of the intraoperative scene;

[0125] The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

[0126] The network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same.

[0127] like Figure 2 As shown, the present invention provides a high telepresence visual perception system for a surgical robot, which is applied to a high telepresence visual perception method for a surgical robot. The system includes:

[0128] A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene;

[0129] A depth estimation unit is used to capture the real-time scene during surgery through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene;

[0130] An intraoperative reconstruction unit is used to generate an intraoperative field point cloud from the intraoperative scene depth map, and to perform 3D reconstruction based on the intraoperative field point cloud to obtain a 3D model of the intraoperative scene;

[0131] The preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model;

[0132] The holographic projection unit is used to superimpose the registered three-dimensional model of the surgical scene onto the real-time scene during the operation through holographic projection, thereby obtaining a highly immersive visual scene during the operation.

[0133] The semi-supervised binocular image depth estimation model is:

[0134] ;

[0135] ;

[0136] Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network.

[0137] The three-dimensional model of the surgical scene was reconstructed using 3D Slicer software.

[0138] The present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, a high-presence visual perception method for a surgical robot is implemented.

[0139] The present invention collects intraoperative real-scene information and integrates it with preoperative image information, enhances the visual perception of the surgical scene, develops a semi-supervised binocular image depth estimation model, realizes accurate three-dimensional reconstruction of the organ surface during surgery, and utilizes preoperative and intraoperative organ surface registration to realize holographic fusion and superposition of preoperative three-dimensional images and intraoperative endoscopic images, providing augmented reality visual perception and providing doctors with advanced visual perception of organ occlusion areas and complete information inside soft tissues.

[0140] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A surgical robot high-presence visual perception system, characterized in that the system include: A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; A depth estimation unit is used to capture the real-time scene during surgery through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene; An intraoperative reconstruction unit is used to generate an intraoperative field point cloud from the intraoperative scene depth map, and to perform 3D reconstruction based on the intraoperative field point cloud to obtain a 3D model of the intraoperative scene; The preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model; The holographic projection unit is used to superimpose the registered 3D model of the surgical scene onto the real-time scene during the operation through holographic projection, thereby obtaining a highly immersive visual scene during the operation. The method for constructing the semi-supervised binocular image depth estimation model includes: The left and right eye images in the binocular image of the intraoperative scene are used to perform bidirectional disparity estimation using a two-branch convolutional neural network to obtain the left and right bidirectional disparity; Establishing a reconstruction loss and a consistency loss for left-right bidirectional disparity, and performing bidirectional self-supervised training on the two-branch convolutional neural network based on the reconstruction loss and the consistency loss for left-right bidirectional disparity to obtain a semi-supervised binocular image depth estimation model; The semi-supervised binocular image depth estimation model is: ; ; Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network; The reconstruction loss is: ; Where, To rebuild the losses, is the left disparity of the i-th binocular image in the dataset used to train the semi-supervised binocular image depth estimation model, is the left eye image of the i-th binocular image in the dataset, is the right eye image of the i-th binocular image in the dataset, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right eye image is for pass The reconstructed left eye image, n is the total number of binocular images in the dataset; The consistency loss is: ; Where, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for parallax gradient.

2. The surgical robot high telepresence visual perception system according to claim 1, characterized in that: The method for constructing the intraoperative scene depth map includes: Input the binocular image of the intraoperative scene into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the binocular image of the intraoperative scene; Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map; The first intraoperative scene depth map is: ; Where h1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax; Obtain the baseline distance and focal length of the binocular camera, and convert the right parallax in the left and right bidirectional parallax to obtain the second intraoperative scene depth map; The second intraoperative scene depth map is: ; Where h2 is the depth map of the second intraoperative scene, b is the baseline distance, f is the focal length, is the right parallax.

3. The surgical robot high-presence visual perception system according to claim 2, characterized in that: The method for generating the field scenic spot cloud during the operation includes: Get the camera intrinsic parameters of the binocular camera; Perform point cloud conversion using the camera intrinsic parameters and the first intraoperative scene depth map to obtain the first intraoperative field point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth value at The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative field point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth of the .

4. The surgical robot high-presence visual perception system according to claim 3, characterized in that: The method for reconstructing the three-dimensional model of the intraoperative scene includes: The ICP point cloud registration method is used to align the first and second intraoperative scene cloud to obtain the intraoperative scene registration point cloud. Performing three-dimensional reconstruction on the intraoperative scene registration point cloud using a Poisson reconstruction method to obtain a three-dimensional model of the intraoperative scene; The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

5. The surgical robot high-presence visual perception system according to claim 1, characterized in that: The network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same, and the three-dimensional model of the surgical scene is obtained by three-dimensional reconstruction using 3D Slicer software.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, a method for highly immersive visual perception of a surgical robot is implemented, comprising the following steps: Collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene. Based on the binocular image of the intraoperative scene, a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation to obtain a depth map of the intraoperative scene. The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; Registering the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model; The registered 3D model of the surgical scene is superimposed on the real-time scene during the operation through holographic projection, resulting in a highly immersive visual scene during the operation. The method for constructing the semi-supervised binocular image depth estimation model includes: The left and right eye images in the binocular image of the intraoperative scene are used to perform bidirectional disparity estimation using a two-branch convolutional neural network to obtain the left and right bidirectional disparity; Establishing a reconstruction loss and a consistency loss for left-right bidirectional disparity, and performing bidirectional self-supervised training on the two-branch convolutional neural network based on the reconstruction loss and the consistency loss for left-right bidirectional disparity to obtain a semi-supervised binocular image depth estimation model; The semi-supervised binocular image depth estimation model is: ; ; Where, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network; The reconstruction loss is: ; Where, To rebuild the losses, is the left disparity of the i-th binocular image in the dataset used to train the semi-supervised binocular image depth estimation model, is the left eye image of the i-th binocular image in the dataset, is the right eye image of the i-th binocular image in the dataset, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right eye image is for pass The reconstructed left eye image, n is the total number of binocular images in the dataset; The consistency loss is: ; Where, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for parallax gradient.

7. The computer-readable storage medium according to claim 6, wherein: The method for constructing the intraoperative scene depth map includes: Input the binocular image of the intraoperative scene into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the binocular image of the intraoperative scene; Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map; The first intraoperative scene depth map is: ; Where h1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax; Obtain the baseline distance and focal length of the binocular camera, and convert the right parallax in the left and right bidirectional parallax to obtain the second intraoperative scene depth map; The second intraoperative scene depth map is: ; Where h2 is the depth map of the second intraoperative scene, b is the baseline distance, f is the focal length, is the right parallax.

8. The computer-readable storage medium according to claim 7, wherein: The method for generating the field scenic spot cloud during the operation includes: Get the camera intrinsic parameters of the binocular camera; Perform point cloud conversion using the camera intrinsic parameters and the first intraoperative scene depth map to obtain the first intraoperative field point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth value at The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative field point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; Where, Depth map The pixel coordinates in , The main distance of the camera is , are the camera principal point coordinates, Depth map middle The depth of the .

9. The computer-readable storage medium according to claim 8, wherein: The method for reconstructing the three-dimensional model of the intraoperative scene includes: The ICP point cloud registration method is used to align the first and second intraoperative scene cloud to obtain the intraoperative scene registration point cloud. Performing three-dimensional reconstruction on the intraoperative scene registration point cloud using a Poisson reconstruction method to obtain a three-dimensional model of the intraoperative scene; The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

10. The computer-readable storage medium according to claim 7, wherein: The network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same, and the three-dimensional model of the surgical scene is obtained by three-dimensional reconstruction using 3D Slicer software.

Citation Information

Patent Citations

  • Deep learning semi-supervised dense matching method and system based on consistency constraint

    CN113780389A

  • Three-dimensional grid model registration fusion system for laparoscopic surgery navigation

    CN116485851A