Surgical robot high telepresence visual perception method and system and storage medium

By combining preoperative three-dimensional reconstruction and intraoperative binocular image depth estimation, an accurate grasp of the tissue situation in the occluded space is achieved, and the problem of difficult high-perspective visual perception in the prior art is solved, and augmented reality visual perception is provided to help doctors perform surgical operations better.

CN120093440AActive Publication Date: 2025-06-06XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510576186.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to intuitively grasp the organizational situation in the occlusion space, and the sense of precautions in the operation process is difficult to achieve.

Method used

By collecting preoperative three-dimensional images and performing three-dimensional reconstruction, combining binocular camera acquisition intraoperative real-time scenes, using a semi-supervised binocular image depth estimation model for depth estimation, generating an intraoperative scene point cloud and performing three-dimensional reconstruction. Finally, the intraoperative scene three-dimensional model is registered with the preoperative three-dimensional model, and superimposing it into the intraoperative real-time scene through holographic projection.

Benefits of technology

It achieves accurate grasp of the tissue conditions in the occlusion space, enhances the doctor's visual perception of the intraoperative environment, and provides complete information on the organ occlusion area and soft tissue inside.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120093440A_ABST
    Figure CN120093440A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical visual perception, in particular to a surgical robot high-immediacy visual perception method and system and a storage medium, and the method comprises the following steps: collecting a preoperative three-dimensional image, and carrying out the three-dimensional reconstruction to obtain a surgical scene three-dimensional model; the method comprises the following steps: acquiring an intra-operative real-time scene through a binocular camera to obtain an intra-operative scene binocular image, and performing binocular image depth estimation by using a semi-supervised binocular image depth estimation model to obtain an intra-operative scene depth map; generating an intraoperative scene point cloud by using the intraoperative scene depth map, and performing three-dimensional reconstruction to obtain an intraoperative scene three-dimensional model; and registering the intraoperative scene three-dimensional model with the operation scene three-dimensional model, and superposing the registered operation scene three-dimensional model into the intraoperative real-time scene through holographic projection to obtain an intraoperative high-telepresence visual scene. According to the invention, intraoperative live-action information is collected and fused with preoperative image information, so that visual perception of an operation scene is enhanced, and advanced visual perception of complete information of an organ shielding area and the interior of soft tissue is provided for doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical visual perception technology, and in particular to a high-presence visual perception method, system and storage medium for a surgical robot. Background Art

[0002] Doctors control surgical robots during laparoscopic surgery. They use the real-time images transmitted to the monitor to understand the real-time situation of the surgery and guide the operation of surgical instruments. Although the real-time images returned by the monitor can reflect the actual situation of the surgery, it is difficult to understand the internal situation of the tissue in the occluded space due to the limited field of view of the endoscope and the occlusion caused by the stacking of tissues. Therefore, it is necessary to rely on the doctor's experience to make intraoperative judgments. Therefore, the surgical effect has extremely high requirements on the doctor's technical level. Even so, subjective judgments can hardly fully guarantee the surgical effect and efficiency.

[0003] In order to realize the real-time surgical scene perception of laparoscopic surgery by doctors, the existing technology combines mechanical perception with laparoscopic visual perception to achieve multi-dimensional perception of real-time scenes and improve the doctor's perception of the surgical environment. However, mechanical perception of tissue occlusion still makes it difficult to intuitively grasp the tissue situation in the occluded space and achieve a high sense of presence during the operation. Summary of the invention

[0004] The purpose of the present invention is to provide a surgical robot high-presence visual perception method, system and storage medium to solve the technical problems in the prior art that it is difficult to intuitively grasp the tissue conditions in the obstructed space and difficult to achieve a high sense of presence during the operation.

[0005] In order to solve the above technical problems, the present invention specifically provides the following technical solutions: A surgical robot high telepresence visual perception method comprises the following steps: Collect preoperative 3D images, and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene, and a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation based on the binocular image of the intraoperative scene to obtain a depth map of the intraoperative scene; The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; The intraoperative scene 3D model is registered with the surgical scene 3D model, and the registered surgical scene 3D model is superimposed on the intraoperative real-time scene through holographic projection to obtain a highly immersive intraoperative visual scene.

[0006] As a preferred solution of the present invention, the method for constructing the semi-supervised binocular image depth estimation model includes: The left and right images in the binocular image of the intraoperative scene are estimated by using a two-branch convolutional neural network to obtain the left and right bidirectional disparity; Establishing a reconstruction loss and a consistency loss of left-right bidirectional disparity, and performing bidirectional self-supervised training on the dual-branch convolutional neural network based on the reconstruction loss and the consistency loss of left-right bidirectional disparity, to obtain a semi-supervised binocular image depth estimation model; The semi-supervised binocular image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network; The reconstruction loss is: ; In the formula, To rebuild the losses, is the left disparity of the i-th stereo image in the dataset used to train the semi-supervised stereo image depth estimation model, is the left eye image of the i-th binocular image in the data set, is the right eye image of the i-th binocular image in the data set, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right image is for pass The reconstructed left eye image, n is the total number of binocular images in the data set; The consistency loss is: ; In the formula, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for The parallax gradient.

[0007] As a preferred solution of the present invention, the method for constructing the intraoperative scene depth map includes: Input the intraoperative scene binocular image into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the intraoperative scene binocular image; Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map; The first intraoperative scene depth map is: ; In the formula, h 1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax; Obtain the baseline distance and focal length of the binocular camera, and obtain the second intraoperative scene depth map based on the right parallax in the left and right bidirectional parallax; The second intraoperative scene depth map is: ; In the formula, h 2 is the depth map of the scene during the second operation, b is the baseline distance, f is the focal length, is the right parallax.

[0008] As a preferred solution of the present invention, the method for generating a scene cloud of the field during surgery includes: Get the camera internal parameters of the binocular camera; The camera intrinsic parameters and the first intraoperative scene depth map are used to perform point cloud conversion to obtain the first intraoperative scene point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth value at ; The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative scene point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth value at .

[0009] As a preferred embodiment of the present invention, the method for reconstructing the three-dimensional model of the intraoperative scene includes: The ICP point cloud registration method is used to align and register the first intraoperative scene cloud and the second intraoperative scene cloud to obtain the intraoperative scene registration point cloud; Reconstructing the intraoperative scene registration point cloud in three dimensions using a Poisson reconstruction method to obtain a three-dimensional model of the intraoperative scene; The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

[0010] As a preferred solution of the present invention, the network structure of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network is the same.

[0011] As a preferred solution of the present invention, the present invention provides a surgical robot high telepresence visual perception system, which is applied to a surgical robot high telepresence visual perception method, and the system includes: A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; A depth estimation unit is used to collect the real-time scene during the operation through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene; An intraoperative reconstruction unit, used to generate an intraoperative field point cloud from an intraoperative scene depth map, and to perform three-dimensional reconstruction based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; A preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model; The holographic projection unit is used to superimpose the registered three-dimensional model of the surgical scene onto the real-time scene during the operation through holographic projection, so as to obtain a highly immersive visual scene during the operation.

[0012] As a preferred solution of the present invention, the semi-supervised binocular image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network.

[0013] As a preferred solution of the present invention, the three-dimensional model of the surgical scene is obtained by three-dimensional reconstruction using 3D Slicer software.

[0014] As a preferred embodiment of the present invention, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer execution instructions. When a processor executes the computer execution instructions, a high-presence visual perception method for a surgical robot is implemented.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention collects intraoperative real-scene information and integrates it with preoperative image information, strengthens the visual perception of the surgical scene, develops a semi-supervised binocular image depth estimation model, realizes accurate three-dimensional reconstruction of the organ surface during surgery, and utilizes preoperative and intraoperative organ surface registration to realize holographic fusion and superposition of preoperative three-dimensional images and intraoperative endoscopic images, providing augmented reality visual perception, and providing doctors with advanced visual perception of organ occluded areas and complete information inside soft tissues. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0017] Figure 1 A flow chart of a surgical robot high telepresence visual perception method provided by an embodiment of the present invention; Figure 2 A block diagram of a surgical robot high telepresence visual perception system provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a semi-supervised binocular image depth estimation model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] like Figure 1 As shown, the present invention provides a method for high telepresence visual perception of a surgical robot, comprising the following steps: Collect preoperative 3D images, and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene, and a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation based on the binocular image of the intraoperative scene to obtain a depth map of the intraoperative scene; The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; The intraoperative scene 3D model is registered with the surgical scene 3D model, and the registered surgical scene 3D model is superimposed on the intraoperative real-time scene through holographic projection to obtain a highly immersive intraoperative visual scene.

[0020] The present invention firstly obtains a three-dimensional image of the organ tissue before the operation by means of CT imaging technology, and then reconstructs the preoperative organ tissue in three dimensions by means of three-dimensional reconstruction technology, thereby obtaining a three-dimensional model of the organ tissue. The surgical operation object is the organ tissue, and the three-dimensional model of the organ tissue is also the three-dimensional model of the surgical scene. Therefore, the three-dimensional model of the surgical scene panoramically displays the organizational structure and morphology of the organ tissue, so that the organizational structure of the organ tissue can be grasped through the three-dimensional model of the surgical scene.

[0021] The present invention uses a binocular camera loaded in a laparoscope to capture the real scene during surgery, and then performs stereoscopic image depth estimation based on the binocular image of the real scene during surgery obtained by the binocular camera to obtain a real-scene depth map during surgery, which is then converted into real-scene point cloud data, and finally a three-dimensional model of the intraoperative organ tissue is obtained through point cloud three-dimensional reconstruction.

[0022] The present invention aligns the three-dimensional model of intraoperative organ tissue with the three-dimensional model of preoperative organ tissue (the three-dimensional model of the surgical scene), and converts the three-dimensional model of preoperative organ tissue into the three-dimensional model space of intraoperative organ tissue, providing a guarantee for holographic projection.

[0023] The present invention utilizes holographic projection to project the three-dimensional model of preoperative organ tissue into the real-time scene during surgery, realizes the holographic fusion and superposition of the preoperative three-dimensional image and the intraoperative endoscopic image, provides augmented reality visual perception, and provides doctors with advanced visual perception of organ occluded areas and complete information inside soft tissues.

[0024] When the present invention performs stereo image depth estimation on real-scene binocular images during surgery, a semi-supervised binocular image depth estimation model is adopted for implementation, wherein the semi-supervised binocular image depth estimation model includes a dual-branch convolutional neural network and bidirectional adaptive supervision, thereby ensuring that accurate left-right parallax can be obtained based on the binocular image.

[0025] Specifically, the dual-branch convolutional neural network includes a first-branch convolutional neural network and a second-branch convolutional neural network, the first-branch convolutional neural network is used to predict the left disparity of the left-eye image relative to the right-eye image, and the second-branch convolutional neural network is used to predict the right disparity of the right-eye image relative to the left-eye image.

[0026] Bidirectional adaptive supervision includes bidirectional supervision of reconstruction loss and bidirectional supervision of consistency loss. In the bidirectional supervision of reconstruction loss, one direction of supervision is the loss between the right eye image and the original right eye image obtained by reconstructing the original left eye image through the left disparity of the left eye image relative to the right eye image (reconstruction loss of the left disparity), and the other direction of supervision is the loss between the left eye image and the original left eye image obtained by reconstructing the right disparity of the right eye image relative to the left eye image (reconstruction loss of the right disparity).

[0027] The bidirectional supervision of the reconstruction loss is quantified by the reconstruction loss between the original left eye image and the original right eye image. The original left eye image and the original right eye image are the label values ​​of the reconstruction loss of the two-branch convolutional neural network, and the label value is the true value. Therefore, the bidirectional supervision of the reconstruction loss belongs to the supervised training with accurate supervision truth value for the two-branch convolutional neural network.

[0028] One direction of supervision in the two-way supervision of disparity consistency loss is the loss between the left disparity predicted by the first branch neural network, the right disparity reconstructed by the left disparity predicted by the first branch neural network, and the right disparity predicted by the second branch neural network (left-right disparity consistency loss), and the other direction of supervision is the loss between the left disparity predicted by the second branch neural network, the right disparity reconstructed by the right disparity predicted by the second branch neural network, and the left disparity predicted by the first branch neural network (right-left disparity consistency loss).

[0029] The bidirectional supervision of the disparity consistency loss is quantified by the reconstruction loss between the left disparity predicted by the first branch neural network and the right disparity predicted by the second branch neural network. The left disparity and the right disparity are the label values ​​of the consistency loss of the two-branch convolutional neural network, but the label value is a predicted value. Therefore, the bidirectional supervision of the reconstruction loss belongs to the unsupervised training of the two-branch convolutional neural network with unsupervised truth value.

[0030] In the above, the combination of reconstruction loss and consistency loss combines supervised training with unsupervised training, and obtains a semi-supervised training two-branch convolutional neural network, constructing a semi-supervised binocular image depth estimation model, which can obtain accurate left disparity and right disparity estimation. The disparity can be directly converted into depth, so accurate stereo depth estimation can be achieved.

[0031] The present invention also superimposes the disparity gradient in the disparity consistency loss, thereby ensuring the left and right disparity consistency with the highest disparity value obtained after training is completed, and at the same time having the smallest disparity gradient, so that the obtained disparity map has the highest density and the disparity is kept smooth locally.

[0032] After obtaining the left and right bidirectional parallax, the present invention converts it into two corresponding real-world point cloud data, and obtains an accurate three-dimensional model of the intraoperative organ by point cloud registration and optimization to construct the surface point cloud of the intraoperative organ.

[0033] The present invention uses a semi-supervised binocular image depth estimation model to perform stereo image depth estimation on the real-scene binocular image during surgery. The semi-supervised binocular image depth estimation model includes a dual-branch convolutional neural network and bidirectional adaptive supervision to ensure that accurate left and right disparity can be obtained based on the binocular image, as follows: The method for constructing a semi-supervised binocular image depth estimation model includes: The left and right images in the binocular image of the intraoperative scene are estimated by using a two-branch convolutional neural network to obtain the left and right bidirectional disparity; The reconstruction loss and consistency loss of left-right bidirectional disparity are established, and the two-branch convolutional neural network is trained in a bidirectional self-supervised manner based on the reconstruction loss and consistency loss of left-right bidirectional disparity to obtain a semi-supervised binocular image depth estimation model. like Figure 3 As shown in Figure 2, the semi-supervised stereo image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network; The reconstruction loss is: ; In the formula, To rebuild the losses, is the left disparity of the i-th stereo image in the dataset used to train the semi-supervised stereo image depth estimation model, is the left eye image of the i-th binocular image in the dataset, is the right eye image of the i-th binocular image in the dataset, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right image is for pass The reconstructed left eye image, n is the total number of binocular images in the dataset; Figure 3 middle, , refers to the original left-eye image By left parallax Reconstructed right eye image , , refers to the original right eye image By right parallax Reconstructed left image .

[0034] The dual-branch convolutional neural network includes a first-branch convolutional neural network and a second-branch convolutional neural network. The first-branch convolutional neural network is used to predict the left disparity of the left-eye image relative to the right-eye image, and the second-branch convolutional neural network is used to predict the right disparity of the right-eye image relative to the left-eye image.

[0035] Bidirectional adaptive supervision includes bidirectional supervision of reconstruction loss and bidirectional supervision of consistency loss. In the bidirectional supervision of reconstruction loss, one direction of supervision is the loss between the right eye image and the original right eye image obtained by reconstructing the original left eye image through the left disparity of the left eye image relative to the right eye image (reconstruction loss of the left disparity), and the other direction of supervision is the loss between the left eye image and the original left eye image obtained by reconstructing the right disparity of the right eye image relative to the left eye image (reconstruction loss of the right disparity).

[0036] The bidirectional supervision of the reconstruction loss is quantified by the reconstruction loss between the original left eye image and the original right eye image. The original left eye image and the original right eye image are the label values ​​of the reconstruction loss of the two-branch convolutional neural network, and the label value is the true value. Therefore, the bidirectional supervision of the reconstruction loss belongs to the supervised training with accurate supervision truth value for the two-branch convolutional neural network.

[0037] The consistency loss is: ; In the formula, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for The parallax gradient.

[0038] Figure 3 middle, , refers to the left parallax By left parallax Reconstructed right parallax , , refers to the right parallax By right parallax Reconstructed left disparity .

[0039] One direction of supervision in the two-way supervision of disparity consistency loss is the loss between the left disparity predicted by the first branch neural network, the right disparity reconstructed by the left disparity predicted by the first branch neural network, and the right disparity predicted by the second branch neural network (left-right disparity consistency loss), and the other direction of supervision is the loss between the left disparity predicted by the second branch neural network, the right disparity reconstructed by the right disparity predicted by the second branch neural network, and the left disparity predicted by the first branch neural network (right-left disparity consistency loss).

[0040] The bidirectional supervision of the disparity consistency loss is quantified by the reconstruction loss between the left disparity predicted by the first branch neural network and the right disparity predicted by the second branch neural network. The left disparity and the right disparity are the label values ​​of the consistency loss of the two-branch convolutional neural network, but the label value is a predicted value. Therefore, the bidirectional supervision of the reconstruction loss belongs to the unsupervised training of the two-branch convolutional neural network with unsupervised truth value.

[0041] The present invention also superimposes the disparity gradient in the disparity consistency loss, thereby ensuring the left and right disparity consistency with the highest disparity value obtained after training is completed, and at the same time having the smallest disparity gradient, so that the obtained disparity map has the highest density and the disparity is kept smooth locally.

[0042] In the above, the combination of reconstruction loss and consistency loss combines supervised training with unsupervised training, and obtains a semi-supervised training two-branch convolutional neural network, constructing a semi-supervised binocular image depth estimation model, which can obtain accurate left disparity and right disparity estimation. The disparity can be directly converted into depth, so accurate stereo depth estimation can be achieved.

[0043] After obtaining the left and right bidirectional parallax, the present invention converts it into two corresponding real-world point cloud data, and obtains an accurate three-dimensional model of the intraoperative organ by point cloud registration and optimization to construct the surface point cloud of the intraoperative organ, as follows: The method for constructing the intraoperative scene depth map includes: Input the intraoperative scene binocular image into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the intraoperative scene binocular image; Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map; The first intraoperative scene depth map is: ; In the formula, h 1 is the first intraoperative scene depth map, b is the baseline distance, f is the focal length, is the left parallax; Obtain the baseline distance and focal length of the binocular camera, and obtain the second intraoperative scene depth map based on the right parallax in the left and right bidirectional parallax; The depth map of the second intraoperative scene is: ; In the formula, h 2 is the depth map of the scene during the second operation, b is the baseline distance, f is the focal length, is the right parallax.

[0044] The generation method of the field scenic spot cloud during the operation includes: Get the camera internal parameters of the binocular camera; The camera intrinsic parameters and the first intraoperative scene depth map are used to perform point cloud conversion to obtain the first intraoperative scene point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth value at ; The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative scene point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth of the place.

[0045] The reconstruction method of the intraoperative scene 3D model includes: The ICP point cloud registration method is used to align and register the first intraoperative scene cloud and the second intraoperative scene cloud to obtain the intraoperative scene registration point cloud; The Poisson reconstruction method is used to reconstruct the intraoperative scene registration point cloud into three dimensions to obtain a three-dimensional model of the intraoperative scene; The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

[0046] The network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same.

[0047] like Figure 2 As shown, the present invention provides a surgical robot high telepresence visual perception system, which is applied to a surgical robot high telepresence visual perception method, and the system includes: A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; A depth estimation unit is used to collect the real-time scene during the operation through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene; An intraoperative reconstruction unit, used to generate an intraoperative field point cloud from an intraoperative scene depth map, and to perform three-dimensional reconstruction based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; A preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model; The holographic projection unit is used to superimpose the registered three-dimensional model of the surgical scene onto the real-time scene during the operation through holographic projection, so as to obtain a highly immersive visual scene during the operation.

[0048] The semi-supervised stereo image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network.

[0049] The three-dimensional model of the surgical scene was reconstructed using 3D Slicer software.

[0050] The present invention provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, a high-presence visual perception method for a surgical robot is implemented.

[0051] The present invention collects intraoperative real-scene information and integrates it with preoperative image information, strengthens the visual perception of the surgical scene, develops a semi-supervised binocular image depth estimation model, realizes accurate three-dimensional reconstruction of the organ surface during surgery, and utilizes preoperative and intraoperative organ surface registration to realize holographic fusion and superposition of preoperative three-dimensional images and intraoperative endoscopic images, providing augmented reality visual perception, and providing doctors with advanced visual perception of organ occluded areas and complete information inside soft tissues.

[0052] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and protection scope of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present application.

Claims

1. A surgical robot high telepresence visual perception method, characterized in that: The following steps are involved: Collect preoperative 3D images, and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; The real-time scene during the operation is captured by a binocular camera to obtain a binocular image of the intraoperative scene, and a semi-supervised binocular image depth estimation model is used to perform binocular image depth estimation based on the binocular image of the intraoperative scene to obtain a depth map of the intraoperative scene; The intraoperative scene depth map is used to generate an intraoperative field point cloud, and a three-dimensional reconstruction is performed based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; The intraoperative scene 3D model is registered with the surgical scene 3D model, and the registered surgical scene 3D model is superimposed on the intraoperative real-time scene through holographic projection to obtain a highly immersive intraoperative visual scene.

2. A surgical robot high telepresence visual perception method according to claim 1, characterized in that: The method for constructing the semi-supervised binocular image depth estimation model includes: The left and right images in the binocular image of the intraoperative scene are estimated by using a two-branch convolutional neural network to obtain the left and right bidirectional disparity; Establishing a reconstruction loss and a consistency loss of left-right bidirectional disparity, and performing bidirectional self-supervised training on the dual-branch convolutional neural network based on the reconstruction loss and the consistency loss of left-right bidirectional disparity, to obtain a semi-supervised binocular image depth estimation model; The semi-supervised binocular image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network; The reconstruction loss is: ; In the formula, To rebuild the losses, is the left disparity of the i-th stereo image in the dataset used to train the semi-supervised stereo image depth estimation model, is the left eye image of the i-th binocular image in the data set, is the right eye image of the i-th binocular image in the data set, is the right disparity of the i-th binocular image in the dataset, for pass The reconstructed right image is for pass The reconstructed left eye image, n is the total number of binocular images in the data set; The consistency loss is: ; In the formula, is the consistency loss, for pass The reconstructed right disparity, for pass The reconstructed left disparity, n is the total number of binocular images in the dataset, for The parallax gradient, for The parallax gradient.

3. A surgical robot high telepresence visual perception method according to claim 2, characterized in that: The method for constructing the intraoperative scene depth map includes: Input the intraoperative scene binocular image into the semi-supervised binocular image depth estimation model to obtain the left and right bidirectional disparity of the intraoperative scene binocular image; Obtaining the baseline distance and focal length of the binocular camera, and converting the left parallax in the left and right bidirectional parallax to obtain a first intraoperative scene depth map; The first intraoperative scene depth map is: ; Where h1 is the first intraoperative scene depth map, b is the baseline distance, and f is the focal length. is the left parallax; Obtain the baseline distance and focal length of the binocular camera, and obtain the second intraoperative scene depth map based on the right parallax in the left and right bidirectional parallax; The second intraoperative scene depth map is: ; Where h2 is the depth map of the scene during the second operation, b is the baseline distance, and f is the focal length. is the right parallax.

4. A surgical robot high telepresence visual perception method according to claim 3, characterized in that: The method for generating the field scenic spot cloud during the operation includes: Get the camera internal parameters of the binocular camera; The camera intrinsic parameters and the first intraoperative scene depth map are used to perform point cloud conversion to obtain the first intraoperative scene point cloud (X1, Y1, Z1), where X1 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y1 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z1 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth value at ; The camera intrinsic parameters and the second intraoperative scene depth map are used to perform point cloud conversion to obtain the second intraoperative scene point cloud (X2, Y2, Z2), where X2 is the X-axis coordinate in the three-dimensional coordinates of the point cloud, Y2 is the Y-axis coordinate in the three-dimensional coordinates of the point cloud, and Z2 is the Z-axis coordinate in the three-dimensional coordinates of the point cloud; in, , , ; In the formula, Depth map The pixel coordinates in , The principal distance of the camera. , are the camera principal point coordinates, Depth map middle The depth of the place.

5. A surgical robot high telepresence visual perception method according to claim 4, characterized in that: The method for reconstructing the three-dimensional model of the intraoperative scene includes: The ICP point cloud registration method is used to align and register the first intraoperative scene cloud and the second intraoperative scene cloud to obtain the intraoperative scene registration point cloud; Reconstructing the intraoperative scene registration point cloud in three dimensions using a Poisson reconstruction method to obtain a three-dimensional model of the intraoperative scene; The image optimization tool in SLAM technology is used to optimize the three-dimensional model of the intraoperative scene.

6. A surgical robot high telepresence visual perception method according to claim 2, characterized in that: The network structures of the first branch convolutional neural network and the second branch convolutional neural network in the dual-branch convolutional neural network are the same.

7. A surgical robot high telepresence visual perception system, characterized in that: A method for high telepresence visual perception of a surgical robot as described in any one of claims 1 to 6, the system comprising: A preoperative reconstruction unit is used to collect preoperative 3D images and perform 3D reconstruction based on the preoperative 3D images to obtain a 3D model of the surgical scene; A depth estimation unit is used to collect the real-time scene during the operation through a binocular camera to obtain a binocular image of the intraoperative scene, and perform binocular image depth estimation based on the binocular image of the intraoperative scene using a semi-supervised binocular image depth estimation model to obtain a depth map of the intraoperative scene; An intraoperative reconstruction unit, used to generate an intraoperative field point cloud from an intraoperative scene depth map, and to perform three-dimensional reconstruction based on the intraoperative field point cloud to obtain a three-dimensional model of the intraoperative scene; A preoperative and intraoperative registration unit is used to register the intraoperative scene 3D model with the surgical scene 3D model to obtain a registered surgical scene 3D model; The holographic projection unit is used to superimpose the registered three-dimensional model of the surgical scene onto the real-time scene during the operation through holographic projection, so as to obtain a highly immersive visual scene during the operation.

8. The surgical robot high telepresence visual perception system according to claim 7, characterized in that: The semi-supervised binocular image depth estimation model is: ; ; In the formula, is the left parallax, is the right parallax, is the left eye image, is the right eye image, CNN1 is the first branch convolutional neural network, and CNN2 is the second branch convolutional neural network.

9. The surgical robot high telepresence visual perception system according to claim 7, characterized in that: The three-dimensional model of the surgical scene is obtained by three-dimensional reconstruction using 3D Slicer software.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Operation quality detection method and device

    CN113288452A

  • Deep learning semi-supervised dense matching method and system based on consistency constraint

    CN113780389A

  • Three-dimensional grid model registration fusion system for laparoscopic surgery navigation

    CN116485851A

  • Autonomous navigation method for tracheal intubation robot based on self-supervised monocular depth estimation

    CN116966381A

  • Semi-supervised 3D target detection method and system in indoor scene, and storage medium

    CN117671239A