3D laparoscopic surgery camera posture estimation method, device, equipment and storage medium
By acquiring a set of images and using the structural similarity index and optical flow map combined with the UNet network for pose estimation, the problem of inaccurate pose estimation of laparoscopic cameras was solved, high-precision pose estimation was achieved in complex surgical environments, and the safety and accuracy of the surgery were improved.
Patent Information
- Application Number
- CN202410777694.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-06-17
AI Technical Summary
The existing technology lacks accurate pose estimation for laparoscopic cameras, leading to increased difficulty and increased risk in surgeries, especially in complex medical scenarios with limited field of view, complex operation of surgical instruments, and lack of three-dimensional information about the surgical scene.
By acquiring a set of abdominal cavity images, the structural similarity index formula is used to calculate the matching cost and depth image, and the optical flow map and UNet network are combined for posture estimation, reducing the dependence on high-performance hardware and improving the accuracy of feature extraction.
While maintaining high precision, the complexity of the equipment is reduced, and surgical data obtained by the binocular system is used to achieve more accurate camera pose estimation, enhance the accuracy and safety of surgical navigation, optimize the surgical process and improve the quality of postoperative evaluation.
Smart Images

Figure CN118674770B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a 3D laparoscopic surgery camera posture estimation method, device, equipment and storage medium. Background Art
[0002] With advances in medical technology, minimally invasive surgery has become a key development direction in modern surgery. Laparoscopic surgery, with its advantages of minimal trauma and rapid recovery, has been widely used in various abdominal surgeries. Compared with open surgery, minimally invasive surgery reduces trauma and associated pain, and significantly shortens postoperative recovery time. 3D laparoscopic surgery captures two images of the surgical area at slightly different angles, then rasterizes and displays them in real time on a specialized 3D monitor. Surgeons wearing 3D glasses to view these stereoscopic images enable precise manipulation of surgical instruments, greatly improving surgical precision and safety. Laparoscopic surgery has been widely used in various abdominal surgeries, not only optimizing the surgical process but also providing patients with higher-quality medical services.
[0003] In highly complex medical scenarios such as laparoscopic surgery, surgical teams often face challenges such as limited field of view, complex instrument manipulation, and a lack of 3D information about the surgical scene. These factors collectively increase the difficulty and risk of the procedure. Furthermore, dynamic changes during surgery, such as tissue movement, bleeding, or smoke, can further impair visual clarity, making the processing and analysis of surgical data particularly challenging. Therefore, accurate posture information enables physicians to correlate local intraoperative image information with global 3D position, understanding instrument location and facilitating diagnosis and treatment. This plays a crucial role in improving surgical safety and efficiency. By extracting key features from laparoscopic video streams and combining them with advanced image processing techniques and deep learning algorithms, the relative position and posture of the laparoscope relative to the previous moment can be estimated. This technology not only improves surgical precision but also provides physicians with rich depth information and spatial positioning capabilities, helping them make more accurate surgical decisions.
[0004] In highly complex medical scenarios such as laparoscopic surgery, the processing and analysis of surgical data presents a host of challenges. Laparoscopic camera pose estimation is the process of accurately determining the specific position and orientation of the laparoscopic camera in three-dimensional space. Accurate pose estimation is crucial for guiding surgeons to perform precise and safe surgical procedures, but current laparoscopic camera pose estimation methods are not sufficiently accurate. Summary of the Invention
[0005] Based on this, it is necessary to address the technical problem that the current laparoscopic camera posture estimation in the existing technology is not accurate enough, and propose a 3D laparoscopic surgery camera posture estimation method, device, equipment and storage medium.
[0006] In a first aspect, a method for estimating a camera pose in 3D laparoscopic surgery is provided, the method comprising:
[0007] Acquiring an image set of the abdominal cavity, the image set including a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0008] Calculating based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost;
[0009] determining a depth image based on the first left image, the first right image, and the matching cost;
[0010] Determining a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and determining a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0011] Calculating optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0012] Determine a third optical flow map based on the first structural similarity index and the first optical flow map, and determine a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0013] Perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network.
[0014] In a second aspect, a 3D laparoscopic surgery camera posture estimation device is provided, the device comprising:
[0015] an acquisition module, configured to acquire an image set of the abdominal cavity, the image set comprising a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0016] A first calculation module, configured to calculate based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost;
[0017] A first determining module, configured to determine a depth image based on the first left image, the first right image, and the matching cost;
[0018] a second determining module, configured to determine a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and to determine a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0019] a second calculation module, configured to calculate optical flow based on the first left image and the second left image to obtain a first optical flow map, and to calculate optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0020] a third determining module, configured to determine a third optical flow map based on the first structural similarity index and the first optical flow map, and to determine a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0021] A posture estimation module is used to perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map and the trained UNet network.
[0022] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned 3D laparoscopic surgery camera posture estimation method when executing the computer program.
[0023] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned 3D laparoscopic surgery camera posture estimation method are implemented.
[0024] The present invention proposes a 3D laparoscopic surgical camera posture estimation method, which obtains an image set of the abdominal cavity, wherein the image set includes a first left-side image taken from the left side, a first right-side image taken from the right side, a second left-side image taken from the left side, and a second right-side image taken from the right side, the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time, and then a matching cost is obtained based on the first left-side image, the first right-side image and a preset structural similarity index formula, and then a depth image is determined based on the first left-side image, the first right-side image and the matching cost. The first left image, the second left image and the structural similarity index formula are used to determine the first structural similarity index, and based on the first right image, the second right image and the structural similarity index formula, the second structural similarity index is determined, and based on the first left image and the second left image, the optical flow is calculated to obtain a first optical flow map, and based on the first right image and the second right image, the optical flow is calculated to obtain a second optical flow map, thereby determining a third optical flow map based on the first structural similarity index and the first optical flow map, and determining a fourth optical flow map based on the second structural similarity index and the second optical flow map, and finally performing posture estimation based on the depth image, the third optical flow map, the fourth optical flow map and the trained UNet network. The present invention can, under the premise of maintaining high accuracy, not require the support of high-performance CPUs, GPUs and high-quality sensors (such as RGB-D cameras, lidars, etc.), reduce the complexity of the equipment, better utilize the surgical data obtained by the binocular system, improve the accuracy of feature extraction, and achieve more accurate camera posture estimation in different surgical scenarios to enhance the accuracy and comprehensiveness of surgical navigation, improve the accuracy and safety of surgery, optimize the surgical process and improve the quality of postoperative evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0026] in:
[0027] Figure 1 A diagram showing an application environment of a method for estimating a camera posture in 3D laparoscopic surgery according to an embodiment;
[0028] Figure 2 is a flowchart of a method for estimating a camera pose in 3D laparoscopic surgery according to one embodiment;
[0029] Figure 3 Schematic diagram of a method for estimating a camera posture in 3D laparoscopic surgery according to one embodiment;
[0030] Figure 4 Schematic diagram of a second convolutional neural network of a method for estimating camera pose in 3D laparoscopic surgery according to one embodiment;
[0031] Figure 5 A visualization result of a test trajectory diagram of a 3D laparoscopic surgery camera pose estimation method according to an embodiment;
[0032] Figure 6 This is another test trajectory visualization result of a 3D laparoscopic surgery camera pose estimation method according to an embodiment;
[0033] Figure 7 is a structural block diagram of a 3D laparoscopic surgery camera posture estimation device according to one embodiment;
[0034] Figure 8 is a structural block diagram of a computer device in one embodiment;
[0035] Figure 9 It is a structural block diagram of a computer device in another embodiment. DETAILED DESCRIPTION
[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0037] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort shall fall within the scope of protection of the present invention.
[0039] The 3D laparoscopic surgery camera posture estimation method provided by the embodiment of the present invention can be applied in the following fields: Figure 1In an application environment, the client 110 communicates with the server 120 through a network. The server 120 can receive a set of images of the abdominal cavity through the client 110, wherein the image set includes a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side. The shooting time of the first left-side image is earlier than the shooting time of the second left-side image, the shooting time of the second right-side image is earlier than the shooting time of the second right-side image, and the shooting time of the first left-side image is the same as that of the first right-side image. Then, the server 120 calculates based on the first left-side image, the first right-side image and a preset structural similarity index formula to obtain a matching cost. Then, the server 120 determines a depth image based on the first left-side image, the first right-side image and the matching cost. Then, the server 120 determines a depth image based on the The first left image, the second left image and the structural similarity index formula are used to determine the first structural similarity index, and the second structural similarity index is determined based on the first right image, the second right image and the structural similarity index formula, and the optical flow is calculated based on the first left image and the second left image to obtain a first optical flow map, and the optical flow is calculated based on the first right image and the second right image to obtain a second optical flow map, so that the server 120 determines the third optical flow map based on the first structural similarity index and the first optical flow map, and determines the fourth optical flow map based on the second structural similarity index and the second optical flow map. Finally, the server 120 performs posture estimation based on the depth image, the third optical flow map, the fourth optical flow map and the trained UNet network. The present invention can maintain high accuracy without the support of high-performance CPU, GPU and high-quality sensors (such as RGB-D camera, lidar, etc.), reduce the complexity of the equipment, make better use of the surgical data obtained by the binocular system, improve the accuracy of feature extraction, and achieve more accurate camera posture estimation in different surgical scenarios, so as to enhance the accuracy and comprehensiveness of surgical navigation, improve the accuracy and safety of surgery, optimize the surgical process and improve the quality of postoperative evaluation. Among them, the client 110 can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server 120 can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0040] See also Figure 2 as well as Figure 3 As shown, Figure 2 A flowchart of a method for estimating the pose of a 3D laparoscopic surgical camera provided by one embodiment of the present invention and a schematic diagram of the structure of the method are provided. The method for estimating the pose of a 3D laparoscopic surgical camera includes the following steps:
[0041] Step S101: Acquire an image set of the abdominal cavity, the image set including a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0042] Step S102: performing calculation based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost;
[0043] The Structural Similarity Index (SSIM) is a metric used to measure the similarity between two images. SSIM simulates how the human visual system perceives image quality, taking into account brightness, contrast, and structural information.
[0044]
[0045] Among them, x and y are the local windows of the two images, μ x and μ y are the mean of image x and y respectively, and are the variances of images x and y, σ xy It is the covariance of images x and y. The SSIM value is between -1 and 1, where 1 means that the two images are exactly the same.
[0046] It's important to note that SSIM is used to calculate the matching cost between feature maps, assisting the network in accurately matching left and right features. The cost volume stores the matching cost of pairing a pixel in feature map f1 with the corresponding pixel in feature map f2. By introducing SSIM, the calculation takes into account brightness, contrast, and structure, better reflecting the visual similarity of images and providing the network with a more reliable matching cost, thereby improving the accuracy of depth maps.
[0047] The matching cost formula is as follows:
[0048] cost(x0,y0)=SSIM(f1(x0,y0),f2(x0+d,y0))
[0049] Among them, x0∈[0,w], y0∈[0,h], d∈[0,d max], f(x,y) represents the feature block of the sliding window at position (x,y), and cost(x0,y0) represents the SSIM value between the feature block at position (x0,y0) and the feature block at position (x0+d,y0) as the cost. As an example, feature map f1 is extracted from the first left image, feature map f2 is extracted from the first right image, and a sliding window is applied to the two feature maps. For each sliding window position, the SSIM value between the feature blocks is calculated as the matching cost. When calculating the disparity, the value in the cost volume is minimized, the optimal disparity for each pixel position is found, and the depth is calculated based on the disparity.
[0050] Step S103: determining a depth image based on the first left image, the first right image, and the matching cost;
[0051] In one embodiment, the step of determining the depth image based on the first left image, the first right image, and the matching cost includes:
[0052] Step S1031: determining a first depth image based on the first left image, the first right image, and the trained depth network;
[0053] The depth network is used to preliminarily extract depth images of the first left image and the first right image as the first depth image.
[0054] Step S1032: Based on the first depth image and the trained first convolutional neural network, obtain the matching cost, and determine the depth image based on the matching cost.
[0055] The first convolutional neural network is used to calculate the matching cost.
[0056] Step S104: determining a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and determining a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0057] Specifically, the optical flow is used to deform the image of the next time frame into the perspective of the previous time frame, and the SSIM between the deformed image and the previous time frame image is calculated. The formula of this principle is as follows:
[0058] W d =I s (x+u x , y+v y )
[0059] ssim s =SSIM(W d ,I d )
[0060] Among them, I s is the source image, (u x , v y ) is the displacement of the optical flow in the x and y directions, W d is the target image after optical flow deformation, SSIM(W d ,I k ) represents the calculation of W d and I d The structural similarity index SSIM between them.
[0061] As an example, the optical flow is used to convert the second left image into the perspective of the first left image to obtain a third left image, and the structural similarity index between the third left image and the first left image is calculated as the first structural similarity index.
[0062] Step S105: calculating optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0063] In one embodiment, the steps of calculating the optical flow based on the first left image and the second left image to obtain the first optical flow map, and calculating the optical flow based on the first right image and the second right image to obtain the second optical flow map include:
[0064] Step S1051: determining a first optical flow map based on the first left image, the second left image, and the trained RAFT optical flow network;
[0065] Step S1052: Determine a second optical flow map based on the first right image, the second right image, and the trained RAFT optical flow network.
[0066] Among them, the RAFT optical flow network is used to calculate the optical flow.
[0067] Step S106: determining a third optical flow map based on the first structural similarity index and the first optical flow map, and determining a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0068] In one embodiment, the step of determining the third optical flow map based on the first structural similarity index and the first optical flow map, and determining the fourth optical flow map based on the second structural similarity index and the second optical flow map includes:
[0069] Step S1061: determining a third optical flow map based on the first structural similarity index, the first optical flow map, and the trained second convolutional neural network;
[0070] Step S1062: Determine a fourth optical flow map based on the second structural similarity index, the second optical flow map, and the trained second convolutional neural network.
[0071] In one embodiment, Figure 4 As shown, the second convolutional neural network includes a plurality of sequentially connected convolution modules, and the convolution module includes an average pooling layer, a first convolutional layer, a maximum pooling layer, a stacking layer, and a second convolutional layer. The output end of the average pooling layer, the output end of the first convolutional layer, and the output end of the maximum pooling layer are respectively connected to the input end of the stacking layer, and the output end of the stacking layer is connected to the input end of the second convolutional layer. The input ends of the average pooling layer, the first convolutional layer, and the maximum pooling layer serve as the input ends of the second convolutional neural network, and the output end of the second convolutional layer serves as the output end of the second convolutional neural network.
[0072] Among them, the second convolutional neural network is used to adjust and optimize the optical flow to obtain more accurate optical flow estimation.
[0073] Step S107: performing posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network.
[0074] In one embodiment, the step of performing pose estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network includes:
[0075] Step S501: performing 2D residual calculation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained two-layer UNet network;
[0076] Step S502: performing 3D residual calculation based on the three-layer UNet network trained on the depth image, the third optical flow map, and the fourth optical flow map;
[0077] Step S503: generating a comprehensive residual function based on the 2D residual and the 3D residual, and optimizing the pose estimation based on minimizing the comprehensive residual function.
[0078] Among them, the 2D residual r 2D (p t ,x) and 3D residual r 3D (p t ,x), the 2D residual function calculates the pixel position difference between the current frame and the previous frame on the 2D projection based on a single depth map; the 3D residual function calculates the alignment error of the point cloud between the two frames in 3D space. 2D (p t ,x) and 3D residual r 3D (p t ,x) is as follows
[0079]
[0080] r 3D (p t ,x)=||exp(p t )π 3D (D t ,x)-π 3D (D t-1 ,x+F t (x))||2
[0081] The 2D and 3D residuals are combined together and each residual is weighted using adaptive weights to generate a comprehensive residual function, which is shown below:
[0082] r(p t ,x)=w 2D (x)r 2D (p t ,x)+w 3D (x)r 3D (p t ,x)
[0083] Among them, π 2D is the projection function from 3D to 2D, π 3D is the reprojection function from 2D to 3D, exp(p t ) is the matrix index from Lie algebra to Lie group, D t is the depth map at time t, F t (x) is the optical flow function, p t is the relative posture at time t, w 2D (x) and w 3D (x) is the pixel-level weight map of 2D and 3D residuals, X and Y represent the image size, F t (x) is an optical flow function used to represent the third optical flow map and the fourth optical flow map, θ 2D is the network parameter of the 2D residual, θ 3D is the network parameter of 3D residual, F t is the optical flow at time t, F′ t It is based on F t The calculated disparity flow, is the left image at time t, D t is the depth map at time t, is the left image at time t-1, D t-1 is the depth map at time t-1.
[0084] The weight mapping is learned by two independent UNet networks, which accept inputs such as images, depth maps and optical flows, and output weight mappings of 2D and 3D residuals. The two-layer UNet network is used to learn w 2D (x), the three-layer UNet network learns w 3D (x).
[0085] Finally, the relative pose is optimized by minimizing the comprehensive residual function. By training the network, minimizing the training loss using the ground truth pose, learning the adaptive weight mapping, and thus optimizing the pose estimation, the expression for minimizing the comprehensive residual function is as follows:
[0086]
[0087] Here, Ω contains the set of all spatial image coordinates x, and the optimization is performed in the Lie algebra vector space.
[0088] It should be noted that the 3D laparoscopic surgical camera pose estimation method of the present invention has been experimented on the public datasets SCAREDMICCAI 2019 dataset and StereoMIS dataset, and has achieved the best results so far. The SCARED dataset contains a 1280×2048 high-resolution video stream shot using the Da Vinci surgical robot, as well as the corresponding three-dimensional position and pose information of the camera. It consists of 7 training datasets and 2 test datasets. Each dataset corresponds to a live pig research subject and contains 4 to 5 key frames. And the real camera pose provides a benchmark for evaluation of the model. The StereoMIS dataset was also shot using the Da Vinci surgical robot, with a total of 16 recording sequences. The sequence duration ranges from 50 seconds to 30 minutes, and includes challenging scenes such as respiratory movement, tissue deformation, resection, bleeding, and smoke. Posture estimation is generally measured by ATE-RMSE (Absolute Trajectory Error-Root Mean Square Error), which can be expressed as:
[0089]
[0090] in, is the estimated location and the true position p i The Euclidean distance between them is , and N is the total number of positions in the trajectory. By reproducing and modifying previous work, the test results of the model on the StereoMIS dataset P2 and P3 are shown in the following table:
[0091]
[0092] Figure 5 、 Figure 6This is the test trajectory visualization result, where Figure 5 This is the test trajectory diagram of the StereoMIS dataset P2_7. Figure 6 This is the test trajectory diagram of the StereoMIS dataset P2_8.
[0093] See also Figure 7 As shown, in one embodiment, a 3D laparoscopic surgery camera posture estimation device is provided, the device comprising:
[0094] an acquisition module 10, configured to acquire an image set of the abdominal cavity, the image set comprising a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0095] A first calculation module 20 is configured to calculate a matching cost based on the first left image, the first right image, and a preset structural similarity index formula;
[0096] A first determining module 30 is configured to determine a depth image based on the first left image, the first right image, and the matching cost;
[0097] a second determining module 40, configured to determine a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and to determine a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0098] A second calculation module 50 is configured to calculate optical flow based on the first left image and the second left image to obtain a first optical flow map, and to calculate optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0099] a third determining module 60, configured to determine a third optical flow map based on the first structural similarity index and the first optical flow map, and to determine a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0100] The posture estimation module 70 is used to perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map and the trained UNet network.
[0101] A first determining module 30 is configured to determine a first depth image based on the first left image, the first right image, and the trained depth network;
[0102] The matching cost is obtained based on the first depth image and the trained first convolutional neural network, and the depth image is determined based on the matching cost.
[0103] A second calculation module 50 is configured to determine a first optical flow map based on the first left image, the second left image, and the trained RAFT optical flow network;
[0104] A second optical flow map is determined based on the first right image, the second right image, and the trained RAFT optical flow network.
[0105] In one embodiment, the third determining module 60 is configured to determine a third optical flow map based on the first structural similarity index, the first optical flow map, and the trained second convolutional neural network;
[0106] A fourth optical flow map is determined based on the second structural similarity index, the second optical flow map, and the trained second convolutional neural network.
[0107] In one embodiment, the second convolutional neural network includes multiple convolutional modules connected in sequence, and the convolutional module includes an average pooling layer, a first convolutional layer, a maximum pooling layer, a stacking layer, and a second convolutional layer. The output end of the average pooling layer, the output end of the first convolutional layer, and the output end of the maximum pooling layer are respectively connected to the input end of the stacking layer, and the output end of the stacking layer is connected to the input end of the second convolutional layer. The input ends of the average pooling layer, the first convolutional layer, and the maximum pooling layer serve as the input ends of the second convolutional neural network, and the output end of the second convolutional layer serves as the output end of the second convolutional neural network.
[0108] In one embodiment, the posture estimation module 70 is further configured to: perform 2D residual calculation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained two-layer UNet network;
[0109] Performing 3D residual calculation based on the three-layer UNet network trained on the depth image, the third optical flow map, and the fourth optical flow map;
[0110] A comprehensive residual function is generated based on the 2D residual and the 3D residual, and the pose estimation is optimized based on minimizing the comprehensive residual function.
[0111] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a 3D laparoscopic surgical camera posture estimation method service side.
[0112] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a 3D laparoscopic surgery camera posture estimation method.
[0113] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0114] Acquiring an image set of the abdominal cavity, the image set including a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0115] Calculating based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost;
[0116] determining a depth image based on the first left image, the first right image, and the matching cost;
[0117] Determining a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and determining a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0118] Calculating optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0119] Determine a third optical flow map based on the first structural similarity index and the first optical flow map, and determine a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0120] Perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network.
[0121] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0122] Acquiring an image set of the abdominal cavity, the image set including a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time;
[0123] Calculating based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost;
[0124] determining a depth image based on the first left image, the first right image, and the matching cost;
[0125] Determining a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and determining a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula;
[0126] Calculating optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating optical flow based on the first right image and the second right image to obtain a second optical flow map;
[0127] Determine a third optical flow map based on the first structural similarity index and the first optical flow map, and determine a fourth optical flow map based on the second structural similarity index and the second optical flow map;
[0128] Perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network.
[0129] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0130] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0131] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0132] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A 3D laparoscopic surgery camera pose estimation method, characterized in that: The 3D laparoscopic surgery camera posture estimation method includes: Acquiring an image set of the abdominal cavity, the image set including a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time; Calculating based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost; determining a depth image based on the first left image, the first right image, and the matching cost; Determining a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and determining a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula; Calculating optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating optical flow based on the first right image and the second right image to obtain a second optical flow map; Determine a third optical flow map based on the first structural similarity index and the first optical flow map, and determine a fourth optical flow map based on the second structural similarity index and the second optical flow map; Perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network.
2. The 3D laparoscopic surgery camera posture estimation method according to claim 1, characterized in that: The step of determining a depth image based on the first left image, the first right image, and the matching cost includes: Determine a first depth image based on the first left image, the first right image, and the trained depth network; The matching cost is obtained based on the first depth image and the trained first convolutional neural network, and the depth image is determined based on the matching cost.
3. The 3D laparoscopic surgery camera posture estimation method according to claim 1, characterized in that: The steps of calculating the optical flow based on the first left image and the second left image to obtain a first optical flow map, and calculating the optical flow based on the first right image and the second right image to obtain a second optical flow map include: Determine a first optical flow map based on the first left image, the second left image, and the trained RAFT optical flow network; A second optical flow map is determined based on the first right image, the second right image, and the trained RAFT optical flow network.
4. The 3D laparoscopic surgery camera posture estimation method according to claim 1, characterized in that: The step of determining a third optical flow map based on the first structural similarity index and the first optical flow map, and determining a fourth optical flow map based on the second structural similarity index and the second optical flow map includes: Determining a third optical flow map based on the first structural similarity index, the first optical flow map, and the trained second convolutional neural network; A fourth optical flow map is determined based on the second structural similarity index, the second optical flow map, and the trained second convolutional neural network.
5. The 3D laparoscopic surgery camera posture estimation method according to claim 4, characterized in that: The second convolutional neural network includes multiple convolution modules connected in sequence, and the convolution module includes an average pooling layer, a first convolutional layer, a maximum pooling layer, a stacking layer, and a second convolutional layer. The output end of the average pooling layer, the output end of the first convolutional layer, and the output end of the maximum pooling layer are respectively connected to the input end of the stacking layer, and the output end of the stacking layer is connected to the input end of the second convolutional layer. The input ends of the average pooling layer, the first convolutional layer, and the maximum pooling layer serve as the input ends of the second convolutional neural network, and the output end of the second convolutional layer serves as the output end of the second convolutional neural network.
6. The 3D laparoscopic surgery camera posture estimation method according to claim 1, characterized in that: The step of performing posture estimation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained UNet network includes: Performing 2D residual calculation based on the depth image, the third optical flow map, the fourth optical flow map, and the trained two-layer UNet network; Performing 3D residual calculation based on the three-layer UNet network trained on the depth image, the third optical flow map, and the fourth optical flow map; A comprehensive residual function is generated based on the 2D residual and the 3D residual, and the pose estimation is optimized based on minimizing the comprehensive residual function.
7. A 3D laparoscopic surgery camera posture estimation device, characterized in that: The 3D laparoscopic surgery camera posture estimation device comprises: an acquisition module, configured to acquire an image set of the abdominal cavity, the image set comprising a first left-side image taken on the left side, a first right-side image taken on the right side, a second left-side image taken on the left side, and a second right-side image taken on the right side, wherein the first left-side image is taken earlier than the second left-side image, the second right-side image is taken earlier than the second right-side image, and the first left-side image and the first right-side image are taken at the same time; A first calculation module, configured to calculate based on the first left image, the first right image, and a preset structural similarity index formula to obtain a matching cost; A first determining module, configured to determine a depth image based on the first left image, the first right image, and the matching cost; a second determining module, configured to determine a first structural similarity index based on the first left image, the second left image, and the structural similarity index formula, and to determine a second structural similarity index based on the first right image, the second right image, and the structural similarity index formula; a second calculation module, configured to calculate optical flow based on the first left image and the second left image to obtain a first optical flow map, and to calculate optical flow based on the first right image and the second right image to obtain a second optical flow map; a third determining module, configured to determine a third optical flow map based on the first structural similarity index and the first optical flow map, and to determine a fourth optical flow map based on the second structural similarity index and the second optical flow map; A posture estimation module is used to perform posture estimation based on the depth image, the third optical flow map, the fourth optical flow map and the trained UNet network.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the 3D laparoscopic surgery camera posture estimation method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the 3D laparoscopic surgery camera pose estimation method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image-feature-based method for reconstruction of abdominal cavity environment map and laparoscope positioning
CN108090954A
Scene flow estimation method and device and scene flow estimation model training method and device
CN113160278A