Surgical robot intraoperative dynamic scene three-dimensional reconstruction method
Through the improved 3D Gaussian rendering model, combining depth maps and tissue masks, the three-dimensional point clouds are extracted and fused, and the motion curve is learned through the flexible deformation model, the problems of low 3D reconstruction efficiency and unclear geometric representation in dynamic surgical scenarios are solved, achieving efficient and accurate 3D reconstruction effect.
Patent Information
- Application Number
- CN202510541408.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing three-dimensional reconstruction techniques have problems of inefficiency and unclear geometric representation when dealing with dynamic surgical scenarios, especially in complex surgical scenarios, which are difficult to achieve efficient and precise reconstruction.
An improved three-dimensional Gaussian rendering model is proposed, including an initialization module, a flexible deformation model and a rendering module. The three-dimensional point cloud is extracted through the depth map and the organization mask, and the initial Gaussian point cloud is fused to generate the initial Gaussian point cloud, and the motion curve of the Gaussian point is learned through the flexible deformation model, and finally a two-dimensional image and depth map are generated through the differentiated rendering algorithm.
It significantly improves the rendering speed, meets the requirements of real-time navigation in the operation, accurately captures dynamic changes in the surgical scene, improves the geometric accuracy of three-dimensional reconstruction, enhances the model's ability to handle depth map noise and errors, and realizes efficient and precise reconstruction of dynamic surgical scenes in surgical robots.
Smart Images

Figure CN120070775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surgical robots, and more specifically, to a method for three-dimensional reconstruction of intraoperative dynamic scenes of a surgical robot. Background Art
[0002] In the field of surgical robots, accurate intraoperative three-dimensional scene reconstruction is crucial for surgical navigation and operation. Existing three-dimensional reconstruction technologies have some problems when dealing with dynamic surgical scenes. For example, traditional stereo matching and SLAM (Simultaneous Localization and Mapping) methods assume that the scene is static and cannot effectively handle the dynamic changes of tissues and tools during the surgical process. While the NeRF (Neural Radiance Field)-based method performs well in static scene reconstruction, it has a long training time, slow rendering speed, and lacks a clear geometric representation in dynamic scenes. In recent years, the 3D Gaussian Splatting (3DGS) technology has received extensive attention due to its high efficiency and high-quality rendering effect. 3DGS represents the scene as a three-dimensional Gaussian point cloud and adopts an efficient differential rendering algorithm to quickly generate high-quality images and depth maps. However, although the recent 3D Gaussian rendering methods have improved the rendering speed, they still have deficiencies in initialization and depth change capture, especially when dealing with complex dynamic surgical scenes. Summary of the Invention
[0003] The object of the present invention is to propose a method for three-dimensional reconstruction of intraoperative dynamic scenes of a surgical robot, and through an improved three-dimensional Gaussian rendering model, to achieve efficient and accurate reconstruction of the intraoperative dynamic three-dimensional scene of the surgical robot.
[0004] To achieve the above object, the present invention proposes a method for three-dimensional reconstruction of intraoperative dynamic scenes of a surgical robot, including: S1: Collect an image sequence of the surgical scene, where the image sequence includes various surgical scenes and image pairs of left and right views with dynamic changes; preprocess the image sequence, generate a depth map for each frame based on the preprocessed image pairs, and extract a tissue mask for each frame; S2: Construct a three-dimensional reconstruction model based on 3DGS, where the three-dimensional reconstruction model includes: Initialization module: used to extract the three-dimensional point cloud for each frame according to the depth map and the tissue mask, and generate a dense initial Gaussian point cloud by fusing multi-frame point clouds through a binary motion mask; Flexible deformation model: used to assign dynamic attributes to each Gaussian point in the initial Gaussian point cloud, and learn the motion curve of each Gaussian point through a Gaussian-type combined basis function; Rendering module: used to convert Gaussian point cloud into 2D image and depth map through differential rendering algorithm; S3: Use the preprocessed image, depth map and tissue mask as training data to train the 3D reconstruction model, and optimize the explicit parameters of Gaussian points and the weights and parameters of the basis functions of the flexible deformation model by minimizing the total loss function; S4: Load the trained 3D reconstruction model in the surgical navigation system, generate a 3D view of the surgical scene in real time, and output for surgical navigation and robot control.
[0005] Optionally, in step S1, the depth map is generated by a stereo matching algorithm, and the tissue mask is extracted by an automatic segmentation algorithm or a pre-trained segmentation model.
[0006] Optionally, in step S2, the calculation formula for extracting the 3D point cloud in each frame based on the depth map and tissue mask is:
[0007] where, P i is the 3D point cloud in the i-th frame, I i is the i-th frame image, D i is the depth map of the current i-th frame, M i is the tissue mask of the i-th frame, K 1 and K 2 are the camera intrinsic matrix and extrinsic matrix respectively, and ⊙ represents element-wise multiplication.
[0008] Optionally, in step S2, the method for generating the binary motion mask includes: Calculate the depth map difference between each frame and the reference frame to generate a binary motion mask, and the reference frame is the first frame; The generation formula of the binary motion mask is: B i = I (| D 0 - D i | > τ ) ∪ (1 - M 0 ) ∩ M i where, B i is the binary motion mask of the i-th frame, I (⋅) is the indicator function,M 0 is the organization mask of the reference frame, M i is the organization mask of the current frame, D 0 is the depth map of the reference frame, τ is the threshold of depth change, used to determine whether the depth change of pixels is significant.
[0009] Optionally, in step S2, a dense initial Gaussian point cloud is generated by fusing multiple frames of point clouds through a binary motion mask. The fusion formula is: P ={ P 0 , P 1 ⊙ B 1 ,…, P T ⊙ B T} where, P represents the fused initial Gaussian point cloud, P 0 represents the 3D point cloud of the first frame, P i represents the i frame's 3D point cloud, B i is the i frame's binary motion mask, T represents the number of frames, i ∈ [1, T].
[0010] Optionally, in step S2, the flexible deformation model represents the complex motion of tissues through a Gaussian-type basis function combination. The position, rotation, and scaling changes of each Gaussian point are controlled by learnable weights, centers, and variances, where: (1) where, t represents time, θ and σ are the learnable center positions and variances; The change in the position of the Gaussian point along the X-axis is represented by a basis function combination as: (2) where, ψ μ,x ( t ; Θ μ,x ) represents the position change of the Gaussian point along the x-axis at time t ; Θ μ,xA set of parameters representing the position change along the x-axis, , represents the contribution of the j -th basis function to the motion curve of the position change along the x-axis, represents the central position of the j -th basis function in time, and represents the width of the j -th basis function in time; represents the j -th Gaussian basis function, which is calculated using formula (1); j is an index variable used to traverse the set of basis functions. The number of basis functions is determined by B and is expressed as j = 1, 2, …, B . Each basis function j has its own parameters, including the weight ω , the center θ and the variance σ ; The position of the Gaussian point along the x-axis at time t is expressed as: μ x ( t ) = μ x (0) + ψ μ,x ( t ) where μ x ( t ) represents the position of the Gaussian point along the x-axis at time t , μ x (0) represents the position of the Gaussian point along the x-axis at the initial time t = 0, ψ μ,x ( t ) represents the change in the position of the Gaussian point along the x-axis at time t and is calculated using formula (2).
[0011] Optionally, in step S3, the process of training the three-dimensional reconstruction model includes: The model updates the dynamic attributes of the Gaussian point cloud according to the input depth map and tissue mask; The flexible deformation model adjusts the attributes of the Gaussian points according to the time series; The rendering module renders the updated Gaussian point cloud into a two-dimensional image and a depth map; Calculate the total loss through the total loss function. The total loss includes photometric loss, normalized depth loss, and total variation loss. The photometric loss is used to optimize the color difference between the rendered image and the actual image, and the normalized depth loss and the total variation loss are used to optimize the difference between the predicted depth map and the actual depth map; Update the model parameters through the backpropagation algorithm according to the calculated total loss to minimize the total loss function.
[0012] Optionally, the calculation formula for the normalized depth regularization loss is:
[0013] Where, is the normalized depth regularization loss, M is the mask, and D norm are the normalized predicted depth map and the actual depth map respectively. ⊙ represents the element-wise multiplication operation, and the mask M is applied to the difference calculation of the depth map.
[0014] Optionally, the calculation formula for the total variation loss is:
[0015] Where, L smooth represents the depth smoothness loss, which is used to remove noise in the depth map and retain depth details, represents the number of valid pixels in the predicted depth map, is the result of Canny edge detection, indicating whether the pixel ( i , j ) is on the edge. If the pixel is on the edge, the value is 0, otherwise it is 1; represents the absolute value of the depth difference between the pixel ( i , j ) and its right pixel ( i +1, j ), represents the absolute value of the depth difference between the pixel ( i , j ) and its lower pixel ( i , j +1).
[0016] Optionally, the total loss function is:
[0017] Where, L colour is the original photometric loss of 3DGS, is the normalized depth loss, L smooth is the total variation loss, λ smooth is the weight coefficient of the smooth loss.
[0018] The beneficial effects of the present invention are as follows: Through the improved 3D Gaussian rendering framework and the adoption of 3D Gaussian rendering technology, the present invention significantly improves the rendering speed and meets the requirements of intraoperative real-time navigation. Through the flexible deformation model and depth optimization steps, it accurately captures the dynamic changes in the surgical scene and improves the geometric accuracy of 3D reconstruction. By introducing normalized depth regularization and unsupervised depth smoothing, it enhances the model's ability to handle depth map noise and errors, improves the stability and reliability of the reconstruction results, and thus realizes the efficient and accurate reconstruction of the intraoperative dynamic surgical scene of the surgical robot.
[0019] The system of the present invention has other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent detailed description, or will be described in detail in the accompanying drawings incorporated herein and the subsequent detailed description. These accompanying drawings and detailed description are used together to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more obvious. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0021] Figure 1 Shows a flowchart of the steps of a method for three-dimensional reconstruction of an intraoperative dynamic scene of a surgical robot according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0023] As Figure 1 shown, a method for three-dimensional reconstruction of an intraoperative dynamic scene of a surgical robot according to the present invention includes: S1: Collect an image sequence of the surgical scene, where the image sequence includes image pairs of various surgical scenes and left and right views with dynamic changes; preprocess the image sequence, generate a depth map for each frame based on the preprocessed image pairs, and extract the tissue mask in each frame; Specifically, a surgical video image sequence is collected using a stereoscopic endoscope (such as a laparoscope or thoracoscope) of a surgical robot, including an image pair of left and right views, ensuring that the data covers various surgical scenarios and dynamic changes. This data includes images under different patients, different surgical stages, and different operation methods, ensuring that the model is exposed to rich variations. The collected images are preprocessed, including operations such as correction and denoising, to improve the accuracy of subsequent processing.
[0024] Then, a depth map of each frame is generated through a stereo matching algorithm, and the depth information of the tissue region is extracted by combining with a tissue mask. The depth map D i can be expressed as: D i =StereoMatching( I i,left , I i,right ) where I i,left and I i,right are the left and right view images of the i th frame respectively, and StereoMatching represents stereo matching.
[0025] The tissue mask of each frame is extracted through an automatic segmentation algorithm or a pre-trained segmentation model M i : M i =Segmentation( I i ) where I i is the input image of the i th frame, including the left and right views, and Segmentation represents segmentation.
[0026] S2: Construct a 3D reconstruction model based on 3DGS. The 3D reconstruction model includes: Initialization module: used to extract the 3D point cloud in each frame according to the depth map and tissue mask, and generate a dense initial Gaussian point cloud by fusing multi-frame point clouds through a binary motion mask; Flexible deformation model: used to assign dynamic attributes to each Gaussian point in the initial Gaussian point cloud, and learn the motion curve of each Gaussian point through a Gaussian-type combined basis function; Rendering module: used to convert the Gaussian point cloud into a 2D image and a depth map through a differential rendering algorithm; In this step, a 3D reconstruction model based on 3DGS is constructed, which mainly includes an initialization module, a flexible deformation model, and a rendering module.
[0027] The initialization module is used to: extract the three-dimensional point cloud in each frame based on the depth map and the tissue mask, design a binary motion mask for each frame, identify regions with significant depth changes, and fuse the point clouds of all frames to form a dense initial Gaussian point cloud. Specifically, the point cloud generation is based on the camera parameters K 1 and K 2 , convert the depth map into a three-dimensional point cloud P i :
[0028]
[0029] where P i is the three-dimensional point cloud in the i-th frame, I i is the i-th frame image, D i is the depth map of the current i-th frame, M i is the tissue mask of the i-th frame, K 1 、 K 2 are the camera intrinsic matrix and the extrinsic matrix respectively, and ⊙ represents element-wise multiplication.
[0030] Then, calculate the depth map difference between each frame and the reference frame (usually the first frame) to generate a binary motion mask B i : B i = I (∣ D 0 - D i ∣> τ )∪(1- M 0 )∩ M i where B i is the binary motion mask of the i-th frame, I (⋅) is the indicator function, M 0 is the tissue mask of the first frame, M i is the tissue mask of the current frame, D 0 is the depth map of the first frame, τis the threshold of depth change, which is used to determine whether the depth change of a pixel is significant.
[0031] After that, multiple frames of point clouds are fused through a binary motion mask to form a dense initial Gaussian point cloud P: P ={ P 0 , P 1 ⊙ B 1 ,…, P T ⊙ B T} Among them, P represents the initial Gaussian point cloud after fusion, P 0 represents the three-dimensional point cloud of the first frame, P i represents the i three-dimensional point cloud of the B i is the binary motion mask of the i frame, T represents the number of frames, i ∈ [1, T].
[0032] The flexible deformation model is used to: represent the dynamic scene by using a deformation field, assign dynamic attributes including position, rotation, and scaling to each Gaussian point in the initial Gaussian point cloud, and learn the motion curve of each Gaussian point through a Gaussian-type combined basis function; specifically, the flexible deformation model represents the dynamic scene by using a deformation field, associates a set of learnable weights ω and parameters Θ, and represents the motion curve of the Gaussian point by using a Gaussian-type basis function.
[0033] Specifically, the flexible deformation model learns the motion curve of each Gaussian point through a Gaussian-type basis function learns the motion curve of each Gaussian point, represents the complex motion of the tissue through the combination of Gaussian-type basis functions , and the position, rotation, and scaling changes of each Gaussian point are controlled by learnable weights, centers, and variances, where: (1) Among them, t represents time, θ and σ are learnable central positions and variances; Each Gaussian point is associated with a set of learnable weights ω and parameters Θ μ ,Θ r ,Θ sRelatedly, these weights and parameters linearly combine basis functions to represent deformations of position, rotation, and scaling, where: Θ μ The set of parameters representing position changes, including position change parameters along the x, y, and z axes Θ μ,x 、Θ μ,y and Θ μ,z , Θ μ,x 、Θ μ,y and Θ μ,z respectively include the corresponding weight, center, and variance parameters; Θ r The set of parameters representing rotation changes, including the weight, center, and variance of rotation; Θ s The set of parameters representing scaling changes, including the weight, center, and variance of scaling; For example, the change in the position of a Gaussian point along the X axis is represented by the combination of basis functions as: (2) where, ψ μ,x ( t ; Θ μ,x ) represents the change in the position of the Gaussian point along the x axis at time t , which is represented by the linear combination of basis functions; Θ μ,x The set of parameters representing the position change along the x axis, , including the weight 、center and variance , represents the contribution of the j -th basis function to the motion curve of the position change along the x axis, represents the central position of the j -th basis function in time, representing the j -th basis function's width in time; represents the j -th Gaussian basis function, whose calculation formula is formula (1); j is an index variable used to iterate through the set of basis functions, and the number of basis functions is determined by B , represented as j =1,2,…, B , and each basis function j has its own parameters, including the learnable weight ω j (controlling the importance of the basis function), the learnable center θ j and variance σ j parameters (controlling the central position and width of the basis function).
[0034] The position of the Gauss point along the x-axis at time t is expressed as: μ x ( t ) = μ x (0) + ψ μ,x ( t ) (3) Wherein, μ x ( t ) represents the position of the Gauss point along the x-axis at time t , μ x (0) represents the position of the Gauss point along the x-axis at the initial time t = 0, ψ μ,x ( t ) represents the change in the position of the Gauss point along the x-axis at time t , calculated from formula (2).
[0035] Similarly, the changes in the position of the Gauss point along the Y-axis and Z-axis are represented by a combination of basis functions as:
[0036] Wherein, ψ μ,y ( t ; Θ μ,y ) represents the change in the position of the Gauss point along the y-axis at time t , , are the weight, center, and variance parameters of the position change in the y-direction; ψ μ,z ( t ; Θ μ,z ) represents the change in the position of the Gauss point along the z-axis at time t , , are the weight, center, and variance parameters of the position change in the z-direction; The basis functions are all formula (1). The position update refers to the form of formula (3).
[0037] For the rotational quaternion components, the dynamic change formula for each component of the quaternion r = r 1 , r 2 , r 3 , r 4 T is:
[0038] Among them, , the updated quaternion needs to be normalized.
[0039] Scaling factor s = s x , s y , s z ] T The dynamic change formula of is:
[0040] Among them, .
[0041] Finally, the position, rotation, and scaling updates of each point cloud at time t are jointly achieved by the above parameters (the initial value plus the change amount).
[0042] The rendering module is used to: convert the Gaussian point cloud into a 2D image and a depth map through a differential rendering algorithm; specifically, the rendering module performs rendering based on the existing 3DGS differential algorithm, and the process is as follows: The influence of each Gaussian point is defined by the Gaussian distribution:
[0043] Among them, G ( x ) represents the value of the Gaussian distribution, which is used to measure the response intensity of point x at the Gaussian point μ , x represents a point in space, usually a three-dimensional vector x , y , z ], μ represents the center position of the Gaussian point, usually a three-dimensional vector μ x , μ y , μ z ], Σ represents the covariance matrix of the Gaussian point, which determines the shape and direction of the Gaussian distribution, and Σ −1 represents the inverse of the covariance matrix, which is used to calculate the exponential part of the Gaussian distribution.
[0044] Project the 3D Gaussian point onto the 2D image plane through the camera parameters, and calculate the projected covariance matrix Σ′: Σ′ = JW Σ W T JT Among them, W is the view transformation matrix, J is the Jacobian matrix of the projection, Σ is the 3D covariance matrix of the Gaussian points, T denotes the transpose.
[0045] Rasterize the projected 2D Gaussian, and calculate the pixel color and depth through alpha blending:
[0046] Among them, denotes the color value of the rendered pixel x , denotes the depth value of the rendered pixel x , N denotes the set of Gaussian points affecting the pixel x , c i denotes the color of the i -th Gaussian point, d i denotes the depth of the i -th Gaussian point; α i denotes the opacity of the i -th Gaussian point at the pixel x , where o i denotes the opacity parameter of the i -th Gaussian point, denotes the 2D Gaussian distribution value of the i -th Gaussian point at the pixel x .
[0047] S3: Use the pre - processed image, depth map, and tissue mask as training data to train the 3D reconstruction model, and optimize the explicit parameters of the Gaussian points and the weights and parameters of the basis functions of the flexible deformation model by minimizing the total loss function to improve the reconstruction accuracy of the model; In this step, the process of training the 3D reconstruction model includes: The model updates the dynamic attributes of the initial Gaussian point cloud according to the input depth map and tissue mask, including position, rotation, and scaling, etc.; The flexible deformation model adjusts the attributes of the Gaussian points according to the time series; The rendering module renders the updated Gaussian point cloud into a 2D image and a depth map; Calculate the total loss through the total loss function. The total loss includes photometric loss, normalized depth loss, and total variation loss. The photometric loss is used to optimize the color difference between the rendered image and the actual image. The normalized depth loss and the total variation loss are used to optimize the difference between the predicted depth map and the actual depth map; According to the calculated total loss, update the model parameters through the backpropagation algorithm to minimize the total loss function. Preferably, use an optimization algorithm (such as the Adam optimizer) to adjust the model parameters to minimize the loss function.
[0048] Specifically, by introducing the normalized depth loss, ensure that the predicted depth map is accurately aligned with the actual depth map; specifically, the calculation formula for the normalized depth regularization loss is:
[0049] Where, is the normalized depth regularization loss, M is the mask, and D norm are the normalized predicted depth map and the actual depth map respectively. ⊙ represents the element-wise multiplication operation, and the mask M is applied to the difference calculation of the depth map.
[0050] Furthermore, use the total variation loss to remove the noise in the depth map and apply the Canny edge detection as the mask to prevent over-smoothing of the edges; the calculation formula for the total variation loss is:
[0051] Where, L smooth represents the depth smoothing loss, which is used to remove the noise in the depth map and retain the depth details, represents the number of valid pixels in the predicted depth map, is the result of the Canny edge detection, indicating whether the pixel ( i , j ) is on the edge. If the pixel is on the edge, the value is 0, otherwise it is 1; represents the absolute value of the depth difference between the pixel ( i , j ) and its right pixel ( i +1, j ), represents the absolute value of the depth difference between the pixel ( i , j ) and its lower pixel ( i , j +1).
[0052] The total loss function includes photometric loss (color loss), depth loss, and depth smoothness loss:
[0053] Among them, L colour is the original photometric loss of 3DGS (calculated by existing methods), is the normalized depth loss, L smooth is the total variation loss, λ smooth is the weight coefficient of the smoothness loss.
[0054] By minimizing the total loss function, the explicit parameters of the Gaussian points and the weights and parameters of the basis functions of the flexible deformation model are optimized to improve the reconstruction accuracy of the model. The weights and parameters of the basis functions are optimized by minimizing the loss function to capture the changes in attributes such as the position, rotation, and scaling of the Gaussian points over time. Through explicit parameter optimization and flexible deformation models, the dynamic change characteristics in the surgical scene are captured to provide an accurate three-dimensional view to assist surgical operations. It is preferred to use the Adam optimizer to jointly optimize the Gaussian point attributes (position, rotation, scaling, transparency, color) and basis function parameters to achieve efficient parameter optimization, and use GPU acceleration to render and dynamically cull low-opacity Gaussian points (such as o i <0.01), parallelize the rasterization process, and accelerate the real-time rendering speed.
[0055] The purposes of training are: (1) Through dynamic tissue deformation modeling (using flexible deformation models): Learn the dynamic deformation patterns (such as pushing, pulling, cutting, suturing, etc.) of tissues in the surgical scene due to tool operations, and accurately simulate complex deformation processes. (2) High-precision three-dimensional reconstruction: Accurately reconstruct the geometric structure and surface details of tissues under the challenges of tool occlusion, lighting changes, and low-texture areas. (3) Real-time rendering ability: Through efficient parameter optimization, ensure that the model can generate three-dimensional scenes in real time during surgery (≥30 FPS) to meet the timeliness requirements of surgical navigation. (4) Improvement of generalization ability: Make the model adapt to different surgical scenes (such as laparoscopy, thoracoscopy) and tissue types (such as soft tissues, vascular structures), and enhance clinical applicability.
[0056] Among them, the role of basis function parameter learning is: By training, adjust the weights ( ω j ), centers ( θ j ), and variances ( σ j ) of the Gaussian-type basis functions so that they can be combined to represent complex deformations. For example, when the number of basis functions B = 17, the variance of the j-th basis function σ j= 1.0:
[0057] These parameters are automatically optimized through training, enabling the model to capture local deformation features at different time points (such as instantaneous displacements caused by tool pushing and pulling). The localization property of the basis functions allows the model to independently learn deformations in different time periods, avoiding distortions caused by global coupling.
[0058] The role of multi-degree-of-freedom modeling is to independently parameterize position (x, y, z), rotation (quaternion), and scaling (each axis), covering all degrees of freedom in three-dimensional space and ensuring accurate expression of complex deformations.
[0059] The role of depth-constrained joint optimization is to avoid optimization instability caused by differences in depth ranges, significantly improve geometric accuracy, enforce the consistency between the rendered depth and the true depth, and suppress noise at the same time. Through the total variation loss (L smooth ), noise is suppressed, the tissue boundary is protected by combining Canny edge detection, and the need for detail retention and smoothing is balanced. At the same time, complementary information is extracted from the stereo endoscope video using photometric loss (color consistency) and depth loss (geometric accuracy) to enhance the model's scene understanding ability.
[0060] The role of occluded area completion is to use multi-frame point cloud fusion to initialize and supplement the areas occluded by the tool in a single frame, improving the reconstruction integrity.
[0061] The role of dynamic mask generation is to screen dynamic areas through a motion mask ( B i ), focus on optimizing the significantly deformed parts, and fuse multi-frame point clouds to solve the single-frame occlusion problem. For example, the tissue occluded by the tool in the first frame may be visible in subsequent frames, and the model generates a more complete initial point cloud by fusing this information.
[0062] The main role of the flexible deformation model is to capture the dynamic change characteristics of Gaussian points during the surgical process and learn the change characteristics in the surgical scene. Specifically: During the surgical procedure, tissues and tools move and deform, causing changes in the properties such as the position, rotation, and scaling of Gaussian points. During the training process, the flexible deformation model learns the patterns and rules of these changes by optimizing the weights and parameters of the basis functions. Surgery is a dynamic process with temporal continuity and dependence. The flexible deformation model can capture the changes of Gaussian points in the time series, thus accurately representing the dynamic characteristics of the surgical scene. In addition, during the surgical procedure, tissues are stretched, compressed, moved, etc., resulting in changes in their shape and structure. The flexible deformation model can learn these deformation characteristics, thereby more accurately reconstructing and rendering the dynamic morphology of tissues. Further, the motion trajectory and mode of surgical tools also have an important impact on the dynamic changes of the surgical scene. The flexible deformation model can learn the motion patterns of tools, including operations such as entry, exit, rotation, shearing, etc., as well as the interaction between the tool and the tissue. As the surgery progresses, the occlusion relationship between the tissue and the tool changes continuously. The flexible deformation model can learn the change characteristics of these occlusion relationships during the training process, helping the model better process and recover the occluded areas.
[0063] In summary, through dynamic deformation modeling, deep constraint optimization, and efficient parameter learning in model training, high-precision real-time reconstruction of the intraoperative dynamic scene is achieved. By adapting the 3DGS technology to the intraoperative scene of surgical robots during the training process, the deficiencies of traditional methods in dynamic modeling, occlusion processing, and real-time performance are solved, providing reliable three-dimensional perception support for surgical navigation and robot operation.
[0064] Preferably, after the training is completed, the trained model parameters are saved, and the performance of the model is evaluated using the validation dataset to ensure the accuracy and robustness of the model in different surgical scenarios.
[0065] S4: Load the trained three-dimensional reconstruction model in the surgical navigation system to generate a three-dimensional view of the surgical scene in real time, and output for surgical navigation and robot control.
[0066] Specifically, load the trained three-dimensional reconstruction model and related modules for image preprocessing, depth map, and tissue mask extraction in the surgical navigation system. During the surgical procedure, images are collected in real time to generate a depth map and a tissue mask and input them into the model. The Gaussian point cloud is updated through the flexible deformation model to capture the dynamic changes of the surgical scene. A real-time three-dimensional view and depth map are generated through the rendering module for doctors to perform surgical navigation and operation. In the training and application process of the method of the present invention, through technical means such as multi-frame data fusion, depth regularization, and unsupervised depth smoothing, the efficiency and accuracy of the intraoperative dynamic three-dimensional reconstruction of the robot are ensured. And in the face of complex surgical scenarios and various interference factors, such as noise, occlusion, etc., the model can maintain stable performance and provide reliable reconstruction results.
[0067] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A method for 3D reconstruction of dynamic scenes during surgery by a surgical robot, characterized in that: include: S1: Acquiring an image sequence of a surgical scene, wherein the image sequence includes various surgical scenes and image pairs of dynamically changing left and right views; Preprocessing the image sequence, generating a depth map of each frame based on the preprocessed image pairs, and extracting a tissue mask in each frame; S2: constructing a three-dimensional reconstruction model based on 3DGS, wherein the three-dimensional reconstruction model includes: Initialization module: used to extract the 3D point cloud in each frame according to the depth map and tissue mask, and generate a dense initial Gaussian point cloud by fusing multiple frame point clouds through binary motion mask; Flexible deformation model: used for assigning dynamic attributes to each Gaussian point in the initial Gaussian point cloud, and learning the motion curve of each Gaussian point through a Gaussian type combined basis function; Rendering module: used to convert Gaussian point cloud into two-dimensional image and depth map through differential rendering algorithm; S3: using the preprocessed image, the depth map and the tissue mask as training data to train the three-dimensional reconstruction model, and optimizing the explicit parameters of the Gaussian points and the weights and parameters of the basis functions of the flexible deformation model by minimizing the total loss function; S4: Load the trained 3D reconstruction model into the surgical navigation system, generate a 3D view of the surgical scene in real time, and output it for surgical navigation and robot control.
2. The method for three-dimensional reconstruction of dynamic scenes during surgery by a surgical robot according to claim 1, characterized in that: In step S1, the depth map is generated by a stereo matching algorithm, and the tissue mask is extracted by an automatic segmentation algorithm or a pre-trained segmentation model.
3. The method for three-dimensional reconstruction of dynamic scenes during surgery by a surgical robot according to claim 1, characterized in that: In step S2, the calculation formula for extracting the three-dimensional point cloud in each frame based on the depth map and tissue mask is: in, P i is the 3D point cloud in the i-th frame, I i is the i-th frame image, D i is the depth map of the current i-th frame, M i is the tissue mask of the i-th frame, K 1. K 2 are the camera intrinsic parameter matrix and extrinsic parameter matrix respectively, and ⊙ represents element-by-element multiplication.
4. The method for 3D reconstruction of dynamic scenes during surgery by a surgical robot according to claim 3, characterized in that: In step S2, the method for generating the binary motion mask includes: Calculate the depth map difference between each frame and a reference frame to generate a binary motion mask, where the reference frame is the first frame; The generation formula of the binary motion mask is: B i = I (∣ D 0- D i ∣> τ )∪(1- M 0)∩ M i in, B i is the binary motion mask of the i-th frame, I (⋅) is the indicator function, M 0 is the tissue mask of the reference frame, M i is the organization mask of the current frame, D 0 is the depth map of the reference frame, τ It is the depth change threshold, which is used to determine whether the depth change of a pixel is significant.
5. The method for 3D reconstruction of dynamic scenes during surgery by a surgical robot according to claim 4, characterized in that: In step S2, multiple frame point clouds are fused by binary motion mask to generate dense initial Gaussian point clouds. The fusion formula is: P ={ P 0, P 1⊙ B 1,…, P T ⊙ B T} in, P represents the initial Gaussian point cloud after fusion, P 0 represents the 3D point cloud of the first frame, P i Indicates i 3D point cloud of the frame, B i It is i The binary motion mask of the frame, T Represents the frame number, i∈[1,T].
6. The method for three-dimensional reconstruction of dynamic scenes during surgery by a surgical robot according to claim 1, characterized in that: In step S2, the flexible deformation model is constructed by Gaussian basis function The combination represents the complex motion of the tissue, and the position, rotation, and scale changes of each Gaussian point are controlled by learnable weights, centers, and variances, where: (1) in, t Indicates time, θ and σ are the learnable center locations and variances; The change of the position of the Gaussian point along the X-axis is expressed by the combination of basis functions: (2) in, ψ μ,x ( t ;Θ μ,x ) represents the Gaussian point at time t The position change along the x-axis; Θ μ,x A set of parameters representing position changes along the x-axis, , Indicates j The contribution of the basis functions to the motion curve of the position change along the x-axis, Indicates j The center position of the basis function in time, indicating the j The width of the basis functions in time; Indicates j Gaussian basis functions are calculated using formula (1); j is an index variable used to traverse the set of basis functions. The number of basis functions is determined by B Decide, expressed as j =1,2,…, B , each basis function j Has its own parameters, including weights ω ,center θ and variance σ ; The position of the Gaussian point along the x-axis at time t When expressed as: μ x ( t )= μ x (0)+ ψ μ,x ( t ) in, μ x ( t ) represents the Gaussian point at time t The position along the x-axis, μ x (0) indicates that the Gaussian point is at the initial time t =0 along the x-axis, ψ μ,x ( t ) represents the Gaussian point at time t The position change along the x-axis is calculated by formula (2).
7. The method for three-dimensional reconstruction of dynamic scenes during surgery by a surgical robot according to claim 1, characterized in that: In step S3, the process of training the three-dimensional reconstruction model includes: The model updates the dynamic properties of the Gaussian point cloud based on the input depth map and tissue mask; The flexible deformation model adjusts the properties of the Gaussian points according to the time series; The rendering module renders the updated Gaussian point cloud into a two-dimensional image and a depth map; Calculating a total loss through a total loss function, the total loss includes a photometric loss, a normalized depth loss, and a total variation loss, the photometric loss is used to optimize the color difference between the rendered image and the actual image, and the normalized depth loss and the total variation loss are used to optimize the difference between the predicted depth map and the actual depth map; According to the calculated total loss, the model parameters are updated through the back-propagation algorithm to minimize the total loss function.
8. The method for three-dimensional reconstruction of dynamic scenes during surgery by a surgical robot according to claim 7, characterized in that: The calculation formula of the normalized depth regularization loss is: in, is the normalized depth regularization loss, M is a mask, and D norm They are the normalized predicted depth map and the actual depth map, ⊙ represents the element-level multiplication operation, and the mask M Applied to the difference calculation of the depth map.
9. The method for 3D reconstruction of dynamic scenes during surgery by a surgical robot according to claim 8, characterized in that: The calculation formula of the total variation loss is: in, L smooth represents the depth smoothing loss, which is used to remove noise in the depth map and preserve depth details. Represents the number of valid pixels in the predicted depth map, is the result of Canny edge detection, indicating the pixel ( i , j ) Whether it is on the edge, if the pixel is on the edge, the value is 0, otherwise it is 1; Represents pixels ( i , j ) and its right pixel ( i +1, j ), Represents pixels ( i , j ) and its lower pixel ( i , j +1) is the absolute value of the depth difference.
10. The method for 3D reconstruction of dynamic scenes during surgery by a surgical robot according to claim 1, characterized in that: The total loss function is: in, L colour is the original photometric loss of 3DGS, is the normalized depth loss, L smooth is the total variation loss, λ smooth is the weight coefficient of the smoothing loss.
Citation Information
Patent Citations
Sparse visual angle three-dimensional reconstruction method based on depth prior information
CN118657888A
Surgical operation force feedback guiding method based on virtual mark tracking and instrument pose
CN119055358A
Laser enhanced vision three-dimensional reconstruction method and system based on Gaussian splashing
CN119180908A
Heart interventional operation scene reconstruction method based on 3D ultrasonic imaging rendering
CN119206038A
Dynamic real-time rendering method for large assembly scene based on three-dimensional Gaussian splashing
CN119229031A
Cited By
Overall dual-constraint dynamic registration navigation system for intraoperative scene
CN120451463A
Intraoperative scene overall dual-constraint dynamic registration navigation system
CN120451463B
Endoscope image dynamic reconstruction method based on point cloud compensation and double-domain deformation
CN120510302A
Monocular video scene dynamic three-dimensional reconstruction method based on optical flow
CN120747366A
Visual tracking method for processing flexible linear object with partial shielding
CN121236107A