Endoscope visual field scene reconstruction method, device, equipment and medium
Through the monocular depth prediction network and dynamic Gaussian growth module, the problems of scene complexity and movement difficulties under laparoscopic vision are solved, and high-quality 3D scene reconstruction of organ tissues under laparoscopic vision is realized, assisting doctors in performing precise surgical operations.
Patent Information
- Application Number
- CN202510086518.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The scenes in laparoscopic field of vision are more complex, and it is difficult to obtain motion through tracking devices. Laparoscopic motion is limited by the characteristics of the instrument, making it difficult to observe the overall picture of the cavity, making it difficult to use structure to restore sparse point clouds from motion algorithms and estimate absolute camera postures.
A monocular depth prediction network is used to estimate the continuous frame RGB organ images to obtain the local depth map, and the inter-pose estimation module is used to estimate and optimize the front and back frame camera transform pose matrix. The dynamic Gaussian growth module extends the Gaussian model to realize the sequential dynamic growth and global alignment of the Gaussian model.
Without relying on SfM preprocessing or other prior information, high-quality 3D scene reconstruction of organ tissue under laparoscopic field of vision is achieved, providing doctors with more three-dimensional 3D shape information of organ tissue in laparoscopic field of vision, assisting doctors in achieving more accurate and effective laparoscopic-guided minimally invasive surgery.
Smart Images

Figure CN120014165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual image processing, and in particular to a method, device, equipment and medium for reconstructing an endoscope field of view scene based on 3D Gaussian rendering. Background Art
[0002] Laparoscopically guided minimally invasive surgery has the advantages of less trauma, faster recovery, and less damage to organ and tissue function, which is crucial for robot-assisted minimally invasive surgery. This technology can help simulate the surgical environment and be used for preoperative planning and augmented reality / virtual reality surgical navigation.
[0003] Due to the complex internal tissue structure of the abdominal cavity, the narrow field of view and the deep anatomical position, it is difficult to achieve high-quality 3D surface reconstruction using traditional methods, limited by the narrow movement space and visual obstacles in the cavity.
[0004] With the emergence of Neural Radiance Fields (NeRF) and Neural Implicit Surfaces (NeuS), the field of scene reconstruction and view synthesis has been greatly developed in recent years. In addition, more recent efforts have focused on representing surgical scenes as radiance fields, which can learn continuous functions that implicitly represent 3D scenes trained from 2D images and paired camera poses. Neural field-based methods leverage deep neural networks to implicitly model complex geometry and appearance, outperforming methods based on discrete representations.
[0005] However, although the neural field-based method has achieved quite good results, a large number of points and rays need to be repeatedly queried from the radiation field during the rendering process of each image. This method requires a significant limit on the rendering speed and brings a huge obstacle to practical intraoperative application.
[0006] To address these limitations, the volume rendering in NeRF can be extended to accommodate high-quality point clouds using the 3D Gaussian Splatting (3DGS) method. 3DGS can describe the scene as an anisotropic representation and render the image using an efficient tile-based rasterizer, thus achieving real-time rendering with excellent reconstruction quality.
[0007] Nevertheless, it is troublesome to obtain motion by tracking devices due to the more complex scenes under the laparoscope field of view. In addition, the motion of the laparoscope is usually limited by the characteristics of the instrument, which makes it difficult for the endoscope to observe the full picture of the cavity. Typical 3DGS mainly relies on the Structure from Motion (SfM) algorithm to initialize the Gaussian position, but it has strict requirements on the motion trajectory and field of view of the laparoscope. In addition, it is difficult for doctors to obtain a panoramic view of the inside of the abdominal cavity by manipulating a handheld laparoscope around the patient's organs. Therefore, the field of view of the laparoscope is usually limited to a limited range of motion, which makes it difficult to recover sparse point clouds and estimate absolute camera poses using the SfM pipeline. Summary of the invention
[0008] In view of the above problems, the present invention provides an endoscope field of view scene reconstruction method, device, equipment and medium for overcoming the above problems or at least partially solving the above problems.
[0009] The present invention provides the following scheme:
[0010] A method for reconstructing an endoscope visual field scene, comprising:
[0011] Acquire a number of continuous frame RGB organ images collected through an endoscope;
[0012] Using a monocular depth prediction network to estimate the RGB organ images of several consecutive frames respectively to obtain several local depth maps;
[0013] Using the local depth maps of two adjacent frames, an inter-frame pose estimation module is used to estimate and optimize the pose matrix of the camera transformation of the previous and next frames;
[0014] Obtaining XYZ parameters of the Gaussian model of the previous frame by calculation according to the local depth map of the previous frame having the intrinsic parameter matrix;
[0015] The coordinate system of the Gaussian model of the previous frame is transformed from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames;
[0016] The RGB organ image of the next frame and the predicted corresponding local depth map are processed by a dynamic Gaussian growth module, and the local Gaussian set appearing in the new area under the viewpoint of the next frame is extended from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image;
[0017] The local Gaussian models corresponding to each frame of the RGB organ image are fused to obtain a global Gaussian model.
[0018] Preferably, using the local depth maps of two adjacent frames and adopting an inter-frame pose estimation module to estimate and optimize the pose matrix of the camera transformation of the previous and next frames includes:
[0019] Using the dense key point matching of the feature matching network and the relative depth calculated by the dense depth estimation network, a 2D-3D displacement field is established using the 3D coordinates of the organ map of the previous frame and the 2D corresponding position of the organ map of the next frame;
[0020] The camera transformation pose matrix of the previous and next frames is obtained by setting the optimization objective to minimize the Gaussian rendering variational gradient of the displacement field.
[0021] Preferably: using the intrinsic matrix and distortion parameter transformation in the local camera coordinate system to obtain the corresponding point cloud;
[0022] The 3D displacement vector field of the organ image of the previous frame is obtained by constructing the local depth map and the point cloud, and Kp N and Kp N+1 Construct a 2D displacement vector field;
[0023] Through variational optimization of the displacement vector field and the energy functional of the determined vector field under different transformation matrix positions, the initial estimation of the transformation matrix position is achieved based on the 2D vector after optimization of the 3D vector field.
[0024] Based on the initial value of the initial estimate of the transformation matrix pose, the energy function target of the variational optimization is determined to be the minimum image similarity difference between the Gaussian rendering image and the original RGB organ image at the current viewing angle as a generalization function, and the optimized front and back frame camera transformation pose matrix is calculated.
[0025] Preferably, using a dynamic Gaussian growth module to process the RGB organ image of the next frame and the predicted corresponding local depth map includes:
[0026] Based on the local depth map of the next frame, a new Gaussian ellipsoid is newly obtained in a blank area under a new viewpoint of the next frame using the local depth information and the corresponding RGB spherical harmonic function information as initial Gaussian attributes;
[0027] Optimizing the properties of the newly created Gaussian ellipsoid based on a differential gradient field optimization algorithm, and using the loss function of the transferred depth and RGB as an optimization target;
[0028] Based on the growing Gaussian set under the local perspective of the next frame, the grown Gaussian is aligned with the global Gaussian, thereby realizing the growth of the Gaussian model under the viewpoint of the next new frame.
[0029] Preferably: in the training process, a strategy of forward reconstruction and reverse optimization is adopted to ensure the coherence of sequence frames and to achieve global alignment of the Gaussian model.
[0030] Preferably: the forward reconstruction and reverse optimization strategies include:
[0031] For the initial frame, the viewpoint is initialized to the local camera coordinate system of the first frame, and the corresponding point cloud is calculated based on the local depth map; the XYZ attributes of the Gaussian model are initialized with the point cloud, and the spherical harmonic function SHS parameters are initialized with the RGB values of the frame;
[0032] In the forward processing of sequential video frames, the inter-frame pose estimation module is used to estimate the inter-frame pose transformation matrix, and the dynamic Gaussian growth module is used to realize the gradual growth of the Gaussian model;
[0033] In the backward process of sequential video frames, cyclic training is implemented by reversing the order of video frames so that the estimated Gaussian parameters are optimized during the forward reconstruction and the gradients are back-propagated.
[0034] An endoscope field of view scene reconstruction device, used to execute the above-mentioned endoscope field of view scene reconstruction method, the device comprising:
[0035] An image acquisition unit, used to acquire a plurality of continuous frame RGB organ images collected through the endoscope;
[0036] A depth map estimation unit, used to estimate the RGB organ images of several consecutive frames respectively using a monocular depth prediction network to obtain several local depth maps;
[0037] A pose transformation matrix determination unit is used to estimate and optimize the camera transformation pose matrix of the previous and next frames by using the local depth maps of two adjacent frames and adopting an inter-frame pose estimation module;
[0038] An initial Gaussian model parameter determination unit, configured to calculate and obtain XYZ parameters of a Gaussian model of a previous frame according to the local depth map of the previous frame having an intrinsic parameter matrix;
[0039] A Gaussian model coordinate system transformation unit, used to transform the coordinate system of the Gaussian model of the previous frame from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames;
[0040] A Gaussian model sequential dynamic growth unit is used to process the RGB organ image of the next frame and the predicted corresponding local depth map using a dynamic Gaussian growth module, and expand the local Gaussian set appearing in the new area under the viewpoint of the next frame from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image;
[0041] The global Gaussian model acquisition unit is used to fuse the local Gaussian models corresponding to each frame of the RGB organ image to obtain a global Gaussian model.
[0042] An endoscope field of view scene reconstruction device, the device comprising a processor and a memory:
[0043] The memory is used to store program code and transmit the program code to the processor;
[0044] The processor is used to execute the above-mentioned endoscope field of view scene reconstruction method according to the instructions in the program code.
[0045] A computer-readable storage medium is used to store program codes, and the program codes are used to execute the above-mentioned endoscope field of view scene reconstruction method.
[0046] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0047] The embodiments of the present application provide a method, device, equipment and medium for reconstructing an endoscopic field of view scene. The method uses the original laparoscopic video stream image to learn and optimize the variational gradient of Gaussian rasterization, and designs a new inter-frame posture estimation module to perform accurate micro-pose estimation. Based on the differential gradient field optimization, a new dynamic Gaussian growth strategy is designed for new Gaussian synthesis within the constrained endoscopic view. Combining the above two modules, a new 3D Gaussian training architecture with forward reconstruction and backward optimization strategies is proposed to sequentially generate Gaussian models and solve the limitations of SfM priors. Improved relative posture estimation and new Gaussian growth strategies bring better new view synthesis quality and geometric reconstruction of organ tissue models under laparoscopic field of view. Thereby, more three-dimensional 3D shape information of organ tissues in the laparoscopic field of view is provided to doctors during surgery, assisting doctors in achieving more accurate and effective laparoscopic guided minimally invasive surgery.
[0048] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 is a flow chart of an endoscope field of view scene reconstruction method provided by an embodiment of the present invention;
[0051] Figure 2 It is a flowchart of monocular laparoscope field of view scene reconstruction based on 3D Gaussian rendering provided by an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of inter-frame pose estimation of a laparoscope provided by an embodiment of the present invention;
[0053] Figure 4 is a dynamic Gaussian model growth flow chart provided by an embodiment of the present invention;
[0054] Figure 5 It is a schematic diagram of the Gaussian training strategy of the forward reconstruction and backward optimization architecture provided by an embodiment of the present invention;
[0055] Figure 6 is a schematic diagram of an endoscope visual field scene reconstruction device provided by an embodiment of the present invention;
[0056] Figure 7 It is a schematic diagram of an endoscope field of view scene reconstruction device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.
[0058] See also Figure 1 , is a method for reconstructing an endoscope visual field scene provided by an embodiment of the present invention, such as Figure 1 As shown, the method may include:
[0059] S101: Acquire a number of continuous frames of RGB organ images collected through an endoscope;
[0060] S102: using a monocular depth prediction network to estimate the RGB organ images of several consecutive frames to obtain several local depth maps;
[0061] S103: using the local depth maps of two adjacent frames to use an inter-frame pose estimation module to estimate and optimize to obtain a front and back frame camera transformation pose matrix; in specific implementation, the embodiment of the present application can provide that using the local depth maps of two adjacent frames to use an inter-frame pose estimation module to estimate and optimize to obtain a front and back frame camera transformation pose matrix includes:
[0062] Using the dense key point matching of the feature matching network and the relative depth calculated by the dense depth estimation network, a 2D-3D displacement field is established using the 3D coordinates of the organ map of the previous frame and the 2D corresponding position of the organ map of the next frame;
[0063] The camera transformation pose matrix of the previous and next frames is obtained by setting the optimization objective to minimize the Gaussian rendering variational gradient of the displacement field.
[0064] Furthermore, the corresponding point cloud is obtained using the intrinsic matrix and distortion parameter transformation in the local camera coordinate system;
[0065] The 3D displacement vector field of the organ image of the previous frame is obtained by constructing the local depth map and the point cloud, and Kp N and Kp N+1 Construct a 2D displacement vector field;
[0066] Through variational optimization of the displacement vector field and the energy functional of the determined vector field under different transformation matrix positions, the initial estimation of the transformation matrix position is achieved based on the 2D vector after optimization of the 3D vector field.
[0067] Based on the initial value of the initial estimate of the transformation matrix pose, the energy function target of the variational optimization is determined to be the minimum image similarity difference between the Gaussian rendering image and the original RGB organ image at the current viewing angle as a generalization function, and the optimized front and back frame camera transformation pose matrix is calculated.
[0068] S104: Calculate and obtain XYZ parameters of the Gaussian model of the previous frame according to the local depth map of the previous frame with the internal parameter matrix;
[0069] S105: transforming the coordinate system of the Gaussian model of the previous frame from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the previous and next frame camera transformation pose matrix;
[0070] S106: Processing the RGB organ image of the next frame and the predicted corresponding local depth map by using a dynamic Gaussian growth module, extending the local Gaussian set appearing in the new area under the viewpoint of the next frame from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image; in specific implementation, the embodiment of the present application can provide the local depth map of the next frame, and obtain a new Gaussian ellipsoid in the blank area under the new viewpoint of the next frame with the local depth information and the corresponding RGB spherical harmonic function information as the initial Gaussian attributes;
[0071] Optimizing the properties of the newly created Gaussian ellipsoid based on a differential gradient field optimization algorithm, and using the loss function of the transferred depth and RGB as an optimization target;
[0072] Based on the growing Gaussian set under the local perspective of the next frame, the grown Gaussian is aligned with the global Gaussian, thereby realizing the growth of the Gaussian model under the viewpoint of the next new frame.
[0073] S107: Fusing the local Gaussian models corresponding to each frame of the RGB organ image to obtain a global Gaussian model.
[0074] In order not to rely on SfM preprocessing or other prior information, based on the relative posture between adjacent frames and the dynamic Gaussian growth module, the embodiments of the present application can provide a strategy of forward reconstruction and reverse optimization during the training process to ensure the continuity of sequence frames and achieve global alignment of the Gaussian model.
[0075] Furthermore, the forward reconstruction and reverse optimization strategies include:
[0076] For the initial frame, the viewpoint is initialized to the local camera coordinate system of the first frame, and the corresponding point cloud is calculated based on the local depth map; the XYZ attributes of the Gaussian model are initialized with the point cloud, and the spherical harmonic function SHS parameters are initialized with the RGB values of the frame;
[0077] In the forward processing of sequential video frames, the inter-frame pose estimation module is used to estimate the inter-frame pose transformation matrix, and the dynamic Gaussian growth module is used to realize the gradual growth of the Gaussian model;
[0078] In the backward process of sequential video frames, cyclic training is implemented by reversing the order of video frames so that the estimated Gaussian parameters are optimized during the forward reconstruction and the gradients are back-propagated.
[0079] In order to solve the technical challenges of 3D surface reconstruction brought by monocular laparoscopic guided minimally invasive surgery, Gaussian model reconstruction is performed without camera pose prior or SfM preprocessing. This method involves relative pose estimation of monocular laparoscope and geometric reconstruction of Gaussian model.
[0080] The main purpose of SfM preprocessing is to extract useful information from the original image and provide accurate basic data for the subsequent 3D reconstruction steps. The accuracy of feature extraction and matching directly affects the quality and effect of subsequent 3D reconstruction. Through preprocessing, noise and false matching can be reduced, and the accuracy and stability of reconstruction can be improved. Preprocessing plays a vital role in the SfM system. It provides reliable basic data for subsequent 3D reconstruction steps, ensuring the accuracy and reliability of 3D reconstruction. Through feature extraction and matching, useful feature points can be effectively extracted from multiple images, and outliers can be excluded through the RANSAC method, thereby improving the accuracy of the basic matrix F. These steps provide a solid foundation for subsequent 3D reconstruction.
[0081] The endoscopic field of view scene reconstruction method provided in the embodiment of the present application can realize a 3D scene reconstruction system of organ tissues under the endoscopic field of view. Without relying on the endoscopic posture tracking device and other tracking markers, it only relies on the endoscopic RGB video information to realize the estimation of the relative posture movement of the endoscope in the front and back frames, and based on the real-time surface reconstruction of the intracavitary scene in the 3D Gaussian rendering mode, realizes the surface three-dimensional reconstruction of the organ tissues in the endoscopic field of view without camera posture prior or motion structure reconstruction preprocessing. This method can assist surgeons in real-time perception of the surface three-dimensional model information of the internal organs and tissues of the abdominal cavity during endoscopic-guided minimally invasive surgery, enrich the spatial perception and stereoscopic perception of the organs and tissues under the endoscopic field of view during the operation, and thus improve the success rate of the operation.
[0082] The method provided in the embodiment of the present application is described in detail below, taking an endoscope as a laparoscope to achieve three-dimensional reconstruction of the surface of internal abdominal organ tissue as an example.
[0083] The embodiment of the present application provides a reconstruction method based on 3D Gaussian splatting to achieve three-dimensional surface reconstruction of intracavitary tissues and organs under the field of view of a monocular laparoscope, thereby assisting doctors in observing a more three-dimensional 3D intracavitary reconstruction scene during surgery, reducing the uncertainty of individual operations and the dependence on the doctor's clinical experience to the greatest extent, and improving the success rate of laparoscopic guided minimally invasive surgery. Gaussian splatting achieves the goal of rendering realistic scenes from small image samples in real time through the rasterization method of Gaussian distribution. Its core lies in the rasterization technology, which constructs the Gaussian distribution of each point and optimizes the parameters to achieve a delicate reproduction of the scene.
[0084] The technical problems to be solved by the method provided in this application include:
[0085] 1. Small relative pose estimation between the front and back frames of laparoscopic video
[0086] To address the challenge of estimating small relative poses between adjacent frames, we designed a novel sequential 2D-3D displacement field variational optimization method to achieve continuous small relative camera pose transformations. By using dense keypoint matching of a feature matching network and relative depth calculated by a dense depth estimation network, a 2D-3D displacement field is established using the 3D coordinates of the previous frame and the 2D corresponding positions of the next frame. The relative pose transformation of the laparoscopic anterior and posterior frames is learned by setting the optimization objective to minimize the Gaussian rendered variational gradient of the displacement field.
[0087] 2. Local Gaussian dynamic growth algorithm under restricted field of view.
[0088] In order to perform Gaussian 3D reconstruction within the restricted endoscopic field of view, a sequential dynamic Gaussian growing module is developed based on the Gaussian dynamic transformation of relative pose to perform Gaussian growing step by step. In each rendering iteration, the entire Gaussian model is transformed into the current laparoscopic camera local coordinate system, and a new Gaussian ellipsoid is used to fit the constrained area in the 2D rendered image, and a new differential gradient field optimization method is proposed to refine the XYZ coordinates of the local Gaussian model in the absence of multi-view geometric constraints.
[0089] 3. Gaussian dynamic training strategy based on sequential video frames.
[0090] In order to not rely on SfM preprocessing or other prior information, based on the relative pose between adjacent frames and the dynamic Gaussian growth module, this method proposes a new 3D Gaussian model reconstruction architecture with forward reconstruction and reverse optimization strategies for frame-by-frame training and generation of Gaussian ellipsoids under new viewpoints. The proposed architecture can continuously input infinite frames and consistently perform Gaussian rasterization operations.
[0091] To perform Gaussian model reconstruction without camera pose prior or SfM preprocessing, this method establishes a new dynamically growing 3D Gaussian rendering (DG-3DGS) architecture by sequentially growing Gaussian models during the movement of the laparoscope.
[0092] Firstly, a 2D-3D displacement field is designed based on dense feature matching and depth prediction to establish spatial feature association, and the relative posture is obtained by minimizing the energy functional of variational optimization of the displacement field.
[0093] Secondly, the local Gaussian model is sequentially grown through Gaussian dynamic transformation and differential gradient field optimization;
[0094] Finally, a global Gaussian model is generated based on the forward reconstruction & backward optimization architecture, thereby achieving surface reconstruction of organ tissue scenes under the laparoscopic field of view.
[0095] like Figure 2 The figure shows the overall process diagram implemented in this application.
[0096] 1. Laparoscopic anterior and posterior inter-frame pose estimation.
[0097] At the beginning of reconstruction, the first frame with the predicted depth map is regarded as the Gaussian initialization step, and the current Gaussian model is in the local camera coordinate system (LCCS) corresponding to the current frame. The XYZ parameters of the initial Gaussian model are calculated according to the depth map with the intrinsic parameter matrix. When the next frame flows in (for example, from View N To View N+1 ), use the laparoscope front and back frame pose estimation module to estimate and optimize the relative transformation pose matrix (such as RN →R N+1 , T N →T N+1 ). The Gaussian model of the current frame is converted to the local camera coordinate system of the next frame through the camera transformation matrix of the previous and next frames obtained by the relative pose estimation module.
[0098] 2. Dynamic Gaussian growth strategy.
[0099] By adding R N →R N+1 , T N →T N+1 Set as the transformation matrix to move the Gaussian model from View N The local camera coordinate system (LCCS) is transformed into View N+1 The local camera coordinate system, the dynamic Gaussian growth module according to the current frame I N+1 And the predicted corresponding monocular depth map D N+1 Processing is performed to remove the local Gaussian set appearing in the new area under the current viewpoint from the Gaussian model G N Expand to G N+1 .
[0100] 3. Gaussian dynamic training strategy for sequential video frames
[0101] Based on the above-mentioned inter-frame pose estimation and dynamic Gaussian growth strategy, the local Gaussian growth module of each frame is used to correspond to the global Gaussian model. During the training process, the forward reconstruction + reverse optimization strategy is adopted to ensure the coherence of the sequence frames and achieve global alignment of the Gaussian model.
[0102] like Figure 3 As shown, the inter-frame pose estimation module provided in an embodiment of the present application is described.
[0103] Step 1: Given the camera viewpoint View N and View N+1 The main goal of this module is to estimate the relative pose transformation R between the right-handed camera coordinate system N →R N+1 , T N →T N+1 The local depth map D is estimated using an end-to-end pre-trained monocular depth prediction network DPT. N , and use the intrinsic matrix K in the local camera coordinate system N and the distortion parameter Dist N Convert the corresponding point cloud Pt N .
[0104] Step 2: Using the local depth map D N And point cloud Pt N Construct a 3D displacement vector field through KpN and Kp N+1 Construct a 2D displacement vector field.
[0105] Step 3: Through variational optimization of the displacement vector field, set the energy functional of the vector field under different RT postures, and realize the initial estimation of the RT posture based on the 2D vector after optimization of the 3D vector field.
[0106] Step 4: Based on the initial value, the energy function target of the variational optimization is set to the minimum image similarity difference (PSNR) between the Gaussian rendered image and the original RGB image at the current viewing angle as the generalization function, and the optimized transformation matrix RT is calculated.
[0107] like Figure 4 As shown, the dynamic Gaussian growth module provided in the embodiment of the present application is described in detail.
[0108] Step 1: At each viewpoint, when the Gaussian model is transformed from View N Switch to View N+1 Due to the incompleteness of the initial Gaussian model, there are model defect areas on the raster view, and the corresponding representation on the rendered image will be the appearance of blank areas. opt and T opt , through the new viewpoint from View N Switch to View N+1 Dynamically convert Gaussian models from G N Convert to G N+1 .
[0109] Step 2: Based on the monocular depth estimation map, a new Gaussian ellipsoid is created in the blank area under the new viewpoint using the local depth information and the corresponding RGB spherical harmonic function information as the initial Gaussian attributes.
[0110] Step 3: Optimize the properties of the newly created Gaussian ellipsoid based on the differential gradient field optimization algorithm, and pass the depth and RGB loss functions as optimization targets.
[0111] Step 4: Align the grown Gaussian with the global Gaussian based on the grown Gaussian set under the local perspective, thereby achieving the growth of the Gaussian model under the new frame viewpoint.
[0112] like Figure 5 As shown, the Gaussian training strategy of the forward reconstruction & backward optimization architecture provided in the embodiment of the present application is described in detail.
[0113] Step 1: For the initial frame, initialize the viewpoint to the local camera coordinate system of the first frame and calculate the corresponding point cloud based on the depth map. The XYZ attributes of the Gaussian model are initialized with the point cloud, and the spherical harmonic function SHS parameters are initialized with the RGB values of the frame.
[0114] Step 2: In the forward processing of sequential video frames, the inter-frame pose estimation module is used to estimate the inter-frame pose transformation matrix, and the dynamic Gaussian growth module is used to realize the gradual growth of the Gaussian model. 0 →View 1 、View 1 →View 2 、…、View N-1 →View N To grow a new local Gaussian set. The forward reconstruction process aims to maintain the integrity of the Gaussian reconstruction process as much as possible.
[0115] Step 3: In the reverse processing of sequential video frames, since the estimated pose represents the transformation matrix between each pair of frames, the gradient of the Gaussian model needs to be continuous during the training process. Therefore, the cyclic training is achieved by reversing the order of the video frames. The reverse training order is: N →View N-1 ,View N-1 →View N-2 ,…,View 1 →View 0 The backward optimization process aims to optimize the estimated Gaussian parameters during the forward reconstruction process and back-propagate the gradients.
[0116] In summary, the endoscopic field of view scene reconstruction method provided in this application uses the original laparoscopic video stream image to learn and optimize the variational gradient of Gaussian rasterization, and designs a new inter-frame pose estimation module for accurate micro-pose estimation. Based on the differential gradient field optimization, a new dynamic Gaussian growth strategy is designed for new Gaussian synthesis within the constrained endoscopic view. Combining the above two modules, a new 3D Gaussian training architecture with forward reconstruction and backward optimization strategies is proposed to sequentially generate Gaussian models and solve the limitations of the SfM prior. Improved relative pose estimation and new Gaussian growth strategies bring better new view synthesis quality and geometric reconstruction of organ tissue models under the laparoscopic field of view. Thereby providing doctors with more three-dimensional 3D shape information of organs and tissues in the laparoscopic field of view during surgery, assisting doctors in achieving more accurate and effective laparoscopic-guided minimally invasive surgery.
[0117] See also Figure 6 , the embodiment of the present application may also provide an endoscope field of view scene reconstruction device, such as Figure 6 As shown, the device for executing the above-mentioned endoscope field of view scene reconstruction method may include:
[0118] An image acquisition unit 601 is used to acquire a plurality of continuous frames of RGB organ images collected by an endoscope;
[0119] A depth map estimation unit 602 is used to estimate the RGB organ images of several consecutive frames respectively using a monocular depth prediction network to obtain several local depth maps;
[0120] A pose transformation matrix determination unit 603 is used to estimate and optimize the pose transformation matrix of the camera of the previous and next frames by using the local depth maps of the two adjacent frames and adopting the inter-frame pose estimation module;
[0121] An initial Gaussian model parameter determination unit 604 is used to calculate and obtain XYZ parameters of the Gaussian model of the previous frame according to the local depth map of the previous frame having an intrinsic parameter matrix;
[0122] A Gaussian model coordinate system transformation unit 605 is used to transform the coordinate system of the Gaussian model of the previous frame from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames;
[0123] The Gaussian model sequential dynamic growth unit 606 is used to process the RGB organ image of the next frame and the predicted corresponding local depth map by using a dynamic Gaussian growth module, and expand the local Gaussian set appearing in the new area under the viewpoint of the next frame from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image;
[0124] The global Gaussian model acquisition unit 607 is used to fuse the local Gaussian models corresponding to each frame of the RGB organ image to obtain a global Gaussian model.
[0125] The embodiment of the present application may also provide an endoscope field of view scene reconstruction device, the device comprising a processor and a memory:
[0126] The memory is used to store program code and transmit the program code to the processor;
[0127] The processor is used to execute the steps of the above-mentioned endoscope field of view scene reconstruction method according to the instructions in the program code.
[0128] like Figure 7 As shown, an endoscope visual field scene reconstruction device provided in an embodiment of the present application may include: a processor 10, a memory 11, a communication interface 12 and a communication bus 13. The processor 10, the memory 11 and the communication interface 12 communicate with each other through the communication bus 13.
[0129] In the embodiment of the present application, the processor 10 may be a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic devices, etc.
[0130] The processor 10 may call a program stored in the memory 11. Specifically, the processor 10 may execute operations in an embodiment of the endoscope field of view scene reconstruction method.
[0131] The memory 11 is used to store one or more programs, which may include program codes, and the program codes include computer operation instructions. In the embodiment of the present application, the memory 11 at least stores programs for implementing the following functions:
[0132] Acquire a number of continuous frame RGB organ images collected through an endoscope;
[0133] Using a monocular depth prediction network to estimate the RGB organ images of several consecutive frames respectively to obtain several local depth maps;
[0134] Using the local depth maps of two adjacent frames, an inter-frame pose estimation module is used to estimate and optimize the pose matrix of the camera transformation of the previous and next frames;
[0135] Obtaining XYZ parameters of the Gaussian model of the previous frame by calculation according to the local depth map of the previous frame having the intrinsic parameter matrix;
[0136] The coordinate system of the Gaussian model of the previous frame is transformed from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames;
[0137] The RGB organ image of the next frame and the predicted corresponding local depth map are processed by a dynamic Gaussian growth module, and the local Gaussian set appearing in the new area under the viewpoint of the next frame is extended from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image;
[0138] The local Gaussian models corresponding to each frame of the RGB organ image are fused to obtain a global Gaussian model.
[0139] In one possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required for at least one function (such as a file creation function, a data reading and writing function), etc.; the data storage area can store data created during use, such as initialization data, etc.
[0140] In addition, the memory 11 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.
[0141] The communication interface 12 may be an interface of a communication module, and is used to connect to other devices or systems.
[0142] Of course, it should be noted that Figure 7 The structure shown does not constitute a limitation on the endoscope field of view scene reconstruction device in the embodiment of the present application. In actual applications, the endoscope field of view scene reconstruction device may include Figure 7 More or fewer components than shown, or combinations of certain components.
[0143] The embodiment of the present application may also provide a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the steps of the above-mentioned endoscope field of view scene reconstruction method.
[0144] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0145] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0146] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0147] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for reconstructing an endoscope visual field scene, characterized in that: include: Acquire a number of continuous frame RGB organ images collected through an endoscope; Using a monocular depth prediction network to estimate the RGB organ images of several consecutive frames respectively to obtain several local depth maps; Using the local depth maps of two adjacent frames, an inter-frame pose estimation module is used to estimate and optimize the pose matrix of the camera transformation of the previous and next frames; Obtaining XYZ parameters of the Gaussian model of the previous frame by calculation according to the local depth map of the previous frame having the intrinsic parameter matrix; The coordinate system of the Gaussian model of the previous frame is transformed from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames; The RGB organ image of the next frame and the predicted corresponding local depth map are processed by a dynamic Gaussian growth module, and the local Gaussian set appearing in the new area under the viewpoint of the next frame is extended from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image; The local Gaussian models corresponding to each frame of the RGB organ image are fused to obtain a global Gaussian model.
2. The method for reconstructing an endoscope field of view according to claim 1, characterized in that: Using the local depth maps of two adjacent frames and adopting the inter-frame pose estimation module to estimate and optimize the camera transformation pose matrix of the previous and next frames includes: Using the dense key point matching of the feature matching network and the relative depth calculated by the dense depth estimation network, a 2D-3D displacement field is established using the 3D coordinates of the organ map of the previous frame and the 2D corresponding position of the organ map of the next frame; The camera transformation pose matrix of the previous and next frames is obtained by setting the optimization objective to minimize the Gaussian rendering variational gradient of the displacement field.
3. The method for reconstructing an endoscope field of view according to claim 2, characterized in that: The corresponding point cloud is obtained using the intrinsic matrix and distortion parameter transformation in the local camera coordinate system; The 3D displacement vector field of the organ image of the previous frame is obtained by constructing the local depth map and the point cloud, and Kp N and Kp N+1 Construct a 2D displacement vector field; Through variational optimization of the displacement vector field and the energy functional of the determined vector field under different transformation matrix positions, the initial estimation of the transformation matrix position is achieved based on the 2D vector after optimization of the 3D vector field. Based on the initial value of the initial estimate of the transformation matrix pose, the energy function target of the variational optimization is determined to be the minimum image similarity difference between the Gaussian rendering image and the original RGB organ image at the current viewing angle as a generalization function, and the optimized front and back frame camera transformation pose matrix is calculated.
4. The method for reconstructing an endoscope field of view according to claim 1, characterized in that: Processing the RGB organ image of the next frame and the predicted corresponding local depth map by using a dynamic Gaussian growth module includes: Based on the local depth map of the next frame, a new Gaussian ellipsoid is newly obtained in a blank area under a new viewpoint of the next frame using the local depth information and the corresponding RGB spherical harmonic function information as initial Gaussian attributes; Optimizing the properties of the newly created Gaussian ellipsoid based on a differential gradient field optimization algorithm, and using the loss function of the transferred depth and RGB as an optimization target; Based on the growing Gaussian set under the local perspective of the next frame, the grown Gaussian is aligned with the global Gaussian, thereby realizing the growth of the Gaussian model under the viewpoint of the next new frame.
5. The method for reconstructing an endoscope field of view according to claim 1, characterized in that: During the training process, the strategies of forward reconstruction and backward optimization are used to ensure the coherence of sequence frames and achieve global alignment of the Gaussian model.
6. The method for reconstructing an endoscope field of view according to claim 5, characterized in that: The forward reconstruction and reverse optimization strategies include: For the initial frame, the viewpoint is initialized to the local camera coordinate system of the first frame, and the corresponding point cloud is calculated based on the local depth map; the XYZ attributes of the Gaussian model are initialized with the point cloud, and the spherical harmonic function SHS parameters are initialized with the RGB values of the frame; In the forward processing of sequential video frames, the inter-frame pose estimation module is used to estimate the inter-frame pose transformation matrix, and the dynamic Gaussian growth module is used to realize the gradual growth of the Gaussian model; In the backward process of sequential video frames, cyclic training is implemented by reversing the order of video frames so that the estimated Gaussian parameters are optimized during the forward reconstruction and the gradients are back-propagated.
7. An endoscope visual field scene reconstruction device, characterized in that: The device is used to perform the endoscope field of view scene reconstruction method according to any one of claims 1 to 6, comprising: An image acquisition unit, used to acquire a plurality of continuous frame RGB organ images collected through the endoscope; A depth map estimation unit, used to estimate the RGB organ images of several consecutive frames respectively using a monocular depth prediction network to obtain several local depth maps; A pose transformation matrix determination unit is used to estimate and optimize the camera transformation pose matrix of the previous and next frames by using the local depth maps of two adjacent frames and adopting an inter-frame pose estimation module; An initial Gaussian model parameter determination unit, configured to calculate and obtain XYZ parameters of a Gaussian model of a previous frame according to the local depth map of the previous frame having an intrinsic parameter matrix; A Gaussian model coordinate system transformation unit, used to transform the coordinate system of the Gaussian model of the previous frame from the local camera coordinate system of the previous frame to the local camera coordinate system of the next frame by using the camera transformation matrix of the previous and subsequent frames; A Gaussian model sequential dynamic growth unit is used to process the RGB organ image of the next frame and the predicted corresponding local depth map using a dynamic Gaussian growth module, and expand the local Gaussian set appearing in the new area under the viewpoint of the next frame from the Gaussian model of the previous frame to the Gaussian model of the next frame, so as to realize the sequential dynamic growth of the Gaussian model to obtain the local Gaussian model corresponding to each frame of the RGB organ image; The global Gaussian model acquisition unit is used to fuse the local Gaussian models corresponding to each frame of the RGB organ image to obtain a global Gaussian model.
8. An endoscope visual field scene reconstruction device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the endoscopic field of view scene reconstruction method described in any one of claims 1-6 according to the instructions in the program code.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the endoscope field of view scene reconstruction method described in any one of claims 1-6.
Citation Information
Cited By
Dynamic scene rapid reconstruction method and device based on deformation prior and meta learning, equipment and storage medium
CN121639928A