Endoscopic surgery scene real-time reconstruction method based on 3D Gaussian

By introducing depth estimation model and basis function modeling of learnable parameters in 3D Gaussian reconstruction technology, the problem of initial point cloud sparseness and dynamic change modeling is solved, and real-time high-quality reconstruction of endoscopic surgical scenarios is achieved.

CN120070755APending Publication Date: 2025-05-30NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510144591.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the existing 3D Gaussian reconstruction technology deals with endoscopic surgical scenarios, the initial point cloud sparseness leads to difficulty in optimization process and cannot effectively model dynamically changing deformable tissue, affecting the application effect.

Method used

A depth estimation model is introduced for global Gaussian initialization, and a strategy of selecting some Gaussian points as control points is proposed. The motion of Gaussian points is modeled through the basis function of learning parameters, replacing the traditional MLP network, and achieving rapid Gaussian deformation.

Benefits of technology

It significantly accelerated the speed of Gaussian reconstruction, achieved real-time reconstruction effect of endoscopic surgical scenarios, and improved the reconstruction quality, supporting monocular and binocular inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070755A_ABST
    Figure CN120070755A_ABST
Patent Text Reader

Abstract

The invention discloses an endoscopic surgery scene real-time reconstruction method based on 3D Gaussian, and particularly relates to the technical field of three-dimensional reconstruction, and the method comprises the following steps: carrying out the preprocessing of an endoscopic surgery video; initializing the overall 3D Gaussian representation, and selecting sparse control points; setting a Gaussian basis function of learnable parameters to model the motion; guiding global Gaussian deformation by using a deformation result of a sparse control point, and standardizing the deformation result through local optimization and global optimization; and optimizing a primary function parameter and a Gaussian point parameter according to the difference between the rendered image and the reference image, and managing a Gaussian control point sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a method for real-time reconstruction of endoscopic surgical scenes based on 3D Gaussian. Background Art

[0002] Reconstructing surgical scenes from endoscopic videos is crucial for robot-assisted minimally invasive surgery (RAMIS). By restoring the 3D models of the observed tissues, these technologies can be used to simulate the surgical environment, assist doctors in preoperative planning, and provide support for AR / VR medical staff training. In addition, reconstruction technologies that can achieve real-time rendering can further broaden their applications during the surgical process, provide surgeons with a complete scene view, assist them in more precisely navigating and manipulating instruments during surgery, and may lay the foundation for the automation of future robotic surgeries.

[0003] Compared with existing three-dimensional reconstruction methods (such as NeRF), the 3D Gaussian reconstruction technology has significant advantages in reconstruction speed and quality. However, this technology relies on Structure from Motion (SfM) to initialize the positions of the Gaussians, and the sparse initial point cloud poses challenges to the subsequent optimization process. In addition, the original design of 3D Gaussian Splatting has limitations in dealing with deformable tissues commonly found during surgery and cannot effectively model these dynamically changing structures, further affecting its application effect in complex surgical scenes. Summary of the Invention

[0004] Therefore, the present invention makes improvements on the basis of the original 3D Gaussian technology. Firstly, a depth estimation model is introduced. By predicting depth information and performing reprojection and combination processing on the input pixels based on the depth map, it can support monocular and binocular inputs, thereby achieving faster global Gaussian initialization. In addition, a strategy of selecting some Gaussian points as control points for the global Gaussian is proposed, and at the same time, the motion of key Gaussian points is modeled through a basis function with learnable parameters, replacing the traditional MLP network, thereby realizing the rapid deformation of the Gaussian. Through the combination of these two methods, not only can the speed of Gaussian reconstruction be greatly accelerated to achieve the effect of real-time reconstruction, but also the real-time reconstruction quality of the endoscopic surgical scene is effectively improved.

[0005] The present invention provides a method for real-time reconstruction of endoscopic surgical scenes based on 3D Gaussian to solve the problems raised in the background art, and the technical solutions are as follows: The method for real-time reconstruction of endoscopic surgical scenes based on 3D Gaussian includes the following steps:

[0006] S1: Preprocess the endoscopic surgical video; S2: Initialize the overall 3D Gaussian representation and select sparse control points;

[0007] S3: Model the motion using Gaussian basis functions with learnable parameters;

[0008] S4: Use the deformation results of sparse control points to guide the global Gaussian deformation, and regularize the deformation results through local and global optimizations;

[0009] S5: Optimize the basis function parameters and Gaussian point parameters according to the difference between the rendered image and the reference image, and manage the Gaussian control point sequence.

[0010] Preferably, in step S1, the preprocessing of the endoscopic surgery video includes the following steps:

[0011] S11: Split each frame in the scene video and save it as an RGB image;

[0012] S12: For the input monocular image sequence, use the monocular depth estimation model to estimate the depth of the RGB image to obtain the corresponding depth image D i ; For the input binocular image sequence, use the stereo depth estimation model to estimate the depth of the RGB image to obtain the corresponding left view depth image D i ;

[0013] S13: Based on the predicted depth image D i , according to the formula P i = K -1 T i D i (I i ⊙M i ) Reproject the image pixels into the world coordinate system to obtain the partial point cloud P i , where ⊙ represents the element-wise product, and the binary mask M i is used to filter out the surgical tool pixels, and K and T i represent the known camera intrinsic and extrinsic parameters respectively;

[0014] S14: Combine the partial point clouds P i at all time steps T to obtain the overall point cloud initialization P = {P 1 , P 2 ,..., P T}.

[0015] Preferably, in step S2, the initialization of the overall 3D Gaussian representation and the selection of sparse control points include the following steps:

[0016] S21: Generate a colored and anisotropic 3D Gaussian representation of the scene according to the point cloud;

[0017] S22: Divide the scene according to the 3D scene and divide the scene into a number of voxels of the same size;

[0018] S23: According to the results of 3D Gaussian initialization, select the patch shapes with dense Gaussian points for further division;

[0019] S24: In each voxel obtained by division, select sparse points as control points to form a set of Gaussian control points \(P = \{(p i \in\mathbb{R} 3 , o i \in\mathbb{R}^+)\}, i\in\{1, 2, \cdots, N p \}, where \(p i \) represents the learnable coordinates of the control point \(i\) in the canonical space, and \(o i \) is the learnable radius parameter of the radial basis function (RBF) kernel, which controls how the influence of the control point on the Gaussian distribution decreases as the distance increases. \(N p \) is the total number of control points, which is much less than the number of Gaussian control points.

[0020] Preferably, in step S2, generating a colored and anisotropic 3D Gaussian representation of the scene based on the point cloud \(P i \) includes the following steps:

[0021] S211: For each Gaussian point \(G\), there is a 3D center point position \(\mu\) and a 3D covariance matrix \(\Sigma\), where the 3D covariance matrix \(\Sigma\) is decomposed as \(\Sigma = RSS T R T \) for optimization. \(R\) is a rotation matrix represented by the quaternion \(q\in SO(3)\), and \(S\) is a scaling matrix represented by the 3D vector \(s\). Each Gaussian function has an opacity value \(\sigma\) for adjusting its influence in rendering and is associated with the spherical harmonic coefficients \(SH\) for view-dependent appearance. Parameterize the scene as a set of Gaussian distributions \(G = \{G i : \mu j , q j , s j , \sigma j , sh j \}\) according to the initialized point cloud \(P j \);

[0022] S212: Project the initialized Gaussian distribution \(G\) onto the 2D image plane and aggregate it using fast alpha blending. The two-dimensional covariance matrix \(\Sigma' = JW\Sigma W T J T \), where \(J\) represents the Jacobian matrix and \(W\) represents the view transformation matrix. By calculating the above formula, the three-dimensional Gaussian point cloud can be projected onto the two-dimensional plane to obtain the two-dimensional coordinate center of the Gaussian point \(\mu' = JW\mu\). The color \(C(u)\) of the pixel \(u\) is calculated using neural point-based alpha blending, expressed as \(C(u)=\sum i∈N T i \alpha i SH(shi , v i ), where N is the set of Gaussian points related to pixel u, and T i represents the cumulative transmittance rate, and represents the weight ratio of the color contribution of the i-th Gaussian neural point. α i represents the transparency of the i-th neural point (which can be regarded as the weight on a specific pixel projection). SH(sh i , v i ) represents the color information using the spherical harmonic function SH, where the parameter sh i is the spherical harmonic function, and v i is the viewing direction. And T i represents the cumulative transmittance rate, that is, the probability that the light from the first neural point to the i - 1-th neural point is not completely absorbed, and is calculated by the formula . And the transparency α i is obtained by calculating the corresponding projected Gaussian G at pixel u i , where σ i is the weight scalar of the i-th Gaussian point, p is the position of pixel u, and μ′ i is the mean of the i-th Gaussian point in the pixel projection;

[0023] S213: According to the RGB image, optimize the Gaussian parameter G = {G j : μ j , q j , s j , σ j}, and adaptively adjust the Gaussian density.

[0024] Preferably, in step S3, the setting of the learnable parameter Gaussian basis function to model the motion includes the following steps:

[0025] S31: Set the Gaussian basis function with learnable parameters where t represents time, θ represents the position of the Gaussian, and σ represents the variance of the Gaussian, that is, the coverage range;

[0026] S32: For each Gaussian point, during the motion, its center position, rotation, and scale are all affected by deformation. To model these complex dynamic changes, a set of additional learnable parameters Θ μ , Θ r , Θ s ; are introduced for each Gaussian point;

[0027] S33: Fix the overall Gaussian distribution G, and only train the colors of the sparse control points P and the corresponding learnable parameters.

[0028] Preferably, in step S4, the global Gaussian deformation is guided by the deformation result of sparse control points, and the deformation result is normalized through local optimization and global optimization, including the following steps:

[0029] S41: For each Gaussian distribution G j , use the KNN nearest neighbor algorithm to find K = 4 nearest sparse control points in the canonical space, denoted as {p k |k ∈ N j};

[0030] S42: Adopt a radial basis function RBF based on a Gaussian kernel to calculate the interpolation weight w jk , and the formula is where d jk represents the distance between the center of the Gaussian distribution G j and the neighboring control point p k , o k is the learning radius parameter of the control point p k , which is used to control the attenuation rate of the Gaussian kernel, while refers to the value of the Gaussian kernel function, representing the distance weight between the Gaussian distribution G j and the control point p k . The larger the value, the closer the two are. During the training process, the interpolation weight w jk is learnable to adapt to complex motions;

[0031] S43: Calculate the motion field of the Gaussian through linear blend skinning (LBS), and interpolate the center μ j of the Gaussian distribution. The update formula for the Gaussian center is: where, r k t represents the rotation matrix of the control point p k at time t, and T k t represents the translation amount of the control point at time t. By interpolating the transformation of the control points, the displacement of the Gaussian at time step t can be calculated;

[0032] S44: Calculate the motion field of the Gaussian through linear blend skinning (LBS), and interpolate the quaternion rotation q j , and the formula is: where r k t represents the rotation matrix of the control point p k at time t, and represents the multiplication of quaternions, which is used to synthesize the rotation transformation.

[0033] Preferably, in step S5, optimizing the basis function parameters and Gaussian point parameters according to the differences between the rendered image and the reference image, and managing the Gaussian control point sequence includes the following steps:

[0034] S71: Update the parameters of the global Gaussian using sparse control points, using the formula C(u) = ∑ i∈N T i α i SH(sh i ,v i ),, and the formula Render the dynamic scene at time step t;

[0035] S72: For each control point p j , calculate its trajectory at different time steps This trajectory contains the position information of the control point at N t = 8 randomly sampled time steps, and determine all neighboring control points within the radius through spherical query Define a local area;

[0036] S72: Introduce a local rigid ARPR loss, randomly sample two time steps t 1 and t 2 , obtain the transformed position, According to the local rigid motion assumption, by minimizing the objective function Obtain the rotation matrix where N c i is the set of all neighboring control points;

[0037] S73: Calculate the ARPR loss according to the formula ;

[0038] S74: Calculate the rendering loss L from the difference between the rendered images at different time steps and the real images render ;

[0039] S75: According to the rendering result, adaptively adjust the density of the sparse control points to achieve state management;

[0040] S76: Calculate the mean absolute error between the rendered image and the real image, and the mean absolute error between the rendered depth image and the real depth image to optimize the parameters of the Gaussian point cloud and the basis function.

[0041] Preferably, in step S75, according to the rendering result, adaptively adjusting the density of the sparse control points to achieve state management includes the following steps:

[0042] S751: According to the rendering result, calculate the total influence of each control point p i ; If the total influence W i is close to 0, prune this sparse control point;

[0043] S752: For regions with poor reconstruction quality and large cumulative deformation amounts, clone the Gaussian point p of the current region k to increase the density of the Gaussian point cloud and enhance the accuracy of modeling;

[0044] S753: After adding p k ’ add it to the sparse control point sequence to improve the reconstruction effect.

[0045] The present invention has the following advantages:

[0046] (1) The present invention can propose processing methods for monocular and binocular video inputs, which are applicable to different types of endoscopic surgery scenarios;

[0047] (2) The present invention realizes real-time reconstruction of dynamic scenes, improves the rendering quality, and ensures accurate reproduction of the surgical scene;

[0048] (3) By introducing 3D Gaussian and sparse control points, the present invention significantly reduces the computational amount, reduces the high-performance dependence on hardware, and thus makes the real-time reconstruction technology of endoscopic surgery more popular. Description of the Drawings

[0049] Figure 1 is the text flow chart of the real-time endoscopic surgery scene reconstruction technology based on 3D Gaussian of the present invention.

[0050] Figure 2 is an example of the 3D point cloud obtained after input preprocessing and initialization.

[0051] Figure 3 is an example of the distribution of the sparse control point sequence selected on the overall 3D Gaussian.

[0052] Figure 4 is the operation flow chart of the real-time endoscopic surgery scene reconstruction technology based on 3D Gaussian of the present invention. Detailed Embodiments

[0053] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] Combined with Figure 1, this embodiment mentions a real-time reconstruction method for endoscopic surgery scenarios based on 3D Gaussian, including the following steps:

[0055] First, process the input data, perform preprocessing, and obtain a preliminary 3D point cloud.

[0056] First, segment the input video data to obtain a series of RGB images. Then, for the input monocular image sequence, use the monocular depth estimation model to estimate the depth of the RGB image to obtain the corresponding depth image; for the input binocular image sequence, use the stereo depth estimation model to estimate the depth of the RGB image to obtain the corresponding left-view depth image. After obtaining the depth image, project the original image pixels into the world coordinate system to obtain a partial point cloud. Finally, integrate the partial point clouds at all time steps to obtain the overall original point cloud.

[0057] Figure 2 is an example of the initial point cloud obtained in this embodiment.

[0058] Second, use the point cloud to initialize the overall 3D Gaussian representation and select sparse control points.

[0059] First, initialize a colored, anisotropic 3D Gaussian according to the point cloud in the previous step. The specific steps for generating the 3D Gaussian are as follows:

[0060] (1) For each Gaussian point G, there is a 3D center point position μ and a 3D covariance matrix ∑, where the 3D covariance matrix ∑ is decomposed into ∑ = RSS T R T is optimized, R is a rotation matrix represented by the quaternion q ∈ SO(3), S is a scaling matrix represented by the 3D vector s, each Gaussian function has an opacity value σ, which is used to adjust its influence in rendering and is associated with the spherical harmonic (SH) coefficient SH for view-dependent appearance. Parameterize the scene as a set of Gaussian distributions G = {G j : μ j , q j , s j , σ j , sh j};

[0061] (2) Project the initialized Gaussian distribution G onto the 2D image plane and aggregate it using fast α-blending. The two-dimensional covariance matrix Σ' = JWΣW T J T , the two-dimensional coordinate center μ′ = JWμ, and the color C(u) of the pixel u is rendered using neural point-based α-blending. where, α iIt is calculated by computing the corresponding projected Gaussian G at pixel u i and is computed as follows

[0062] (3) According to the RGB image, optimize the Gaussian parameters G = {G j : μ j , q j , s j , σ j}, adaptively adjust the Gaussian density, which involves duplicating the Gaussian at positions with insufficient reconstruction and pruning the Gaussian with almost zero transparency at positions with over-reconstruction.

[0063] Then, divide the Gaussian scene into several patches according to the complexity of the scene. Uniformly sample Gaussian points within each patch as sparse control points. Finally, a sequence of sparse control points can be obtained.

[0064] In this step, accept manual marking, that is, medical staff can actively mark the parts that need to be focused on before the operation, and denser control points will be selected for these parts. If there is no manual marking, the patches will be automatically divided according to the Gaussian density.

[0065] The number of selected sparse control points is much less than the number of the original 3D Gaussians, as specifically shown in Figure 3 the following figure.

[0066] Third, set the Gaussian basis function with learnable parameters to model the motion.

[0067] Fix the overall Gaussian distribution and only train and optimize the sparse control point sequence extracted in the previous steps and the basis function parameter group representing the motion transformation of the sparse control points. The specific optimization steps are explained as follows:

[0068] (1) Set the Gaussian basis function with learnable parameters

[0069] (2) For each sparse control point k, introduce a set of additional learnable parameters Θ μ , Θ r , Θ s ;

[0070] (3) By jointly training the parameters of the sparse control points and the parameters of the basis function, make the basis function correctly reflect the dynamic changes of the Gaussian points in the scene;

[0071] Through the joint optimization of the sparse control point sequence and the basis function parameters, a preliminary trained dynamic representation model can be obtained, which can complete the prediction of Gaussian deformation in the subsequent reconstruction steps and reduce the computational amount.

[0072] Fourth, use the deformation results of sparse control points to guide the global Gaussian deformation, and standardize the deformation results through local optimization and global optimization.

[0073] First, for each Gaussian distribution G j , use the KNN (K-Nearest Neighbor) algorithm to find K = 4 nearest sparse control points in the canonical space, denoted as {p k |k ∈ N j}}. The effect of the KNN algorithm is as Figure 4 shown. Then, use the radial basis function (RBF) based on the Gaussian kernel to calculate the interpolation weights w jk , and the formula is where d jk represents the distance between the center of the Gaussian distribution G j and the neighboring control point p k , and o k is the learning radius parameter of the control point p k . During the training process, these interpolation weights are adjustable to adapt to complex motions, and after optimization, the video frames can be reconstructed more precisely. After that, use linear blend skinning (LBS) to calculate the motion field of the Gaussian. The specific calculation steps are as follows:

[0074] (1) Interpolate the center μ j of the Gaussian distribution through linear blend skinning (LBS). The update formula for the Gaussian center is: Calculate the displacement of the Gaussian at time step t through the transformation of the interpolated control points;

[0075] (2) Interpolate the quaternion rotation q j through linear blend skinning (LBS), and the formula is: where represents the quaternion multiplication used to synthesize the rotation transformation;

[0076] After synthesizing the displacement transformation and the rotation transformation, the motion deformation result of the 3D Gaussian can be obtained.

[0077] Fifth, optimize the basis function parameters and Gaussian point parameters according to the differences between the rendered image and the reference image, and manage the Gaussian control point sequence.

[0078] First, in the previous step, use the sparse control points to complete the parameter update of the global Gaussian, and perform the rendering of the dynamic scene at time step t according to the formula and the formula . Then, for each control point p j , calculate its trajectory at different time steps This trajectory contains the control points at randomly selected Nt = position information of 8 time steps, and all neighboring control points within the radius are determined through spherical query Define a local area. Then, calculate the local rigidity (ARPR) loss and the rendering loss L respectively render Backpropagation realizes global joint optimization.

[0079] Finally, according to the rendering result, adaptively adjust the density of sparse control points to achieve state management. Specifically, the steps of state management include:

[0080] (1) According to the rendering result, calculate the total influence of each control point p i If the total influence W is close to 0, prune this sparse control point; i If the total influence W

[0081] (2) For areas with poor reconstruction quality and large cumulative deformation, clone the Gaussian points p of the current area k to increase the density of the Gaussian point cloud and enhance the accuracy of modeling;

[0082] (3) After adding p k ’, add it to the sparse control point sequence to improve the reconstruction effect.

[0083] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A real-time reconstruction method of endoscopic surgery scene based on 3D Gaussian, comprising the following steps: S1: Preprocessing of endoscopic surgery video; characterized in that: S2: Initialize the overall 3D Gaussian representation and select sparse control points; S3: Setting Gaussian basis functions with learnable parameters to model motion; S4: Use the deformation results of sparse control points to guide the global Gaussian deformation, and standardize the deformation results through local optimization and global optimization; S5: Optimize basis function parameters and Gaussian point parameters according to the difference between the rendered image and the reference image, and manage the Gaussian control point sequence.

2. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 1, characterized in that: In step S1, the endoscopic surgery video is preprocessed, including the following steps: S11: Segment each frame in the scene video and save it as an RGB image, denoted as I i ; S12: For the input monocular image sequence, use the monocular depth estimation model to perform depth estimation on the RGB image to obtain the corresponding depth image D i ; For the input binocular image sequence, the stereo depth estimation model is used to estimate the depth of the RGB image to obtain the corresponding left view depth image D i ; S13: Based on the predicted depth image D i According to the formula P i =K -1 T i D i (I i ⊙M i ) Reproject the image pixels into the world coordinate system to obtain a partial point cloud P i , where ⊙ represents the element-wise product and the binary mask M i Used to filter out surgical tool pixels, K and T i Represent the known intrinsic and extrinsic parameters of the camera respectively; S14: All the partial point clouds P of the time step T i Combined together, the overall point cloud initialization P = {P1, P2, ..., P T }.

3. The real-time reconstruction method of endoscopic surgery scene based on 3D Gaussian according to claim 2 is characterized in that: In step S2, the initialization of the overall 3D Gaussian representation and the selection of sparse control points include the following steps: S21: Generate a colorful, anisotropic 3D Gaussian representation of the scene based on the point cloud; S22: dividing the 3D scene into a number of voxels of the same size; S23: According to the result of 3D Gaussian initialization, the patch shape with dense Gaussian points is screened out for further division; S24: In each voxel obtained by division, sparse points are selected as control points to form a set of Gaussian control points P = {(p i ∈R 3 ,o i ∈R+)},i∈{1,2,···,N p }, where p i represents the learnable coordinates of control point i in the canonical space, o i is the learnable radius parameter of the radial basis function RBF kernel, N p is the total number of control points.

4. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 3, characterized in that: In step S2, according to the point cloud P i , generating a colorful, anisotropic 3D Gaussian representation of a scene includes the following steps: S211: For each Gaussian point G, there is a 3D center point position μ and a 3D covariance matrix Σ. The Gaussian function is expressed as Among them, the 3D covariance matrix Σ is decomposed into Σ = RSS T R T to be optimized, R is the rotation matrix represented by the quaternion q∈SO(3), S is the scaling matrix represented by the 3D vector s, and each Gaussian function has an opacity value σ used to adjust its influence in the rendering and is associated with the spherical harmonic coefficients SH for view-dependent appearance. According to the initialized point cloud P i The scene is parameterized as a set of Gaussian distributions G = {G j :μ j ,q j ,s j , σ j ,sh j }, for each Gaussian point G j , respectively with the center position μ j , the rotation matrix q j , scaling matrix s j , representing the coefficient σ of the transparency of the Gaussian point j and the spherical harmonic function sh representing the color j ; S212: Project the initialized Gaussian distribution G onto the 2D image plane and aggregate it using fast alpha blending. The two-dimensional covariance matrix Σ′=JWΣW T J T , where J represents the Jacobi matrix and W represents the view transformation matrix. The three-dimensional Gaussian point cloud is projected onto a two-dimensional plane to obtain the two-dimensional coordinate center of the Gaussian point μ′=JWμ; the color C(u) of pixel u is calculated using α-blending based on the neural point, expressed as C(u)=∑ i∈N T i α i SH i ,v i ), where N is the set of Gaussian points associated with pixel u, T i represents the transmission accumulation rate, which represents the weight ratio of the i-th Gaussian neural point's contribution to color, α i Indicates the transparency of the i-th neural point, SH(sh i ,v i ) uses spherical harmonic function SH to represent color information, where the parameter sh i is a spherical harmonic function, v i is the viewing direction, and T i Represents the cumulative transmittance, that is, the probability that the light from the first neural point to the i-1th neural point is not completely absorbed, through the formula Calculated, and the transparency α i is calculated by computing the corresponding projected Gaussian G at pixel u i Come get, where σ i is the weight scalar of the i-th Gaussian point, p is the position of pixel u, μ′ i is the mean value of the i-th Gaussian point in the pixel projection; S213: According to the RGB image, optimize the Gaussian parameter G={G j :μ j ,q j ,s j , σ j }, adaptively adjust the Gaussian density.

5. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 4, characterized in that: In step S3, the Gaussian basis function with learnable parameters is set to model the motion, including the following steps: S31: Setting Gaussian basis functions for learnable parameters Where t represents time, θ represents the position of Gaussian, and σ represents the variance of Gaussian, that is, the coverage range; S32: For each Gaussian point, its center position, rotation and scale will be affected by deformation during the movement. In order to model these complex dynamic changes, an additional set of learnable parameters is introduced for each Gaussian point, {Θ μ ,Θ r ,Θ s }; S33: Fix the overall Gaussian distribution G and only train the color of the sparse control point P and the corresponding learnable parameters.

6. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 4, characterized in that: In step S4, the deformation result of the sparse control points is used to guide the global Gaussian deformation, and the deformation result is standardized by local optimization and global optimization, including the following steps: S41: For each Gaussian distribution G j , use the KNN nearest neighbor algorithm to find the K nearest sparse control points in the standard space, denoted as {p k |k∈N j }; S42: Use the radial basis function RBF based on the Gaussian kernel to calculate the interpolation weight w jk , the formula is, in d jk Represents Gaussian distribution G j The center and neighborhood control points p k The distance between k is the control point p k The learning radius parameter is used to control the decay rate of the Gaussian kernel, and Refers to the value of the Gaussian kernel function, indicating the Gaussian distribution G j and control point p k The distance weight between them, the larger the value, the closer the two are; during the training process, the interpolation weight w jk are learnable to adapt to complex movements; S43: Calculate the Gaussian motion field through linear mixed skinning LBS, and the center μ of the Gaussian distribution j Interpolation, the update formula of Gaussian center is: Among them, r k t Indicates that at time t, the control point p k The rotation matrix, T k t Represents the translation of the control point at time t; by interpolating the transformation of the control point, the displacement of Gaussian at time step t is calculated; S44: Calculate the Gaussian motion field through linear blend skinning LBS, and rotate the quaternion q j Interpolation is performed, the formula is: Among them, r k t Indicates that at time t, the control point p k The rotation matrix of Represents the multiplication of quaternions, used to synthesize rotation transformations.

7. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 3, characterized in that: In step S5, the optimizing of basis function parameters and Gaussian point parameters according to the difference between the rendered image and the reference image and managing the Gaussian control point sequence comprises the following steps: S71: Use sparse control points to update the parameters of the global Gaussian, using the formula C(u)=∑ i∈N T i α i SH i ,v i ),, and the formula Render the dynamic scene at time step t; S72: For each control point p j , calculate its trajectory p at different time steps trag t , the trajectory contains control points in N randomly selected t The position information of the time step is obtained, and all neighboring control points N within the radius are determined by spherical query c i Delimitation of local areas; S72: Introduce local rigid ARPR loss, randomly sample two time steps t1 and t2, and obtain the transformed position, According to the local rigid motion assumption, by minimizing the objective function Get the rotation matrix Where N c i is the set of all neighboring control points; S73: According to the formula Calculate ARPR loss; S74: The rendering loss L is calculated by the difference between the rendered image and the real image at different time steps. render ; S75: Adaptively adjust the density of sparse control points according to the rendering results to achieve state management; S76: Calculate the mean absolute error between the rendered image and the real image, and the mean absolute error between the rendered depth image and the real depth image to optimize the parameters of the Gaussian point cloud and the parameters of the basis function.

8. The method for real-time reconstruction of endoscopic surgery scenes based on 3D Gaussian according to claim 7, characterized in that: In step S75, the density of the sparse control points is adaptively adjusted according to the rendering result, and the state management is implemented, including the following steps: S751: Calculate each control point p according to the rendering result i Total influence If the total influence W i Close to 0, prune the sparse control point; S752: For areas with poor reconstruction quality and large cumulative deformation, clone the Gaussian points p in the current area. k Increase the density of Gaussian point cloud; S753: cloned p k ', add it to the sparse control point sequence to improve the reconstruction effect.

Citation Information

Patent Citations

  • Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function

    CN118505541A

  • Modeling point cloud data using hierarchies of gaussian mixture models

    US20170249401A1

Cited By

  • Overall dual-constraint dynamic registration navigation system for intraoperative scene

    CN120451463A

  • Intraoperative scene overall dual-constraint dynamic registration navigation system

    CN120451463B

  • Endoscope image dynamic reconstruction method based on point cloud compensation and double-domain deformation

    CN120510302A

  • Dynamic scene reconstruction method using Kalman filter to guide Gaussian rendering

    CN121074215A