Improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision
By combining depth and curvature supervision techniques, the 3DGS scene reconstruction and rendering method is optimized, and the problems of Gaussians occlusion and high-gradient area processing in the prior art are solved, achieving higher quality and efficiency rendering effects.
Patent Information
- Application Number
- CN202510145733.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing 3DGS technology is prone to Gaussians occlusion problems of wrong opacity or scale during the rendering process, resulting in artifacts in the rendering results; at the same time, the ADC algorithm generates a large number of small and dense Gaussians when processing high-gradient areas, ignoring the contribution of small gradient areas to rendering.
Using an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision, the real depth of the sparse point cloud generated by SFM and the predicted depth generated by fast microrasterization are calculated as a loss function to optimize the geometric constraints of the Gaussians position; at the same time, the scale is adjusted according to the distance from the Gaussians center to the camera, and the curvature is combined as a constraint condition to optimize the Gaussians distribution.
It effectively avoids wrong occlusion relationships and rendering artifacts, improves the realism of the scene; ensures the correct position and scale of Gaussians in the space, improves the rendering quality; by reducing redundant small Gaussians, the calculation complexity is reduced and performance bottlenecks are avoided.
Smart Images

Figure CN120014172A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of three-dimensional image reconstruction, and in particular relates to an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision. Background Art
[0002] Reconstructing complex surfaces and geometries and synthesizing new perspectives have always attracted much attention in computer graphics and computer vision, and have very important applications in intelligent driving, virtual reality, and cultural relics protection. With the introduction of Neural Radiance Fields (NeRF), NeRF uses multilayer perceptron (MLP) parameters to achieve continuous implicit encoding of scenes, and has achieved impressive results in surface geometry reconstruction and new perspective synthesis. Although NeRF has opened a new direction in geometry reconstruction and new perspective synthesis, such as A Multiscale Representation for Anti-Aliasing Neural Radiance Fields (Mip-NeRF), it has a good rendering quality, but it has a huge time cost. Instant neural graphics primitives with a multiresolution hash encoding (Instant-NGP) proposes a multiresolution hash encoding to replace MLP, and reduces some quality in exchange for faster rendering speed, but it is not well adapted to large-scale scenes. NeRF-based methods cannot solve the problem of fast and effective real-time rendering.
[0003] Recently, 3D Gaussian Splatting (3DGS) takes a learnable Gaussian function as its starting point and the design of an explicit radiation field, which is rendered to screen space through splash-based rasterization. The emergence of 3DGS as a new method, based on NeRF, achieves state-of-the-art visual quality while maintaining competitive training time, bringing existing scene representation and rendering to a new level. The innovation of 3DGS lies in its unique fusion of the advantages of differentiable pipelines and point-based rendering techniques. By using a learnable 3D Gaussian function to represent the scene, it retains the strong fitting ability of the continuum radiation field, which is critical for high-quality image synthesis, while avoiding the computational overhead associated with NeRF-based methods. During the optimization process, 3DGS uses a single adaptive density control (ADC) algorithm to manage points, which uses the average gradient to determine the area that needs to be processed. Usually during the optimization process, the complexity of the scene geometry varies greatly, and a single ADC algorithm cannot well identify and process all points, which may lead to incorrect rendering. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision provided by the present invention solves the following problems: (1) During rendering along the viewing direction, Gaussians with incorrect opacity or scale will occlude the Gaussians behind them, eventually causing artifacts in the rendering results.
[0005] (2) The ADC algorithm of 3DGS uses the average gradient threshold based on the image gradient to adjust the splitting and cloning of Gaussians, which will result in a large number of small and dense Gaussians in high-gradient areas and ignore the Gaussians that contribute to rendering in areas with smaller gradients.
[0006] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision, comprising the following steps: S1. Obtain a sparse corresponding set of image features based on the real-time image collected by the camera, use the SfM technology to generate a sparse point cloud representation of the scene, and estimate the camera's motion posture to obtain the camera extrinsic parameters; S2. Using sparse point clouds, the scene is represented as a set of Gaussians, where the center position is the spatial position of the point cloud, the constant of the spherical harmonic function is the color of the corresponding pixel of the point cloud, the high-order coefficient is set to zero, the three-axis scale is set to the average distance of the three point clouds adjacent to the center position, and the axis is represented by a quaternion, which is initialized as , opacity is set to 0.1 and the number of iterations is initialized; S3, combined with the camera external parameters, use the perspective transformation matrix to project the Gaussians distribution onto the two-dimensional plane under the camera perspective, obtain the 2D Gaussians projected onto the image plane, and obtain the predicted depth from it; S4, sorting the 2D Gaussians by depth, determining the color of each pixel to be rendered according to the specific point through which the viewing direction passes, and generating a rendered image; S5, determine whether the number of iterations reaches the preset iteration threshold, if so, enter S6, if not, calculate the loss function according to the real depth, predicted depth, rendered image and RGB real image, update the optimized Gaussians distribution according to the calculation result, increase the number of iterations by 1 and return to S3; S6. Output the final rendered image.
[0007] Further: In S2, the method of representing the scene as a group of Gaussians distribution is specifically: The sparse point cloud is combined with the camera extrinsic parameters and converted into the screen space coordinate system. The distance from each point in the sparse point cloud to the camera center is obtained and used as the true depth of each point. The invisible points under the camera's perspective are filtered through the view cone, and the point with the smallest camera depth corresponding to the true depth of the remaining points is determined as the surface point, and the surface points are initialized as a set of Gaussians distribution; The Gaussians distribution includes several Gaussians, each of which includes a 3D Gaussian mean, a covariance matrix, opacity, and a spherical harmonic function. The specific expression is:
[0008] In the formula, is any point in 3D space, , is the mean vector, , R is a matrix, is the covariance matrix, , used to maintain The positive semidefinite property of .
[0009] Further: In S4, the color of the pixel is obtained by the Gaussians passed through by the viewing direction, and the pixel p Color The specific expression is:
[0010] In the formula, is a learnable, view-dependent color representation using spherical harmonics, For the evaluation from i Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point through which the viewing direction passes. For the evaluation from j Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point that the viewing direction passes through. Represents the occlusion effect of the previous Gaussians on the current Gaussian.
[0011] Further: In S5, the loss function The specific expression is:
[0012] In the formula, is the first hyperparameter, which is used to control the weight of SSIM structural similarity loss in the loss function. is the second hyperparameter, which is used to control the weight of the depth loss in the loss function. is the mean absolute error loss between the real image and the rendered image, is the SSIM structural similarity loss based on real images and rendered images, The depth of the camera corresponding to the center distance of the downsampled Gaussians and the actual depth chamfer distance obtained from the sparse point cloud are expressed as follows:
[0013] In the formula, is the set of predicted depth components, is the set of true depths of each point in the sparse point cloud, d i is the prediction depth of the specific Gaussian, d j is the true depth of a specific point in the sparse point cloud, is the L1 norm.
[0014] Further: in S5, after the number of iterations reaches 500, adaptive density control is performed every 100 iterations to update and optimize the Gaussians distribution. The method of adaptive density control is specifically as follows: For each Gaussian, track the magnitude of the position gradient of all rendered views, generate the average position gradient of all rendered views, and perform a split or clone operation in response to the gradient of the position being greater than the average gradient threshold during the iteration. The specific method is as follows: Determine whether the scale of the current Gaussian is less than or equal to the scale threshold. If so, perform a clone operation; if not, perform a split operation.
[0015] Further: in S5, after the number of iterations reaches 500, a curvature-based Gaussians optimization is performed every 100 iterations to update and optimize the uniformity of the Gaussians distribution. The curvature-based Gaussians optimization method is specifically as follows: During the iteration process, the number of Gaussians in each quadrant is counted, and the quadrant with the highest density is selected as the dense quadrant. A Gaussian is randomly selected in the dense quadrant as the center position, and the Euclidean distance from other Gaussians in the dense quadrant to the center is calculated. The Gaussian whose Euclidean distance is less than the distance threshold is selected, and the selected Gaussian is optimized.
[0016] Further: The method for optimizing the selected Gaussian is specifically as follows: Calculate the curvature of the selected Gaussian and delete the Gaussian whose curvature is greater than the curvature threshold. The specific method for calculating the curvature is: Use a fast nearest neighbor search algorithm based on KD tree to efficiently find the neighbors of all Gaussians in the range , and calculate the covariance matrix at the center of the Gaussian. The covariance matrix The specific expression is:
[0017] In the formula, n is the number of other Gaussians in the search range centered on this Gaussian, The coordinates of the center position of the Gaussian. To find the coordinate values of the remaining Gaussians within the range by the proximity search algorithm, T is the transposition symbol; Based on 3 3, the covariance matrix, determines the first to third eigenvalues through the relationship between the eigenvalues and the corresponding eigenvectors, and then calculates the curvature; Among them, the expression of the relationship between the eigenvalue and the corresponding eigenvector is specifically:
[0018] In the formula, v is the feature vector, is the characteristic value; Calculating curvature The specific expression is:
[0019] In the formula, is the first eigenvalue, is the second eigenvalue, is the third eigenvalue.
[0020] Further: in S5, a rescaling operation is performed every 300 iterations or every 1000 iterations, and the method of the rescaling operation is specifically as follows: The scale of Gaussians is controlled by combining the predicted depth with a scaling factor, where the scaling factor The specific expression is:
[0021] In the formula, For each Gaussian prediction depth, is the minimum predicted depth, is the maximum predicted depth, is a hyperparameter.
[0022] The beneficial effects of the present invention are as follows: the present invention provides an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision, which has the following effects compared with the prior art: (1) Introduction of loss function: By combining the true depth of the sparse point cloud generated by SFM and the predicted depth generated by fast differentiable rasterization, the chamfer distance is calculated and used as the loss function. This optimization strengthens the geometric constraints on the position of Gaussians, ensuring the correct position of Gaussians in space, thereby avoiding incorrect occlusion relationships and rendering artifacts, and improving the realism of the scene.
[0023] (2) Rescaling operation: The scale of the Gaussians is adjusted according to the distance from the center of the Gaussians to the camera. This operation ensures that the Gaussians have the correct occlusion relationship during the rendering process. Especially when rendering in the view direction, the scale and occlusion relationship of the Gaussians are accurately maintained, avoiding the appearance of artifacts and improving the rendering quality.
[0024] (3) Curvature as a constraint: To address the problem of small and dense Gaussians generated in high-gradient areas, the improved ADC algorithm combines curvature as a new constraint, especially the high-curvature Gaussians that suddenly appear in the smooth area. These Gaussians often do not conform to the geometric characteristics of the area and therefore need to be removed. Through this operation, redundant small Gaussians are reduced, the computational complexity is reduced, and the negative impact on model optimization is avoided.
[0025] (4) Improving rendering quality and performance: By performing more precise geometric constraints and scale adjustments on Gaussians, the occlusion relationship between Gaussians during the rendering process and their performance in three-dimensional space are improved, ultimately improving the accuracy of scene reconstruction, especially in the rendering of complex scenes, avoiding performance bottlenecks caused by unreasonable Gaussians distribution.
[0026] In summary, compared with the existing technology, the advantage of the present invention is that it solves the limitations of geometric constraint processing in the existing 3DGS technology through joint supervision of depth and curvature, avoids erroneous occlusion and excessive small Gaussians that may occur in traditional methods, and significantly optimizes the quality and efficiency of scene rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of the improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision of the present invention. DETAILED DESCRIPTION
[0028] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0029] like Figure 1 As shown, in one embodiment of the present invention, an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision includes the following steps: S1. Obtain a sparse corresponding set of image features based on the real-time image collected by the camera, use the SfM technology to generate a sparse point cloud representation of the scene, and estimate the camera's motion posture to obtain the camera extrinsic parameters; S2. Using sparse point clouds, the scene is represented as a set of Gaussians, where the center position is the spatial position of the point cloud, the constant of the spherical harmonic function is the color of the corresponding pixel of the point cloud, the high-order coefficient is set to zero, the three-axis scale is set to the average distance of the three point clouds adjacent to the center position, and the axis is represented by a quaternion, which is initialized as , opacity is set to 0.1 and the number of iterations is initialized; S3, combined with the camera external parameters, use the perspective transformation matrix to project the Gaussians distribution onto the two-dimensional plane under the camera perspective, obtain the 2D Gaussians projected onto the image plane, and obtain the predicted depth from it; S4, sorting the 2D Gaussians by depth, determining the color of each pixel to be rendered according to the specific point through which the viewing direction passes, and generating a rendered image; S5, determine whether the number of iterations reaches the preset iteration threshold, if so, enter S6, if not, calculate the loss function according to the real depth, predicted depth, rendered image and RGB real image, update the optimized Gaussians distribution according to the calculation result, increase the number of iterations by 1 and return to S3; S6. Output the final rendered image.
[0030] In this embodiment, in order to solve the limitations of geometric constraints in the existing model optimization process and improve the rendering quality and performance, the present invention obtains the real depth and the distance of the Gaussians generated by fast differentiable rasterization from the sparse point cloud generated by SFM, and the depth of the camera corresponding to the distance is used as the predicted depth, and the chamfer distance calculated by the predicted depth and the real depth is used as the loss function to strengthen the geometric constraints on the position of the Gaussians. At the same time, according to the distance of the center of the Gaussians from the center of the camera, the scale is reset to ensure that when rendering along the viewing direction, the Gaussians have the correct occlusion relationship and have a more accurate frequency representation. In the dense area after the ADC algorithm is optimized, the present invention introduces curvature as a new constraint. For example, the high curvature Gaussian that suddenly appears in the smooth area indicates that the Gaussian is not suitable for the geometric features of the area, and it is removed to prevent a large number of small and dense Gaussians from increasing the computational complexity and hindering the model optimization.
[0031] The S2 is specifically: The sparse point cloud is combined with the camera extrinsic parameters and converted into the screen space coordinate system. The distance from each point in the sparse point cloud to the camera center is obtained and used as the true depth of each point. The invisible points under the camera's perspective are filtered through the view cone, and the point with the smallest camera depth corresponding to the true depth of the remaining points is determined as the surface point, and the surface points are initialized as a set of Gaussians distribution; The Gaussians distribution includes several Gaussians, each of which includes a 3D Gaussian mean, a covariance matrix, opacity, and a spherical harmonic function. The specific expression is:
[0032] In the formula, is any point in 3D space, , is the mean vector, , R is a matrix, is the covariance matrix, , used to maintain The positive semidefinite property of , which can be expressed as:
[0033] In the formula, R is a matrix, specifically representing a 3 3's rotation matrix, which is converted from quaternion, S For 3 3's scaling matrix, which describes the scaling transformation in 3D space.
[0034] In this embodiment, for the parameters of Gaussian, the covariance matrix is generated by the scale and the quaternion q, and the opacity and spherical harmonics are used to fit the appearance and color in the viewing direction.
[0035] In S3, the perspective transformation matrix is used in combination with the camera external parameters. Project the Gaussians distribution onto the two-dimensional plane from the camera's perspective:
[0036] In the formula, is the Jacobian matrix of the projective transformation, which is used to calculate the covariance matrix in the two-dimensional image space.
[0037] In S4, the color of the pixel is obtained by the Gaussians passed through by the viewing direction. p Color The specific expression is:
[0038] In the formula, is a learnable, view-dependent color representation using spherical harmonics, For the evaluation from i Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point through which the viewing direction passes. For the evaluation from j Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point through which the viewing direction passes. Used to indicate the occlusion effect of the previous Gaussians on the current Gaussian.
[0039] In this embodiment, It is not necessarily the center of the Gaussian. It is obtained based on the specific point through which the viewing direction passes. And the specific point through which the viewing direction passes Distance , and The exponential decay coefficient is defined as , which is used to calculate the opacity of specific points passing through Gaussians, which is equivalent to a decay index for the opacity of the center position , where c is the opacity corresponding to the center position of the Gaussian, The definition is as follows:
[0040] In the formula, yes x The variance of the coordinates, x The dispersion of the coordinates, for y The variance of the coordinates, y The dispersion of the coordinates, For x and y The covariance between the coordinates, denoted x and y Linear relationship of coordinates.
[0041] In S5, the loss function The specific expression is:
[0042] In the formula, is the first hyperparameter, which is used to control the weight of SSIM structural similarity loss in the loss function. is the second hyperparameter, which is used to control the weight of the depth loss in the loss function. is the mean absolute error loss between the real image and the rendered image, is the SSIM structural similarity loss based on real images and rendered images, The chamfer distance is the distance between the camera depth and the real depth obtained from the sparse point cloud based on the downsampled Gaussians center distance. The specific expression is:
[0043] In the formula, is the set of predicted depth components, is the set of true depths of each point in the sparse point cloud, d i is the prediction depth of the specific Gaussian, d j is the true depth of a specific point in the sparse point cloud, is the L1 norm.
[0044] In this embodiment, based on the obtained loss function result, back propagation is performed to update and optimize each Gaussian parameter of the Gaussians distribution in S2, which includes the 3D Gaussian mean, the covariance matrix generated by the scale and the quaternion q, and each Gaussian also includes opacity and spherical harmonics.
[0045] In S5, after the number of iterations reaches 500, adaptive density control is performed every 100 iterations to update and optimize the Gaussians distribution. The method of adaptive density control is specifically as follows: For each Gaussian, track the magnitude of the position gradient of all rendered views, generate the average position gradient of all rendered views, and perform a split or clone operation in response to the gradient of the position being greater than the average gradient threshold during the iteration. The specific method is as follows: Determine the current Gaussian scale s Is it less than or equal to the scale threshold? If so, it means that the coverage density of the Gaussian is low, and the cloning operation can be used to enhance the density. If not, it means that the coverage area density of the Gaussian is high, and the splitting operation can be used to split the large Gaussian into multiple small Gaussians.
[0046] In this embodiment, if the average position gradient of each Gaussian exceeds a predefined average gradient threshold, it is considered that the Gaussian cannot adequately represent the corresponding 3D region. In areas with insufficient geometric features (i.e., under-reconstructed areas), the density is increased by cloning existing 3D Gaussians. In over-reconstructed areas, a large 3D Gaussian may cover an area that should be represented by multiple small Gaussians. In this case, the large 3D Gaussian is split into two smaller Gaussians to better capture the details.
[0047] In S5, after the number of iterations reaches 500, a curvature-based Gaussians optimization is performed every 100 iterations to update and optimize the uniformity of the Gaussians distribution. The curvature-based Gaussians optimization method is specifically as follows: During the iteration process, the number of Gaussians in each quadrant is counted, and the quadrant with the highest density is selected as the dense quadrant. A Gaussian is randomly selected in the dense quadrant as the center position, and the Euclidean distance from other Gaussians in the dense quadrant to the center is calculated. The Gaussian whose Euclidean distance is less than the distance threshold is selected, and the selected Gaussian is optimized.
[0048] The specific method for optimizing the selected Gaussian is: Calculate the curvature of the selected Gaussian and delete the Gaussian whose curvature is greater than the curvature threshold. The specific method for calculating the curvature is: Use a fast nearest neighbor search algorithm based on KD tree to efficiently find the neighbors of all Gaussians in the range , and calculate the covariance matrix at the center of the Gaussian. The covariance matrix The specific expression is:
[0049] In the formula, n is the number of other Gaussians in the search range centered on this Gaussian, The coordinates of the center position of the Gaussian. To find the coordinate values of the remaining Gaussians within the range by the proximity search algorithm, T is the transposition symbol; Based on 3 3, the covariance matrix, determines the first to third eigenvalues through the relationship between the eigenvalues and the corresponding eigenvectors, and then calculates the curvature; Among them, the expression of the relationship between the eigenvalue and the corresponding eigenvector is specifically:
[0050] In the formula, v is the feature vector, is the characteristic value; Calculating curvature The specific expression is:
[0051] In the formula, is the first eigenvalue, is the second eigenvalue, is the third eigenvalue.
[0052] In this embodiment, a fast nearest neighbor search algorithm based on a KD tree is used to efficiently find all neighborhood points of all Gaussians within the range, and the covariance matrix of the Gaussian is calculated. The curvature of each Gaussian position is calculated using the eigenvalue obtained from the covariance matrix. When it is greater than a given curvature threshold, the present invention will remove it, such as high curvature points on a smooth surface, to reduce the appearance of a large number of small and dense Gaussians in the dense quadrant.
[0053] In S5, a rescaling operation is performed every 300 iterations or every 1000 iterations. The specific method of the rescaling operation is: The scale of Gaussians is controlled by combining the predicted depth with a scaling factor, where the scaling factor The specific expression is:
[0054] In the formula, For each Gaussian prediction depth, is the minimum predicted depth, is the maximum predicted depth, is a hyperparameter.
[0055] In this embodiment, a scaling factor is defined To optimize the scale of Gaussians, the distance from the center of Gaussians to the center of the camera is used as the control basis and normalized. In order to further control the scaling factor, the present invention introduces a hyperparameter , so that the scaling factor is limited to .
[0056] Small Gaussians represent high frequency information, and large Gaussians do the opposite. This scaling factor can reduce the scale of Gaussians closer to the camera plane and increase the scale of Gaussians farther from the camera plane, making their frequency representation more accurate.
[0057] The beneficial effects of the present invention are as follows: the present invention provides an improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision, which has the following effects compared with the prior art: (1) Introduction of loss function: By combining the true depth of the sparse point cloud generated by SFM and the predicted depth generated by fast differentiable rasterization, the chamfer distance is calculated and used as the loss function. This optimization strengthens the geometric constraints on the position of Gaussians, ensuring the correct position of Gaussians in space, thereby avoiding incorrect occlusion relationships and rendering artifacts, and improving the realism of the scene.
[0058] (2) Rescaling operation: The scale of the Gaussians is adjusted according to the distance from the center of the Gaussians to the camera. This operation ensures that the Gaussians have the correct occlusion relationship during the rendering process. Especially when rendering in the view direction, the scale and occlusion relationship of the Gaussians are accurately maintained, avoiding the appearance of artifacts and improving the rendering quality.
[0059] (3) Curvature as a constraint: To address the problem of small and dense Gaussians generated in high-gradient areas, the improved ADC algorithm combines curvature as a new constraint, especially the high-curvature Gaussians that suddenly appear in the smooth area. These Gaussians often do not conform to the geometric characteristics of the area and therefore need to be removed. Through this operation, redundant small Gaussians are reduced, the computational complexity is reduced, and the negative impact on model optimization is avoided.
[0060] (4) Improving rendering quality and performance: By performing more precise geometric constraints and scale adjustments on Gaussians, the occlusion relationship between Gaussians during the rendering process and their performance in three-dimensional space are improved, ultimately improving the accuracy of scene reconstruction, especially in the rendering of complex scenes, avoiding performance bottlenecks caused by unreasonable Gaussians distribution.
[0061] In summary, compared with the existing technology, the advantage of the present invention is that it solves the limitations of geometric constraint processing in the existing 3DGS technology through joint supervision of depth and curvature, avoids erroneous occlusion and excessive small Gaussians that may occur in traditional methods, and significantly optimizes the quality and efficiency of scene rendering.
[0062] In the description of the present invention, it is necessary to understand that the orientation or positional relationship indicated by the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", and "third" are used only for descriptive purposes, and cannot be understood as indicating or implying the relative importance or the number of implicitly specified technical features. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of the features.
Claims
1. An improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision, characterized in that: The following steps are involved: S1. Obtain a sparse corresponding set of image features based on the real-time image collected by the camera, use the SfM technology to generate a sparse point cloud representation of the scene, and estimate the camera's motion posture to obtain the camera extrinsic parameters; S2. Using sparse point clouds, the scene is represented as a set of Gaussians, where the center position is the spatial position of the point cloud, the constant of the spherical harmonic function is the color of the corresponding pixel of the point cloud, the high-order coefficient is set to zero, the three-axis scale is set to the average distance of the three point clouds adjacent to the center position, and the axis is represented by a quaternion, which is initialized as , opacity is set to 0.1, and the number of iterations is initialized; S3, combined with the camera external parameters, use the perspective transformation matrix to project the Gaussians distribution onto the two-dimensional plane under the camera perspective, obtain the 2D Gaussians projected onto the image plane, and obtain the predicted depth from it; S4, sorting the 2D Gaussians by depth, determining the color of each pixel to be rendered according to the specific point through which the viewing direction passes, and generating a rendered image; S5, determine whether the number of iterations reaches the preset iteration threshold, if so, enter S6, if not, calculate the loss function according to the real depth, predicted depth, rendered image and RGB real image, update the optimized Gaussians distribution according to the calculation result, increase the number of iterations by 1 and return to S3; S6. Output the final rendered image.
2. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1 is characterized in that: In S2, the method of representing the scene as a set of Gaussians distribution is specifically as follows: The sparse point cloud is combined with the camera extrinsic parameters and converted into the screen space coordinate system. The distance from each point in the sparse point cloud to the camera center is obtained and used as the true depth of each point. The invisible points under the camera's perspective are filtered through the view cone, and the point with the smallest camera depth corresponding to the true depth of the remaining points is determined as the surface point, and the surface points are initialized as a set of Gaussians distribution; The Gaussians distribution includes several Gaussians, each of which includes a 3D Gaussian mean, a covariance matrix, opacity, and a spherical harmonic function. The specific expression is: In the formula, is any point in 3D space, , is the mean vector, , R is a matrix, is the covariance matrix, , used to maintain The positive semidefinite property of .
3. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1 is characterized in that: In S4, the color of the pixel is obtained by the Gaussians passed through by the viewing direction. p Color The specific expression is: In the formula, is a learnable, view-dependent color representation using spherical harmonics, For the evaluation from i Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point through which the viewing direction passes. For the evaluation from j Gaussian projection to pixel p The 2D Gaussian gets the opacity of the specific point that the viewing direction passes through.
4. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1 is characterized in that: In S5, the loss function The specific expression is: In the formula, is the first hyperparameter, which is used to control the weight of SSIM structural similarity loss in the loss function. is the second hyperparameter, which is used to control the weight of the depth loss in the loss function. is the mean absolute error loss between the real image and the rendered image, is the SSIM structural similarity loss based on real images and rendered images, The depth of the camera corresponding to the center distance of the downsampled Gaussians and the actual depth chamfer distance obtained from the sparse point cloud are expressed as follows: In the formula, is the set of predicted depth components, is the set of true depths of each point in the sparse point cloud, d i is the prediction depth of the specific Gaussian, d j is the true depth of a specific point in the sparse point cloud, is the L1 norm.
5. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1, characterized in that: In S5, after the number of iterations reaches 500, adaptive density control is performed every 100 iterations to update and optimize the Gaussians distribution. The method of adaptive density control is specifically as follows: For each Gaussian, track the magnitude of the position gradient of all rendered views, generate the average position gradient of all rendered views, and perform a split or clone operation in response to the gradient of the position being greater than the average gradient threshold during the iteration. The specific method is as follows: Determine whether the scale of the current Gaussian is less than or equal to the scale threshold. If so, perform a clone operation; if not, perform a split operation.
6. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1, characterized in that: In S5, after the number of iterations reaches 500, a curvature-based Gaussians optimization is performed every 100 iterations to update and optimize the uniformity of the Gaussians distribution. The curvature-based Gaussians optimization method is specifically as follows: During the iteration process, the number of Gaussians in each quadrant is counted, and the quadrant with the highest density is selected as the dense quadrant. A Gaussian is randomly selected in the dense quadrant as the center position, and the Euclidean distance from other Gaussians in the dense quadrant to the center is calculated. The Gaussian whose Euclidean distance is less than the distance threshold is selected, and the selected Gaussian is optimized.
7. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 6 is characterized in that: The specific method for optimizing the selected Gaussian is: Calculate the curvature of the selected Gaussian and delete the Gaussian whose curvature is greater than the curvature threshold. The specific method for calculating the curvature is: Use a fast nearest neighbor search algorithm based on KD tree to efficiently find the neighbors of all Gaussians in the range , and calculate the covariance matrix at the center of the Gaussian. The covariance matrix The specific expression is: In the formula, n is the number of other Gaussians in the search range centered on this Gaussian, The coordinates of the center position of the Gaussian. To find the coordinate values of the remaining Gaussians within the range by the proximity search algorithm, T is the transposition symbol; Based on 3 3, the covariance matrix, determines the first to third eigenvalues through the relationship between the eigenvalues and the corresponding eigenvectors, and then calculates the curvature; Among them, the expression of the relationship between the eigenvalue and the corresponding eigenvector is specifically: In the formula, v is the feature vector, is the characteristic value; Calculating curvature The specific expression is: In the formula, is the first eigenvalue, is the second eigenvalue, is the third eigenvalue.
8. The improved 3DGS scene reconstruction and rendering method based on depth and curvature supervision according to claim 1 is characterized in that: In S5, a rescaling operation is performed every 300 iterations or every 1000 iterations. The specific method of the rescaling operation is: The scale of Gaussians is controlled by combining the predicted depth with a scaling factor, where the scaling factor The specific expression is: In the formula, For each Gaussian prediction depth, is the minimum predicted depth, is the maximum predicted depth, is a hyperparameter.
Citation Information
Patent Citations
Indoor complex scene high-fidelity real-time rendering method based on three-dimensional Gaussian representation
CN118096988A
Inverse rendering method and device based on 3D Gaussian, equipment and storage medium
CN118644605A
Methods, systems, and computer program products for processing three-dimensional image data to render an image from a viewpoint within or beyond an occluding region of the image data
US20090103793A1