Three-dimensional scene reconstruction rendering method and system

By constructing a splitting control function and adaptively adjusting the Gaussian low-pass filter, the problems of insufficient accuracy and efficiency in the existing 3D Gaussian ellipsoid splitting strategy are solved, high-precision 3D scene reconstruction and rendering are achieved, aliasing artifacts are reduced, and the calculation process is optimized.

CN120635277AActive Publication Date: 2025-09-12JIANGNAN UNIV

Patent Information

Application Number
CN202510746162.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing 3D Gaussian ellipsoid splitting strategy only relies on the edge response intensity of the image, ignoring local texture changes, scale changes and spatial distribution relationships, resulting in reduced 3D scene reconstruction rendering accuracy and poor computational efficiency.

Method used

By constructing a splitting control function and comprehensively considering edge strength, local variance, normalized gradient amplitude, dot product of gradient vector and normal vector, and Gaussian kernel function value, the key edge parts in the two-dimensional image can be accurately identified. The density and size of the Gaussian sphere are adjusted according to different regions. A Gaussian low-pass filter is introduced to adaptively adjust the filter parameters to optimize the splitting and pruning process of the 3D Gaussian ellipsoid.

Benefits of technology

It effectively alleviates the aliasing artifact phenomenon, improves the rendering accuracy and computational efficiency of 3D scene reconstruction, and ensures a balance between rendering quality and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635277A_ABST
    Figure CN120635277A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a three-dimensional scene reconstruction rendering method and system. The method comprises the following steps: generating a corresponding 3D Gaussian ellipsoid for each sparse point cloud of a scene to be reconstructed; obtaining a re-projection area of each 3D Gaussian ellipsoid on the two-dimensional image of different visual angles; taking each re-projection area as a center, and taking all re-projection areas within a set pixel distance range as adjacent re-projection areas of the re-projection area; the method is based on the sum of Gaussian kernel function values from the center of a re-projection area of each 3D Gaussian ellipsoid on a two-dimensional image of different visual angles to the center of each adjacent re-projection area of the area, the edge strength, the local variance, the normalized gradient magnitude, and the absolute value of the dot product of a gradient vector and a normal vector. Calculating a split control function value of each re-projection area; and comparing the splitting control function value of each re-projection area with a set threshold value, and controlling the 3D Gaussian ellipsoid to split. According to the invention, the reconstruction rendering precision of the three-dimensional scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a three-dimensional scene reconstruction and rendering method and system. Background Art

[0002] 3D scene reconstruction and rendering aims to recover the 3D geometry and appearance of a scene from 2D image data. It achieves 3D modeling of the scene through geometric analysis, feature matching, and data fusion of multi-view images. Rendering techniques then transform the 3D model into a 2D image that reflects the scene's geometry, lighting effects, material properties, and other 3D information, intuitively presenting the scene's details. This technology is widely used in film and television production, game development, virtual reality, and other fields. Compared to traditional neural radiance field (NeRF) methods, 3D Gaussian sputtering (3DGS) significantly improves rendering efficiency while maintaining high-fidelity rendering quality. 3DGS utilizes Structure from Motion (SfM) technology to process a series of 2D images of the reconstructed 3D scene. SfM matches and tracks feature points in multi-view images to estimate the camera's pose and the 3D position of the sparse point cloud in the scene, generating a sparse point cloud and camera parameters. An initialized 3D Gaussian ellipsoid is created for each sparse point cloud. Each 3D Gaussian ellipsoid consists of position, covariance matrix, color, and transparency. By optimizing the parameters of the 3D Gaussian ellipsoid, the geometric shape and texture information of the scene can be accurately modeled, so that the Gaussian ellipsoid can better approximate the real scene and obtain a three-dimensional scene model. The optimized 3D Gaussian ellipsoid is projected onto a two-dimensional plane using differentiable tile rasterization technology to generate a two-dimensional rendering for displaying the three-dimensional scene model.

[0003] Although the 3DGS method has made significant progress in rendering efficiency and scene representation, it still faces the challenge of aliasing artifacts when processing complex scenes. Aliasing artifacts are usually caused by the aliasing of low-frequency sampling and high-frequency information, especially in high-contrast areas and detailed edges. When the sampling frequency cannot match the high-frequency details in the scene (such as sharp object edges and complex textures), information loss or mismatch will form step-like or flickering visual artifacts. This phenomenon not only reduces the visual quality of the image, but also affects the accuracy and realism of the rendering algorithm. To address this problem, some studies have made splitting decisions on the 3D Gaussian ellipsoid based on the edge information of the image. The 3D Gaussian ellipsoid at the edge is split, and multiple new Gaussians are used to increase the number of Gaussian splits at the edge, thereby refining the sampling density of the local area and making the sampling frequency more closely match the distribution of high-frequency information, thereby reducing the aliasing of low-frequency sampling and high-frequency information, and to some extent alleviating the defects of aliasing artifacts. However, current technologies typically rely solely on the edge response strength of an image (such as the gradient magnitude or the output of an edge detection algorithm) to determine the splitting of the 3D Gaussian ellipsoid, ignoring other important features in the image, such as local texture variations, scale variations, and spatial distribution relationships. For example, although some areas exhibit strong edge features in edge detection, due to low texture complexity or density, excessive 3D Gaussian ellipsoid splitting may lead to overfitting, which not only fails to further improve rendering accuracy but also increases unnecessary computational overhead. For areas with rich texture details but weak edge gradients, if the splitting decision is made based solely on edge strength, aliasing will result due to insufficient sampling. Therefore, a 3D Gaussian ellipsoid splitting strategy is still needed that can effectively balance rendering quality and computational efficiency while resolving aliasing. Summary of the Invention

[0004] To this end, the technical problem to be solved by the present invention is to overcome the defect that the existing 3D Gaussian ellipsoid splitting strategy only relies on the edge response intensity of the image, ignores other important features in the image such as local texture changes, scale changes and spatial distribution relationships, and is difficult to deal with the residual aliasing artifacts, resulting in reduced accuracy of three-dimensional scene reconstruction rendering.

[0005] To solve the above technical problems, the present invention provides a three-dimensional scene reconstruction and rendering method, comprising:

[0006] Step S1: constructing an initial 3D Gaussian ellipsoid set based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed;

[0007] Step S2: obtaining the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; taking each reprojection area as the center, taking all reprojection areas within a set pixel distance range as adjacent reprojection areas of the reprojection area;

[0008] Step S3: Calculate the splitting control function value of each reprojection area based on the sum of the Gaussian kernel function values ​​from the center of each reprojection area to the centers of each adjacent reprojection area in the area, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector at the center of the reprojection area;

[0009] Step S4: Determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, split the 3D Gaussian ellipsoid corresponding to the reprojection area, iteratively update the 3D Gaussian ellipsoid set, and return to execute step S2; if they are all less than or equal to the set threshold, use the current 3D Gaussian ellipsoid set as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

[0010] Preferably, the step of using the current 3D Gaussian ellipsoid set as a scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed includes:

[0011] Based on the size of the reprojected area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different viewing angles, the gradient amplitude and transparency of each pixel point, the 3D Gaussian ellipsoid set is pruned, and the pruned 3D Gaussian ellipsoid set is used to generate a three-dimensional model of the scene to be reconstructed;

[0012] According to the camera extrinsic parameters, the pruned 3D Gaussian ellipsoid set is projected onto the target two-dimensional image plane. After each reprojection area on the target two-dimensional image plane is processed by a Gaussian low-pass filter, the parameters of each processed reprojection area are rasterized through differentiable tiles to generate a two-dimensional rendering of the scene to be reconstructed.

[0013] Preferably, each reprojection region on the target two-dimensional image plane is processed by a Gaussian low-pass filter to obtain parameters of each processed reprojection region, including:

[0014] Calculate the Gaussian low-pass filter parameters based on the width, height, and number of pruned 3D Gaussian ellipsoids of the target 2D image;

[0015] By adding Gaussian low-pass filter parameters to the diagonal elements of the covariance matrix of each reprojection area on the target two-dimensional image plane, the parameters of each processed reprojection area are obtained;

[0016] Among them, the formula for calculating the parameters of the Gaussian low-pass filter is:

[0017]

[0018] Where s is the Gaussian low-pass filter parameter, H is the height of the target two-dimensional image, W is the width of the target two-dimensional image, and K is the number of pruned 3D Gaussian ellipsoids.

[0019] Preferably, the step of adding Gaussian low-pass filter parameters to diagonal elements of the covariance matrix of each reprojection region on the target two-dimensional image plane to obtain parameters of each processed reprojection region comprises:

[0020]

[0021] in, is the position in the kth reprojection area on the target two-dimensional image plane The reprojection area parameters after processing, exp(·) is an exponential function with a natural constant as the base, is any position in the kth reprojection area on the target two-dimensional image plane, μ′ k is the center coordinate of the kth reprojection area on the target 2D image plane, (·) T is the transpose operation, Σ′ k is the covariance matrix of the kth reprojection region, s is the Gaussian low-pass filter parameter, and I is the identity matrix.

[0022] Preferably, the edge strength of the center of the current reprojection area is obtained by weighted summing the amplitudes of the gradients of the three RGB channels at the center of the current reprojection area;

[0023] Calculate the absolute value of the difference between the pixel value at the center of the current reprojection area and the pixel values ​​of other pixels in the area, and take the average as the local variance of the center of the current reprojection area;

[0024] The gradient magnitude at the center of the current reprojection area is normalized to obtain a normalized gradient magnitude at the center of the current reprojection area.

[0025] Preferably, the calculation formula for the normalized gradient amplitude at the center of the current reprojection area is:

[0026]

[0027] Among them, S gauss (x, y) is the normalized gradient magnitude at the center of the current reprojection area, is the gradient vector at the center of the current reprojection area, ∈ is the regularization term, ‖·‖ is the Euclidean norm, indicating the amplitude, is the gradient vector of the center of the current reprojection area, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, and y is the vertical coordinate of the center of the current reprojection area.

[0028] Preferably, the calculation formula for the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area is:

[0029]

[0030] Among them, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of the adjacent reprojection areas in the area, N is the number of adjacent reprojection areas to the center of the current reprojection area, exp(·) is an exponential function with a natural constant as the base, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, y is the vertical coordinate of the center of the current reprojection area, (x i ,y i ) is the coordinate of the center of the i-th adjacent reprojection area of ​​the current reprojection area center, ‖·‖ 2 is the square of the L2 norm, i is the index of the adjacent reprojection area, and σ is the parameter that controls the influence range of the 3D Gaussian ellipsoid.

[0031] Preferably, the splitting control function value of the current reprojection area is calculated based on the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area of ​​the area, the edge strength, the local variance, the normalized gradient amplitude, and the absolute value of the dot product of the gradient vector and the normal vector. The formula is:

[0032]

[0033] Among them, S(x, y) is the splitting control function value of the current reprojection area, E img (x, y) is the edge strength of the center of the current reprojection area, α1 is the local variance weight, is the local variance of the center of the current reprojection area, α2 is the normalized gradient amplitude weight, S gauss (x, y) is the normalized gradient amplitude at the center of the current reprojection area, α3 is the gradient direction matching weight, D direction (x, y) is the absolute value of the dot product of the gradient vector and the normal vector at the center of the current reprojection area, α4 is the spatial coupling effect weight, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area.

[0034] Preferably, the process of obtaining the sparse point cloud of the scene to be reconstructed includes:

[0035] A set of two-dimensional images of the scene to be reconstructed from different perspectives are used through the SfM algorithm to generate a sparse point cloud of the scene to be reconstructed.

[0036] The present invention also provides a three-dimensional scene reconstruction and rendering system, comprising:

[0037] An initial generation module is used to construct an initial 3D Gaussian ellipsoid set based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed;

[0038] The reprojection area mapping module is used to obtain the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; with each reprojection area as the center, all reprojection areas within a set pixel distance range are regarded as adjacent reprojection areas of the reprojection area;

[0039] a splitting control function calculation module, configured to calculate the splitting control function value of each reprojection region based on the sum of the Gaussian kernel function values ​​from the center of each reprojection region to the centers of each adjacent reprojection region of the region, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector of the center of the reprojection region;

[0040] The generation module is used to determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, the 3D Gaussian ellipsoid corresponding to the reprojection area is split, and the 3D Gaussian ellipsoid set is iteratively updated, and the step of executing the reprojection area mapping module is returned; if they are all less than or equal to the set threshold, the current 3D Gaussian ellipsoid set is used as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

[0041] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0042] The present invention discloses a method and system for reconstructing and rendering a three-dimensional scene. The present invention determines whether to split the 3D Gaussian ellipsoid by constructing a splitting control function and comparing the splitting control function value of the reprojected area of ​​each 3D Gaussian ellipsoid on a two-dimensional image at different viewing angles with the size of a set threshold. The function can more accurately identify the real and critical edge parts in the two-dimensional image by calculating the edge strength of the center of the reprojected area, enhance the sampling of high-frequency details, and avoid insufficient or excessive splitting due to edge misjudgment. The introduction of the local variance of the center of the reprojected area can reflect the texture complexity of the area, guide the Gaussian ellipsoid to increase the density of the Gaussian ellipsoid in the area with rich details, and improve the splitting accuracy. The normalized gradient amplitude of the center of the reprojected area is closely related to the boundary and structural size of the object in the image. The size of the Gaussian ellipsoid can be adaptively adjusted according to different image areas, such as detail areas, edge areas, etc., effectively avoiding excessive refinement in large-scale areas, while increasing the density of the Gaussian ellipsoid in the detail area to improve the accuracy of reconstruction. Edges in images usually have clear directionality, especially at the contours of objects or structural boundaries. By calculating the absolute value of the dot product of the gradient vector and the normal vector at the center of the reprojection area, the directional changes in the edge area can be accurately identified, thereby enhancing the splitting strength in the edge area and reducing false splitting caused by insufficient directional changes. Considering that the 3D Gaussian ellipsoid does not exist independently in space, by calculating the sum of the Gaussian kernel function values ​​from the center of each reprojection area to the center of each adjacent reprojection area in the area, the information of the surrounding Gaussian spheres can be fully considered in the splitting process, avoiding excessive or insufficient splitting caused by isolated decisions, and maintaining the consistency and stability of the scene representation. Through multi-level perception of high-frequency information, the splitting decision of the 3D Gaussian ellipsoid is matched with the actual high-frequency distribution in the scene, effectively alleviating the jagged artifact phenomenon in the two-dimensional rendering and improving the rendering accuracy of the three-dimensional scene reconstruction.

[0043] The present invention calculates Gaussian low-pass filter parameters based on the width, height and number of pruned 3D Gaussian ellipsoids of the target two-dimensional image, accurately controls the coverage area of ​​the processed reprojection area parameters on the target two-dimensional image plane, and ensures that each reprojection area covers the necessary minimum pixel area within the effective gradient propagation range by adding the calculated Gaussian low-pass filter parameters to the diagonal elements of the covariance matrix. Compared with the traditional fixed Gaussian low-pass filter parameter values, the present invention can adaptively set the Gaussian low-pass filter parameter values ​​according to the width, height and number of pruned 3D Gaussian ellipsoids of the target two-dimensional image, automatically increase the Gaussian low-pass filter parameters in high-resolution or Gaussian sparse areas, enhance the smoothing effect of the processed reprojection area parameters to suppress jagged artifacts, and automatically reduce the Gaussian low-pass filter parameters in low-resolution or Gaussian dense areas, retain the detailed features of the reprojection area to avoid excessive blurring, and effectively improve the rendering accuracy of the two-dimensional rendering image. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0045] Figure 1 It is a flowchart of the steps of a three-dimensional scene reconstruction and rendering method of the present invention.

[0046] Figure 2 It is a flow chart of a three-dimensional scene reconstruction and rendering method of the present invention.

[0047] Figure 3 This is a schematic diagram comparing the process of the present invention and the traditional pruning method. Figure 3 (a) is a flowchart of the traditional pruning method. Figure 3 (b) is a schematic flow chart of the pruning method of the present invention.

[0048] Figure 4 This is a schematic diagram comparing the effects of a three-dimensional scene reconstruction rendering method of the present invention and other methods on the MipNerf360 dataset.

[0049] Figure 5 This is a schematic diagram comparing the effects of a three-dimensional scene reconstruction rendering method of the present invention and other methods in a low-frequency and high-frequency mixed scene. DETAILED DESCRIPTION

[0050] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0051] Reference Figure 1 、 Figure 2 As shown, this embodiment provides a 3D scene reconstruction and rendering method, including the following steps:

[0052] Step S1: Based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed, an initial set of 3D Gaussian ellipsoids is constructed; each 3D Gaussian ellipsoid is defined by position (mean), covariance matrix and opacity α, while the directional appearance component (color) of the radiation field is represented by spherical harmonics SH;

[0053] In this embodiment, specifically, the process of acquiring the sparse point cloud of the scene to be reconstructed includes:

[0054] A set of two-dimensional images of the scene to be reconstructed from different perspectives are used through the SfM algorithm to generate a sparse point cloud of the scene to be reconstructed.

[0055] Step S2: obtaining the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; taking each reprojection area as the center, taking all reprojection areas within a set pixel distance range as adjacent reprojection areas of the reprojection area;

[0056] Wherein, with the center position of each reprojection area as the center, the distance between all pixel points in the adjacent reprojection areas of the current reprojection area and the center of the current reprojection area is less than or equal to the set pixel distance.

[0057] In this embodiment, specifically, each 3D Gaussian ellipsoid is mapped onto a two-dimensional image at different perspectives through multi-perspective projection, thereby obtaining a reprojection area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different perspectives. Each reprojection area on a two-dimensional image corresponds to a 3D Gaussian ellipsoid.

[0058] Step S3: Calculate the splitting control function value of each reprojection area based on the sum of the Gaussian kernel function values ​​from the center of each reprojection area to the centers of each adjacent reprojection area in the area, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector at the center of the reprojection area;

[0059] In this embodiment, specifically, the edge strength of the center of the current reprojection area is obtained by weighted summing the amplitudes of the gradients of the three RGB channels at the center of the current reprojection area;

[0060] The calculation formula for the edge strength of the center of the current reprojection area is:

[0061]

[0062] Among them, E img (x, y) is the edge strength of the center of the current reprojection area, α is the gradient weight of the R channel, β is the gradient weight of the G channel, and γ is the gradient weight of the B channel. is the gradient vector of the R channel at (x, y), is the gradient vector of the G channel at (x, y), is the gradient vector of the B channel at (x, y), (x, y) is the coordinate of the center of the current reprojection area, x is the abscissa of the center of the current reprojection area, y is the ordinate of the center of the current reprojection area, ‖·‖ is the Euclidean norm, indicating the amplitude.

[0063] By calculating the edge strength at the center of the reprojected area, we can more accurately identify the real and critical edge parts in the two-dimensional image, enhance the sampling of high-frequency details, and avoid insufficient or excessive splitting due to edge misjudgment.

[0064] Calculate the absolute value of the difference between the pixel value at the center of the current reprojection area and the pixel values ​​of other pixels in the area, and take the average as the local variance of the center of the current reprojection area;

[0065] The local variance of the center of the current reprojection area is calculated as:

[0066]

[0067] in, is the local variance of the center of the current reprojection area, M is the number of pixels in the current reprojection area except the center, I(x, y) is the pixel value of the center of the current reprojection area, I(x j ,y j ) is the pixel value of the j-th pixel point outside the center of the current reprojection area.

[0068] The local variance of the center of the reprojected area reflects the texture complexity of the area in the image. In the texture-complex area of ​​the image, even if the edge detection algorithm fails to produce a strong response, the change in local density can still indicate that the area requires more 3D Gaussian ellipsoids for fine modeling. By introducing this factor, the density of 3D Gaussian ellipsoids can be increased in areas with rich details, thereby improving the splitting accuracy.

[0069] Normalize the gradient amplitude at the center of the current reprojection area to obtain the normalized gradient amplitude at the center of the current reprojection area;

[0070] The calculation formula for the normalized gradient amplitude at the center of the current reprojection area is:

[0071]

[0072] Among them, S gauss (x, y) is the normalized gradient magnitude at the center of the current reprojection area, is the gradient vector at the center of the current reprojection area, ∈ is a regularization term to avoid division by zero errors, ‖·‖ is the Euclidean norm, indicating the amplitude, is the gradient vector of the center of the current reprojection area, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, and y is the vertical coordinate of the center of the current reprojection area.

[0073] The normalized gradient magnitude at the center of the reprojected region reflects the scale variation of the 3D Gaussian ellipsoid, which is closely related to the boundaries and structural dimensions of objects in the image. By considering the scale variation of the Gaussian sphere, the size of the Gaussian sphere can be adaptively adjusted according to different image regions (such as detail areas and edge areas). This approach effectively avoids over-refinement in large-scale areas while increasing the density of the 3D Gaussian ellipsoid in detail areas, thereby improving the accuracy of 3D model reconstruction.

[0074] The calculation formula for the absolute value of the dot product of the gradient vector and the normal vector at the center of the current reprojection area is:

[0075]

[0076] Among them, D direction (x, y) is the absolute value of the dot product of the gradient vector and the normal vector at the center of the current reprojection area, is the gradient vector at the center of the current reprojection area, is the normal vector of the center of the current reprojection area, and |·| is the absolute value;

[0077] The edges in an image usually have clear directionality, especially at the contours or structural boundaries of an object. By calculating the matching degree between the gradient direction of the center of the reprojected area and the normal vector, the directional changes in the edge area can be accurately identified, thereby enhancing the splitting strength in the edge area, effectively improving the splitting accuracy and reducing false splitting caused by insufficient directional changes.

[0078] The calculation formula for the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area is:

[0079]

[0080] Among them, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of the adjacent reprojection areas in the area, N is the number of adjacent reprojection areas to the center of the current reprojection area, exp(·) is an exponential function with a natural constant as the base, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, y is the vertical coordinate of the center of the current reprojection area, (x i ,y i ) is the coordinate of the center of the i-th adjacent reprojection area of ​​the current reprojection area center, ‖·‖ 2 is the square of the L2 norm, i is the index of the adjacent reprojection area, and σ is the parameter that controls the influence range of the 3D Gaussian ellipsoid.

[0081] Gaussian spheres do not exist independently in space, and the interaction between adjacent Gaussian spheres may affect the splitting decision. The sum of the Gaussian kernel function values ​​from the center of the reprojection area to the centers of each adjacent reprojection area in the area can reflect the spatial coupling effect of the 3D Gaussian ellipsoid, so that the information of the surrounding 3D Gaussian ellipsoids can be taken into account in the splitting decision process, avoiding excessive or insufficient splitting, avoiding the problem of local over-refinement or coarsening in traditional methods, and improving the coordination and consistency of 3D model reconstruction.

[0082] In this embodiment, preferably, the splitting control function value of the current reprojection area is calculated based on the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area of ​​the area, the edge strength, the local variance, the normalized gradient amplitude, and the absolute value of the dot product of the gradient vector and the normal vector. The formula is:

[0083]

[0084] Among them, S(x, y) is the splitting control function value of the current reprojection area, E img (x, y) is the edge strength of the center of the current reprojection area, α1 is the local variance weight, is the local variance of the center of the current reprojection area, α2 is the normalized gradient amplitude weight, S gauss (x, y) is the normalized gradient amplitude at the center of the current reprojection area, α3 is the gradient direction matching weight, D direction (x, y) is the absolute value of the dot product of the gradient vector and the normal vector at the center of the current reprojection area, α4 is the spatial coupling effect weight, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area.

[0085] The present invention uses the edge strength at the center of the current reprojection area as the basic weight to ensure that splitting occurs preferentially in high-contrast edge areas, directly processing the high-frequency sources of aliasing artifacts, ensuring the stability of the basic edge detection mechanism, while allowing each dimensional feature to provide positive or negative correction when necessary. This design enables the splitting strategy to inherit the advantages of traditional edge detection and accurately locate high-frequency information areas through the synergistic effect of multi-dimensional features.

[0086] By designing a splitting control function, this paper considers not only edge strength but also local density variations, scale variations, directional information, and the spatial coupling effect between 3D Gaussian ellipsoids. This multi-dimensional approach overcomes the limitations of traditional methods that rely solely on edge detection and provides a more comprehensive and accurate basis for splitting decisions.

[0087] Step S4: Determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, split the 3D Gaussian ellipsoid corresponding to the reprojection area, iteratively update the 3D Gaussian ellipsoid set, and return to execute step S2; if they are all less than or equal to the set threshold, use the current 3D Gaussian ellipsoid set as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

[0088] like Figure 3 As shown, Figure 3This is a schematic diagram comparing the process of the present invention and the traditional pruning method. Figure 3 (a) is a flowchart of the traditional pruning method. Figure 3 (b) is a schematic flow chart of the pruning method of the present invention.

[0089] In this embodiment, preferably, the step of using the current 3D Gaussian ellipsoid set as a scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed includes:

[0090] Based on the size of the reprojected area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different viewing angles, the gradient amplitude and transparency of each pixel point, the 3D Gaussian ellipsoid set is pruned, and the pruned 3D Gaussian ellipsoid set is used to generate a three-dimensional model of the scene to be reconstructed;

[0091] According to the camera extrinsic parameters, the pruned 3D Gaussian ellipsoid set is projected onto the target two-dimensional image plane (i.e., splatting). After processing each reprojected area on the target two-dimensional image plane through a Gaussian low-pass filter, the parameters of each processed reprojected area are rasterized through differentiable tiles to generate a two-dimensional rendering of the scene to be reconstructed.

[0092] In this embodiment, optionally, pruning the 3D Gaussian ellipsoid set based on the size of the reprojected area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different viewing angles, the gradient amplitude and transparency of each pixel point includes:

[0093] Determine whether the average gradient amplitude of the current 3D Gaussian ellipsoid reprojection area is less than the set gradient threshold. If so, remove the current 3D Gaussian ellipsoid.

[0094] Determine whether the average transparency of the current 3D Gaussian ellipsoid reprojection area is less than the set transparency threshold. If so, remove the current 3D Gaussian ellipsoid.

[0095] Determine whether the size of the current 3D Gaussian ellipsoid reprojection area is smaller than the set size threshold. If so, remove the current 3D Gaussian ellipsoid.

[0096] Compared with the traditional pruning method, which only performs pruning based on a single pruning condition, such as when the Gaussian weight is lower than a threshold, it is easy to cause the Gaussian in the key detail area to be mistakenly deleted or redundant Gaussian remains. The present invention prunes the 3D Gaussian ellipsoid set according to the size of the reprojection area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different perspectives, the gradient amplitude and transparency of each pixel point, thereby achieving refined screening of Gaussians, suppressing 3D Gaussian ellipsoids with low transparency, reducing redundant model elements that contribute little to the rendering results, retaining more Gaussians in areas with high gradient amplitude, ensuring the sampling density of high-frequency details, eliminating Gaussians with too small a reprojection area, avoiding invalid calculations, and retaining large-size Gaussians to maintain the macroscopic representation of the scene structure, which can effectively avoid the aliasing problem caused by the aliasing of high-frequency information and low-frequency sampling.

[0097] The 3D Gaussian ellipsoid is determined by the full 3D covariance matrix defined in world space, centered at the point (mean) μ, and its spatial distribution is as follows:

[0098]

[0099] In the rendering blending process, the Gaussian radiation value is modulated by the weight factor. To achieve rendering, the 3D Gaussian ellipsoid needs to be projected onto the 2D image plane. Given the view transformation matrix W, the covariance matrix Σ in the camera coordinate system ′ It can be derived by the affine approximate Jacobian matrix J of the projection transformation, the formula is: Σ ′ =JWΣW T J T , ignoring Σ ′ After the third dimension is obtained, a two-dimensional variance matrix can be obtained, whose structure is consistent with the projection method based on the plane point normal.

[0100] Another approach is to optimize the covariance matrix Σ to obtain a 3DGS representation of the radiation field. However, the covariance matrix is ​​physically meaningful only when it is positive semidefinite. Gradient descent is commonly used to optimize all parameters, but this approach is difficult to constrain to generate valid matrices, and the update steps and gradients can easily produce invalid covariance matrices.

[0101] Therefore, the present invention chooses a more intuitive but equally expressive representation method for optimization. The covariance matrix of 3DGS is similar to describing the shape of an ellipsoid. Given a scaling matrix S and a rotation matrix R, the corresponding Σ can be found using the formula: Σ = RSS T R TIn order to optimize these two factors independently, by storing the scale and rotation components separately, the scale is represented by a 3D vector s, and the rotation is represented by a quaternion q, which can be easily converted into their respective matrices and combined together. It is ensured that q is normalized, which can effectively avoid the problem of invalid matrix generation and obtain a valid unit quaternion.

[0102] In order to avoid the significant computational overhead caused by automatic differentiation during training, the gradients of each parameter are explicitly derived to avoid the computational overhead of automatic differentiation, so that the 3D Gaussian ellipsoid can be adaptively optimized into a compact representation of complex geometric structures in the scene (such as planes, curved surfaces, and anisotropic objects).

[0103] In this embodiment, preferably, each reprojection area on the target two-dimensional image plane is processed by a Gaussian low-pass filter to obtain parameters of each processed reprojection area, including:

[0104] Calculate the Gaussian low-pass filter parameters based on the width, height, and number of pruned 3D Gaussian ellipsoids of the target 2D image;

[0105] By adding Gaussian low-pass filter parameters to the diagonal elements of the covariance matrix of each reprojection area on the target two-dimensional image plane, the parameters of each processed reprojection area are obtained;

[0106] Among them, the formula for calculating the parameters of the Gaussian low-pass filter is:

[0107]

[0108] Where s is the Gaussian low-pass filter parameter, H is the height of the target two-dimensional image, W is the width of the target two-dimensional image, and K is the number of pruned 3D Gaussian ellipsoids.

[0109] The Gaussian low-pass filter is a common image processing tool used to suppress high-frequency components, thereby reducing noise and artifacts caused by over-sharpening. Its core concept is to perform local weighted averaging within an image to achieve smoothing. Unlike traditional mean filters, the Gaussian filter weights pixel neighborhoods according to a normal distribution, giving smaller weights to pixels farther from the center, resulting in a more natural smoothing of image details.

[0110] Jagged edges are caused by high-frequency aliasing, which is particularly noticeable when the image sampling rate is insufficient. Because edge regions often contain rich high-frequency information, low-frequency sampling cannot effectively capture these details, resulting in information folding and the formation of artifacts.

[0111] Some studies have found that when the reprojected area in the projected target two-dimensional image plane is smaller than the size of a single pixel, using them directly will cause visual artifacts. Therefore, the scale of the reprojected area in the projected target two-dimensional image plane is expanded by adding a small value to the diagonal elements of the covariance matrix. The formula is as follows:

[0112]

[0113] in, is the position in the kth reprojection area on the target two-dimensional image plane The reprojection area parameters after processing, exp(·) is an exponential function with a natural constant as the base, is any position in the kth reprojection area on the target two-dimensional image plane, μ′ k is the center coordinate of the kth reprojection area on the target 2D image plane, (·) T is the transpose operation, Σ′ k is the covariance matrix of the kth reprojection region, s is the Gaussian low-pass filter parameter, and I is the identity matrix.

[0114] This process can also be understood as the reprojection area in the target two-dimensional image plane obtained by projection With a mean μ = 0, variance Convolution is performed with the Gaussian low-pass filter h, that is, This proves to be a key step in preventing aliasing. After convolution with the low-pass filter, the reprojected area on the target 2D image plane is approximated as a circle, and the radius of this circle is given by the 2D covariance matrix Σ′ k +sI is defined as three times the larger eigenvalue.

[0115] In this study, the Gaussian low-pass filter parameter s is a predefined value, usually set to 0.3. However, the present invention notes that the Gaussian low-pass filter parameter s can ensure that each reprojection region must cover a minimum area in the target two-dimensional image plane. Since the Gaussian only receives gradients within a few standard deviations, learning from a wider area is crucial for the Gaussian learning scene to obtain rough structural information. Therefore, the present invention calculates the Gaussian low-pass filter parameter s to allow the Gaussian to cover a wider area in the early stages of training, and gradually learns from more local areas, ensuring that the minimum area that each reprojection region must cover in the target two-dimensional image plane is greater than 9πs. 2 , s is the area of ​​the target two-dimensional image. The Gaussian low-pass filter makes the image smoother by suppressing high-frequency components, thereby effectively alleviating the jagged edge problem.

[0116] The present invention is different from the traditional global filtering strategy. Instead, it combines the 3D Gaussian ellipsoid splitting strategy, Gaussian low-pass filtering and dynamic pruning mechanism. By adaptively adjusting the scope of the filter, it only acts on the area that needs to be processed, thereby avoiding unnecessary blurring of the non-aliased areas. Specifically, the traditional Gaussian filter usually applies a uniform smoothing process to the entire image, while the 3D Gaussian ellipsoid splitting strategy dynamically splits the Gaussian according to the edge intensity, texture complexity, edge directionality and other features of the two-dimensional projection image (such as the dot product of the gradient vector and the normal vector), generates more sub-Gaussians in high-frequency areas (such as sharp edges, complex textures), improves the local sampling density, and reduces the potential risk of aliasing from the source. The Gaussian dynamic pruning method performs adaptive evaluation on multiple dimensions such as gradient changes, transparency information and screen space size, and only applies filtering operations in high-contrast edge areas or other areas susceptible to aliasing. In this way, the method can effectively reduce aliasing and ensure that the key details of the image are retained, thereby improving the quality of the final rendering result.

[0117] This embodiment compares the 2D rendering with the ground truth, calculates the loss function, updates the parameters of the 3DGS through backpropagation, and feeds the adaptive density control module to further optimize the distribution of the point cloud.

[0118] To validate the proposed 3D scene reconstruction and rendering method, the NeRF-Synthetic dataset was selected as the primary dataset for the experiment. This dataset, proposed by Mildenhall et al. in their seminal work on NeRF, has been widely used in research on view synthesis and 3D scene reconstruction. The NeRF-Synthetic dataset contains eight high-quality synthetic scenes (e.g., Chair, Drums, Ficus, Hotdog, Lego, Materials, Mic, and Ship). Each scene consists of 100 RGB images captured from different viewpoints at a resolution of 800x800. Each image is equipped with precise camera parameters (including position, rotation, and focal length), providing a reliable foundation for model training and evaluation. The scenes in the dataset are highly specialized and complex, featuring rich lighting variations, complex material reflections, and fine geometric structures. This fully tests the model's ability to handle high-frequency details, lighting variations, and complex geometry. Furthermore, the image backgrounds in the dataset are transparent (alpha channel), allowing researchers to flexibly set the background (e.g., solid color or real background) as needed, thereby better simulating real-world application scenarios.

[0119] Although originally designed for NeRF, the NeRF-Synthetic dataset, with its high-quality multi-view images and precise camera parameters, is equally applicable to 3DGS research. By explicitly representing scenes as a series of learnable 3D Gaussian distributions, 3DGS is capable of efficiently rendering complex scenes, achieving significant improvements in rendering speed and quality. The complex lighting and geometry in the NeRF-Synthetic dataset provide an ideal testing environment for 3DGS, enabling validation of the model's performance in high-frequency detail and edge smoothing. Furthermore, the dataset's transparent background design allows 3DGS to more flexibly handle background information, preventing background interference from affecting rendering results. The NeRF-Synthetic dataset was chosen for its widespread recognition, high-quality data, diversity, and transparent background design. This dataset is a benchmark in the field of 3D rendering and has been used in numerous research works, ensuring the reliability and comparability of experimental results. Its high-resolution images and precise camera parameters provide a solid foundation for optimizing and evaluating 3DGS, while the diverse scene types (such as objects, buildings, and natural scenes) validate the model's generalization capabilities across diverse scenarios. Through experiments on this dataset, we can comprehensively evaluate the rendering capabilities of 3DGS in complex scenes and make fair comparisons with existing methods.

[0120] This paper selects Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) as evaluation indicators to comprehensively evaluate the performance of the proposed 3D scene reconstruction and rendering method in the deblurring task. These three evaluation indicators have different advantages and can reflect the quality of the model-restored image from multiple dimensions.

[0121] PSNR is a classic evaluation metric widely used to measure restoration quality in image reconstruction tasks. It reflects the level of image error by calculating the mean squared error (MSE) between the restored image and the original image. Due to its ubiquity and historical background, PSNR remains an indispensable basic metric in the field of image deblurring. SSIM, a perceptually based structural similarity metric, assesses image similarity based on brightness, contrast, and structure. Compared to PSNR, SSIM better simulates the human eye's perception of image quality. SSIM is considered to be closer to the human visual system's cognitive process than PSNR, and therefore more accurately reflects the preservation of image detail and structure when assessing image quality. For deblurring tasks, SSIM effectively assesses the structural consistency of the restored image, avoiding the limitations of relying solely on pixel-level error. LPIPS is a new metric recently proposed for image perceptual quality assessment. LPIPS better reflects the subjective perception of the human visual system, particularly when dealing with complex image content, outperforming traditional PSNR and SSIM. LPIPS is more sensitive to the preservation of image detail and texture, making it particularly suitable for evaluating subtle perceptual differences in image restoration. The combination of these three enables this embodiment to perform a more comprehensive and accurate evaluation of the performance of LPD-3DGS (LowPass Dynamic Prune-3D Gaussian Splatting).

[0122] As shown in the figure, Figure 4 This figure shows a comparison of the performance of a 3D scene reconstruction and rendering method proposed in this invention and other methods on the MipNerf360 dataset. The 3D scene reconstruction and rendering method proposed in this invention (Ours) is compared with GroundTruth, Plenoxels, INCP, and Mip-NeRF.

[0123] As shown in Table 1, Table 1 is a comparison of the PSNR values ​​of the three-dimensional scene reconstruction rendering method of the present invention and the experimental results of Plenoxels, INGP-Base, Mip-NeRF, Point-NeRF, and 3DGS.

[0124] Table 1

[0125]

[0126]

[0127] As shown in Table 1, the proposed method achieves the highest or second highest value in all eight test scenarios, with an average PSNR of 33.79, surpassing cutting-edge methods such as 3DGS (33.64) and Point-NeRF (33.30).

[0128] like Figure 5 As shown, Figure 5 This figure compares the performance of a 3D scene reconstruction rendering method proposed in the present invention with other methods in a mixed low-frequency and high-frequency scene. The 3D scene reconstruction rendering method proposed in the present invention (Ours) is compared with GroundTruth and MSGS in a mixed low-frequency and high-frequency scene.

[0129] As shown in Table 2, Table 2 is a comparison of the PSNR values ​​of a three-dimensional scene reconstruction and rendering method of the present invention and the MSGS experimental results.

[0130] Table 2

[0131]

[0132] according to Figure 5 As shown in Table 2, the superiority of the three-dimensional scene reconstruction and rendering method proposed in the present invention is achieved in low-frequency and high-frequency mixed scenes.

[0133] Plenoxels and INGP are the fastest recent NeRF methods. Plenoxels is a sparse voxel grid-based method for efficiently rendering and optimizing implicitly represented 3D scenes. It encodes the scene into a voxel grid and stores density and color information represented by spherical harmonic coefficients in each voxel. INGP is a method proposed by NVIDIA for fast training of neural scene representations. It uses multi-resolution hash encoding and a small MLP network. Mip-NeRF is an improved NeRF method that better handles anti-aliasing problems and scale changes by introducing the idea of ​​MipMapping (multi-level texture sampling). Experimental results show that the method proposed in this invention is superior to the 3DGS method.

[0134] In order to evaluate the performance contribution of each module in the Gaussian low-pass filter, 3D Gaussian ellipsoid splitting strategy, and dynamic pruning strategy of the present invention, the present invention separately tested the improvement of network performance by the Gaussian low-pass filter, 3D Gaussian ellipsoid splitting strategy, and dynamic pruning strategy, as shown in Table 3, which shows the results of the ablation experiment of each module.

[0135] Table 3

[0136]

[0137] This paper makes several key improvements to the traditional 3DGS method to improve its performance in image anti-aliasing, especially in processing high-contrast edges and complex scenes. Compared with the classic 3DGS method, the method proposed in this paper shows significant advantages in multiple core indicators. It is then experimentally compared with the existing MS-GS method, proving that the three-dimensional scene reconstruction and rendering method proposed in this paper still has advantages in scenes with mixed high and low frequencies.

[0138] The present invention introduces a low-pass filter to suppress high-frequency noise, thereby smoothing the image edges more naturally during the rendering process. Experimental results show that after the introduction of the low-pass filter, the PSNR and SSIM indicators of the image are significantly improved, especially in the recovery of details in complex edge areas. Compared with the original 3DGS method, the addition of the low-pass filter effectively alleviates the aliasing phenomenon in high-contrast areas, and at the same time improves the smoothness of the image.

[0139] The present invention also introduces a dynamic pruning strategy to more flexibly control the application area of ​​the Gaussian filter. Dynamic pruning adaptively determines whether filtering is needed in specific areas by deeply analyzing multidimensional features such as pixel gradient, transparency, and screen space size. This adaptive strategy enables the Gaussian filter to maintain a global smoothing effect while avoiding excessive blurring of detailed areas. Experiments show that compared with the original 3DGS method, the introduction of dynamic pruning significantly enhances the model's generalization ability in different scenarios, especially in rendering high-contrast details and complex backgrounds.

[0140] By integrating a Gaussian low-pass filter, a 3D Gaussian ellipsoid splitting strategy, and a dynamic pruning strategy into the reconstruction and rendering process, this paper significantly improves detail preservation, effectively addressing the detail loss problem caused by Gaussian filtering uniformity in traditional methods. Compared to the original 3DGS method, the proposed joint optimization framework demonstrates higher accuracy when processing detail-rich images, especially in rendering edge transition areas, achieving superior visual quality and significantly improving the overall rendering detail and visual fidelity.

[0141] This second embodiment provides a three-dimensional scene reconstruction and rendering system, including:

[0142] An initial generation module is used to construct an initial 3D Gaussian ellipsoid set based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed;

[0143] The reprojection area mapping module is used to obtain the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; with each reprojection area as the center, all reprojection areas within a set pixel distance range are regarded as adjacent reprojection areas of the reprojection area;

[0144] a splitting control function calculation module, configured to calculate the splitting control function value of each reprojection region based on the sum of the Gaussian kernel function values ​​from the center of each reprojection region to the centers of each adjacent reprojection region of the region, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector of the center of the reprojection region;

[0145] The generation module is used to determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, the 3D Gaussian ellipsoid corresponding to the reprojection area is split, and the 3D Gaussian ellipsoid set is iteratively updated, and the step of executing the reprojection area mapping module is returned; if they are all less than or equal to the set threshold, the current 3D Gaussian ellipsoid set is used as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

[0146] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0148] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0150] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A three-dimensional scene reconstruction and rendering method, characterized in that: include: Step S1: constructing an initial 3D Gaussian ellipsoid set based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed; Step S2: obtaining the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; taking each reprojection area as the center, taking all reprojection areas within a set pixel distance range as adjacent reprojection areas of the reprojection area; Step S3: Calculate the splitting control function value of each reprojection area based on the sum of the Gaussian kernel function values ​​from the center of each reprojection area to the centers of each adjacent reprojection area in the area, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector at the center of the reprojection area; Step S4: Determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, split the 3D Gaussian ellipsoid corresponding to the reprojection area, iteratively update the 3D Gaussian ellipsoid set, and return to execute step S2; if they are all less than or equal to the set threshold, use the current 3D Gaussian ellipsoid set as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

2. A 3D scene reconstruction and rendering method according to claim 1, characterized in that: The method of using the current 3D Gaussian ellipsoid set as a scene densification result to generate a 3D model and a 2D rendering of the scene to be reconstructed includes: Based on the size of the reprojected area of ​​each 3D Gaussian ellipsoid on the two-dimensional image at different viewing angles, the gradient amplitude and transparency of each pixel point, the 3D Gaussian ellipsoid set is pruned, and the pruned 3D Gaussian ellipsoid set is used to generate a three-dimensional model of the scene to be reconstructed; According to the camera extrinsic parameters, the pruned 3D Gaussian ellipsoid set is projected onto the target two-dimensional image plane. After each reprojection area on the target two-dimensional image plane is processed by a Gaussian low-pass filter, the parameters of each processed reprojection area are rasterized through differentiable tiles to generate a two-dimensional rendering of the scene to be reconstructed.

3. The 3D scene reconstruction and rendering method according to claim 2, wherein: Each reprojection area on the target two-dimensional image plane is processed by a Gaussian low-pass filter to obtain the parameters of each processed reprojection area, including: Calculate the Gaussian low-pass filter parameters based on the width, height, and number of pruned 3D Gaussian ellipsoids of the target 2D image; By adding Gaussian low-pass filter parameters to the diagonal elements of the covariance matrix of each reprojection area on the target two-dimensional image plane, the parameters of each processed reprojection area are obtained; Among them, the formula for calculating the parameters of the Gaussian low-pass filter is: Where s is the Gaussian low-pass filter parameter, H is the height of the target two-dimensional image, W is the width of the target two-dimensional image, and K is the number of pruned 3D Gaussian ellipsoids.

4. The 3D scene reconstruction and rendering method according to claim 3, wherein: The method adds Gaussian low-pass filter parameters to the diagonal elements of the covariance matrix of each reprojection area on the target two-dimensional image plane to obtain the parameters of each processed reprojection area, including: in, is the position in the kth reprojection area on the target two-dimensional image plane The reprojection area parameters after processing, exp(·) is an exponential function with a natural constant as the base, is any position in the kth reprojection area on the target two-dimensional image plane, μ ′ k is the center coordinate of the kth reprojection area on the target 2D image plane, (·) T is the transpose operation, Σ ′ k is the covariance matrix of the kth reprojection region, s is the Gaussian low-pass filter parameter, and I is the identity matrix.

5. The 3D scene reconstruction and rendering method according to claim 1, wherein: The edge strength of the center of the current reprojection area is obtained by weighted summing the amplitudes of the RGB three-channel gradients at the center of the current reprojection area; Calculate the absolute value of the difference between the pixel value at the center of the current reprojection area and the pixel values ​​of other pixels in the area, and take the average as the local variance of the center of the current reprojection area; The gradient magnitude at the center of the current reprojection area is normalized to obtain a normalized gradient magnitude at the center of the current reprojection area.

6. The 3D scene reconstruction and rendering method according to claim 1, characterized in that: The calculation formula for the normalized gradient amplitude at the center of the current reprojection area is: Among them, S gauss (x, y) is the normalized gradient magnitude at the center of the current reprojection area, is the gradient vector at the center of the current reprojection area, ∈ is the regularization term, ‖·‖ is the Euclidean norm, indicating the amplitude, is the gradient vector of the center of the current reprojection area, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, and y is the vertical coordinate of the center of the current reprojection area.

7. The 3D scene reconstruction and rendering method according to claim 1, wherein: The calculation formula for the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area is: Among them, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of the adjacent reprojection areas in the area, N is the number of adjacent reprojection areas to the center of the current reprojection area, exp(·) is an exponential function with a natural constant as the base, (x, y) is the coordinate of the center of the current reprojection area, x is the horizontal coordinate of the center of the current reprojection area, y is the vertical coordinate of the center of the current reprojection area, (x i ,y i ) is the coordinate of the center of the i-th adjacent reprojection area of ​​the current reprojection area center, ‖·‖ 2 is the square of the L2 norm, i is the index of the adjacent reprojection area, and σ is the parameter that controls the influence range of the 3D Gaussian ellipsoid.

8. The 3D scene reconstruction and rendering method according to claim 1, wherein: The splitting control function value of the current reprojection area is calculated based on the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area, the edge strength, the local variance, the normalized gradient amplitude, and the absolute value of the dot product of the gradient vector and the normal vector. The formula is: Among them, S(x, y) is the splitting control function value of the current reprojection area, E img (x, y) is the edge strength of the center of the current reprojection area, α1 is the local variance weight, is the local variance of the center of the current reprojection area, α2 is the normalized gradient amplitude weight, S gauss (x, y) is the normalized gradient amplitude at the center of the current reprojection area, α3 is the gradient direction matching weight, D direction (x, y) is the absolute value of the dot product of the gradient vector and the normal vector at the center of the current reprojection area, α4 is the spatial coupling effect weight, C space (x, y) is the sum of the Gaussian kernel function values ​​from the center of the current reprojection area to the centers of each adjacent reprojection area in the area.

9. The 3D scene reconstruction and rendering method according to claim 1, wherein: The process of obtaining the sparse point cloud of the scene to be reconstructed includes: A set of two-dimensional images of the scene to be reconstructed from different perspectives are used through the SfM algorithm to generate a sparse point cloud of the scene to be reconstructed.

10. A three-dimensional scene reconstruction and rendering system, characterized in that: include: An initial generation module is used to construct an initial 3D Gaussian ellipsoid set based on the 3D Gaussian ellipsoid corresponding to each sparse point cloud of the scene to be reconstructed; The reprojection area mapping module is used to obtain the reprojection area of ​​each 3D Gaussian ellipsoid in the 3D Gaussian ellipsoid set on the two-dimensional image at different viewing angles; with each reprojection area as the center, all reprojection areas within a set pixel distance range are regarded as adjacent reprojection areas of the reprojection area; a splitting control function calculation module, configured to calculate the splitting control function value of each reprojection region based on the sum of the Gaussian kernel function values ​​from the center of each reprojection region to the centers of each adjacent reprojection region of the region, as well as the edge strength, local variance, normalized gradient amplitude, and absolute value of the dot product of the gradient vector and the normal vector of the center of the reprojection region; The generation module is used to determine whether the splitting control function values ​​of all reprojection areas are less than or equal to the set threshold; if there is a reprojection area with a splitting control function value greater than the set threshold, the 3D Gaussian ellipsoid corresponding to the reprojection area is split, and the 3D Gaussian ellipsoid set is iteratively updated, and the step of executing the reprojection area mapping module is returned; if they are all less than or equal to the set threshold, the current 3D Gaussian ellipsoid set is used as the scene densification result to generate a three-dimensional model and a two-dimensional rendering of the scene to be reconstructed.

Citation Information

Patent Citations

  • Virtual human arbitrary view angle rendering method and system based on three-dimensional Gaussian spattering

    CN118736092A

  • Underwater three-dimensional reconstruction method and system and storage medium

    CN119516079A

  • Three-dimensional Gaussian spattering reconstruction method and related equipment

    CN119693544A

  • Space target point cloud three-dimensional reconstruction and rendering method

    CN119919587A

  • Method for 3D scene dense reconstruction based on monocular visual slam

    US20200273190A1

Cited By

  • Static scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing

    CN121414978A