Three-dimensional scene model pruning method based on rendering importance

Through the three-dimensional scene model pruning method based on render importance, the geometry with low render importance is eliminated, and the problem of redundant data in three-dimensional reconstruction is solved, computing efficiency and storage utilization are improved, and it is suitable for efficient reconstruction of static and dynamic environments.

CN120451400AActive Publication Date: 2025-08-08BEIJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510553211.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology has high computational complexity, high memory consumption when processing redundant data, and is difficult to achieve real-time and efficientness. Especially in large-scale and dynamic environments, redundant data affects the quality of reconstruction.

Method used

Through the three-dimensional scene model pruning method based on the render importance, keyframes and key points are initialized, target frames are randomly selected for rendering optimization, render importance is calculated and geometry below the threshold is eliminated, and geometry is used to measure render importance, so as to achieve automatic pruning of geometry.

Benefits of technology

On the premise of ensuring that the rendering effect remains unchanged, the number of redundant geometry is significantly reduced, the computing and storage needs are reduced, the rendering efficiency is improved, and three-dimensional reconstruction of different scales and dynamic environments are adapted to three-dimensional reconstructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451400A_ABST
    Figure CN120451400A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional scene model pruning method based on rendering importance, and belongs to the technical field of computer vision and robots. The method comprises the following steps: receiving key frames and key points which need to be reconstructed as input, initializing the key points into geometries in a 3D scene, and constructing a key frame set according to a time sequence; one key frame is randomly selected as a target frame for rendering, and parameters of the geometry are optimized through back propagation, so that a rendering result is close to a real scene as much as possible; the rendering importance of all geometries in the current scene is calculated, the geometries with the rendering importance lower than a preset threshold value are removed, and a three-dimensional scene rendering result after pruning is obtained; and the rendering importance is obtained by multiplying an opacity factor, a visual angle factor and a zoom factor of the geometry in the current three-dimensional scene and then normalizing. On the premise of ensuring that the rendering effect is not affected, the number of redundant geometries is effectively reduced, meanwhile, the calculation overhead and the storage requirement are reduced, and the rendering efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and robotics technology, and in particular relates to a three-dimensional scene model pruning method based on rendering importance. Background Art

[0002] The application of 3D reconstruction technology to real-world scenes faces many challenges, one prominent issue being the cumulative redundancy of scene geometry. When reconstructing a 3D scene, traditional methods typically construct a scene model using sparse point clouds or other geometric representations. However, as information accumulates during the reconstruction process, these methods often generate a large number of overlapping and repeated feature points or geometric elements, resulting in a waste of computational resources and potentially introducing noise and unnecessary details. Especially in large-scale and dynamic environments, models often generate redundant data. This redundancy not only leads to storage and computational burdens but can also compromise the quality of the final reconstruction.

[0003] Therefore, efficiently processing and optimizing this redundant data, avoiding unnecessary duplication and ineffective computation, is key to improving 3D reconstruction performance. Solving this problem requires not only considering how to effectively represent each part of the scene, but also pruning and optimizing redundant information to improve computational efficiency, reduce storage requirements, and ultimately achieve more accurate 3D reconstruction.

[0004] Existing classic 3D reconstruction methods, such as feature point matching, bundle adjustment, point cloud compression, and multi-view geometry optimization, can effectively handle redundant geometry, but they often suffer from high computational complexity, large memory consumption, and loss of accuracy. This is especially true in large-scale datasets or large-scale scenarios, where achieving real-time and high efficiency is often difficult. Deep learning-based methods, while demonstrating greater capabilities in redundant data processing, require large amounts of labeled data during training. Furthermore, the model's high computational overhead and adaptability to large-scale environments remain bottlenecks restricting its application. Therefore, maintaining high accuracy while reducing computational resources and storage requirements has become a key issue in improving 3D reconstruction efficiency. Summary of the Invention

[0005] To address the above problems, the present invention proposes a 3D scene model pruning method based on rendering importance, which aims to remove redundant information through the "rendering importance" of geometric bodies in the scene, thereby significantly improving computational efficiency and reducing storage requirements while ensuring the accuracy of 3D reconstruction.

[0006] The present invention provides a 3D scene model pruning method based on rendering importance, comprising the following steps:

[0007] In step 1, the initialization module receives the keyframes and keypoints to be reconstructed, initializes the keypoints to geometric objects in the 3D scene, and maintains the keyframes in chronological order as a keyframe set. At this point, the geometric objects in the scene serve as the basic units of the scene and begin to construct a 3D representation.

[0008] In step 2, the optimization module randomly selects a frame from the key frame set as the target frame and uses the target frame as the rendering benchmark; the scene is rendered according to the camera pose of the target frame, and compared with the real scene to calculate the rendering error loss; finally, the parameters of the geometric body are optimized through back propagation to improve the scene rendering effect, so that the rendering result is as close as possible to the real image of the target frame.

[0009] Step 3: After each backpropagation, the rendering importance-based pruning module calculates the rendering importance of all geometric bodies in the current 3D scene. According to the set threshold τ, the geometric bodies with rendering importance lower than the threshold τ are eliminated to obtain the pruned 3D scene rendering result; determine whether there are new key points added. If so, continue to step 1, otherwise, end the current 3D scene rendering task.

[0010] Among them, the opacity factor α, the viewing angle factor V and the scaling factor S of the geometric body in the current three-dimensional scene are multiplied to obtain the rendering importance d = α·V·S, and then normalized to obtain the rendering importance score I = N(d) of the geometric body. The normalization function x max and x min The maximum value d of the rendering importance of all geometric bodies in the current 3D scene max and the minimum value d min .

[0011] In the step 1, in the 3D Gaussian rendering scene, the initialization module initializes the input key points into 3D Gaussian geometry, and the 3D Gaussian geometry is represented by Gaussian spheres. The parameters of each Gaussian sphere include position, covariance matrix, opacity, and size scaling value.

[0012] In step 3, the opacity factor α of the geometric body is calculated as follows:

[0013]

[0014] Among them, w a is the opacity weight, set w a =2.0; opacity is the opacity of the geometry;

[0015] The view factor V of the geometry is calculated as follows:

[0016]

[0017] Among them, θ is the angle between the center of the geometry and the axis of the camera frustum, It is the vector from the geometry to the camera in the camera coordinate system of the current 3D scene. is the camera direction of the current 3D scene; V min is the minimum frustum weight, set V min =0.3; w v Is the frustum weight, set w v =1.0;

[0018] The scaling factor S of the geometry is calculated as follows:

[0019]

[0020] Among them, w s is the scaling weight, set w s = 0.5; scale geometry size scaling value, N (scale) is the size scaling value using the normalization function N (x) processing, when processing x max and x min They correspond to the maximum and minimum values of the size scaling values of all geometries in the current 3D scene.

[0021] The advantages and positive effects of the present invention are:

[0022] (1) While ensuring that rendering quality is not affected, the method of the present invention effectively reduces the number of redundant geometric entities, while also reducing computational overhead and storage requirements, improving rendering efficiency, and enhancing the real-time processing capabilities of 3D scene reconstruction. Furthermore, the method of the present invention is adaptable to scenes of varying scales, enabling efficient scene reconstruction in both static and dynamic environments.

[0023] (2) The method of the present invention generates geometric bodies from key points, generates a three-dimensional scene by randomly extracting key frames, reversely optimizes the geometric body parameters, and then designs the opacity factor, viewing angle factor and scaling factor of the optimized geometric body to effectively characterize the rendering importance of the geometric body, and automatically removes the geometric body based on the threshold value according to the rendering importance. The present invention improves the calculation method of the rendering importance, and organically combines the overall method to reduce computing resources and storage requirements, achieve more effective removal of redundant geometric bodies, and maintain the stability of the quality of the reconstructed scene, thereby improving the efficiency of three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flowchart of the overall implementation of the 3D scene model pruning method based on rendering importance according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The technical solution provided by the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0026] The present invention discloses a rendering importance-based 3D scene model pruning method. The method first receives the keyframes to be reconstructed and their corresponding keypoints as input. The keypoints are initialized as 3D Gaussian bodies, and the input keyframe set is used as the basis for mapping. A keyframe is then randomly selected as the rendering target frame. This frame is then rendered and the parameters of the 3D Gaussian body are trained to ensure that the rendering result is as close to the real scene as possible. In this way, the system gradually builds a 3D representation of the scene.

[0027] In order to solve the problem of redundant scene geometry, especially in large-scale or dynamic environments, the system of the present invention defines and calculates the rendering importance of all geometries after training. Specifically, rendering importance is a measure of the contribution of the geometry in the scene. Geometries with lower importance usually contribute less to the reconstruction results, so they can be considered for elimination. To achieve this goal, the method of the present invention sets an importance threshold τ, and geometries below this threshold will be automatically removed. In this way, not only can the quality of scene reconstruction be maintained, but the rendering rate can also be greatly improved.

[0028] Take 3D Gaussian rendering as an example, Figure 1 As shown, the 3D scene model pruning method based on rendering importance according to an embodiment of the present invention includes the following three steps: The system corresponding to the method of the present invention includes an initialization module, an optimization module, and a pruning module based on rendering importance.

[0029] Step 1: Initialization. The initialization module performs initialization operations based on the input key point data, constructs the geometric bodies in the 3D scene, obtains the initial 3D scene, and organizes and maintains the input key frames into a key frame set in a certain way.

[0030] In this embodiment of the present invention, when reconstructing a 3D Gaussian scene, the initialization module processes the input keypoints and keyframes, initializing the keypoints into 3D Gaussian bodies. The received keyframes are then organized into a set of keyframes in chronological order. At this point, the 3D Gaussian bodies serve as the basic units of the scene, and a 3D representation begins to be constructed.

[0031] Step 1.1, initialize the 3D Gaussian geometry.

[0032] First, the input three-dimensional key points are initialized as 3D Gaussian geometries. These geometries are not just simple points, but a Gaussian distribution described by a covariance matrix. The covariance matrix reflects the spatial uncertainty of the point, which is usually related to the local environment of the point, such as the texture of the image and the camera's perspective. The key points received by the method of the present invention represent objects in a three-dimensional scene. They are all initialized as a geometry based on the position of the key points. The position of the geometry is then optimized in the optimization module.

[0033] In three-dimensional reconstruction, 3D Gaussian geometries can be used to represent irregular or complex scene shapes, especially in dynamic or sparse environments. These Gaussian geometries are represented by Gaussian spheres, and the parameters of each Gaussian sphere include the three-dimensional coordinates of the geometric position, covariance information, rotation matrix, opacity, scaling, color or appearance parameters, etc. The color or appearance parameters are encoded using surface texture SH spherical harmonics. The parameters of the Gaussian sphere will be optimized through training. The three-dimensional coordinates and covariance information of the Gaussian sphere initialized in the embodiment of the present invention are expressed as follows:

[0034] p=(x,y,z)

[0035]

[0036] Among them, p is the position of the 3D Gaussian geometry, represented by the center position of the Gaussian sphere, and initially corresponds to the three-dimensional spatial coordinates of the key point; the covariance matrix Σ is a 3×3 matrix that represents the uncertainty of the key point position, including the uncertainty of the three-dimensional point in different directions. xy ,σ xz ,σ zy Represents the covariance between the three coordinate axes x, y, and z. Represents the variance in the directions of the three coordinate axes x, u, and z.

[0037] Step 1.2: Maintain a keyframe set.

[0038] The 3D scene uses an ordered queue to maintain a collection of received keyframes. Each keyframe typically contains information such as the keyframe pose and keyframe feature points. As new keyframes are added, the system continuously expands this collection; at the same time, the queue periodically deletes expired keyframes.

[0039] In step 2, the optimization module randomly selects a target frame from the set of keyframes maintained in step 1 and uses it as the rendering baseline. The 3D Gaussian scene is rendered based on the camera pose of the target frame and compared with the real scene to calculate the rendering error loss. Finally, backpropagation is used to optimize the parameters of the 3D Gaussian geometry, ensuring that the rendered result is as close to the real image of the target frame as possible.

[0040] Step 2.1: Select a target frame. Randomly select a frame from the key frame set as the target frame to calculate the error between the current scene and the real scene. This embodiment of the present invention adopts a random selection strategy, that is, randomly select a frame from the key frame set as the target frame.

[0041] Step 2.2: Render the scene. Based on the pose of the selected target frame, first calculate the camera projection matrix P as follows:

[0042] P=K·[R|t]

[0043] Among them, K is the intrinsic parameter matrix of the camera, R and t are the rotation matrix and translation vector respectively.

[0044] Then, by projecting each Gaussian sphere in the 3D Gaussian scene, the position on the two-dimensional image is obtained. This is done by matrix multiplication of the three-dimensional position (x, y, z) of each Gaussian sphere with the camera projection matrix to obtain the two-dimensional coordinates (u, v):

[0045]

[0046] After projection, the position of each Gaussian sphere in the scene can be obtained in the image.

[0047] Step 2.3, calculate the rendering error loss. The goal of the loss function is to minimize the error between the rendered scene and the real image. Specifically, the embodiment of the present invention compares the rendering result of the target frame with the real image and calculates the loss function as follows:

[0048]

[0049] Among them, I r is the rendered image of the target frame, I gt is the real image of the target frame, SSIM(I r ,I gt ) represents the structural similarity between two images, and λ is a weight factor for balancing.

[0050] Once the total loss is calculated Backpropagation can be used to calculate the gradient of the loss with respect to the 3D Gaussian geometry parameters such as position, covariance, opacity, and scale, and these parameters are updated using the Adam optimization algorithm. By minimizing the loss function, the optimization process gradually adjusts the parameters of the Gaussian sphere to make the rendering closer to the true value.

[0051] In step 3, after each backpropagation step, the system calculates the rendering importance of all 3D Gaussian objects in the scene. Based on these values, an importance threshold τ is set. Objects with importance below the threshold τ are automatically removed. This process effectively removes redundant objects, maintaining stable 3D Gaussian scene quality while reducing computational overhead and improving rendering efficiency.

[0052] Step 3.1, after each iterative training, the rendering importance-based pruning module of the present invention calculates a current rendering importance for all geometric bodies according to the characteristics of all geometric bodies in the current scene, including three-dimensional coordinates, opacity, size, texture, etc., and then performs a pruning operation.

[0053] In the embodiment of the present invention, the factors α, V and S of the geometric body in the current three-dimensional scene are multiplied to obtain the rendering importance d, and then the maximum value d of the rendering importance of all geometric bodies in the current three-dimensional scene is calculated. max and the minimum value d min , normalize d and normalize the rendering importance to the value range [0.1,1].

[0054] The embodiment of the present invention calculates the rendering importance score I of a certain geometric body as follows:

[0055] I=N(α·V·S)

[0056] Where: α is the opacity factor, V is the viewing angle factor, S is the scaling factor; N(x) is the normalization function, and the normalization calculation of the input value s is as follows:

[0057] Here x, x max and x min Corresponding to d, d max and d min .

[0058] Opacity is crucial for the effect of geometry during rendering. The system sorts geometry based on depth during rendering, and then renders it sequentially based on transmittance, or residual opacity. Therefore, geometry with greater opacity has a greater impact on the scene, while geometry with less opacity has a smaller impact on the scene rendering and is therefore less important. The opacity factor α is calculated in this embodiment of the present invention as follows:

[0059]

[0060] where w a = 2.0 is the opacity weight, opacity is the opacity of the 3D Gaussian geometry.

[0061] When reconstructing three-dimensional objects, the objects near the camera center are often of primary interest. Therefore, the method of the present invention establishes a view frustum factor to represent the relative distance of a geometric object from the camera center. Specifically, this factor is represented by the cosine value of the angle θ between the center of the geometric object and the axis of the camera's view frustum. A larger cosine value indicates that the geometric object is closer to the camera center and therefore more important. The embodiment of the present invention calculates the view factor V as follows:

[0062]

[0063] in, is the vector from the 3D Gaussian geometry to the camera in the camera coordinate system of the current scene, is the camera direction of the current scene, V min =0.3 is the minimum frustum weight, w v is the view cone weight. In the embodiment of the present invention, w is set v =1.0.

[0064] The scaling factor S of the geometry is calculated as follows:

[0065]

[0066] Among them, N (scale) is the normalization of the size scaling value of the geometry. When using N (x) normalization, x max and x min They correspond to the maximum and minimum values of the size scaling values of all geometric bodies in the current 3D scene. s = 0.5 is the scaling weight, and scale is the scaling value of the 3D Gaussian geometry. Both opacity and scale are assigned an initial value to be optimized during the initialization phase and then optimized in the optimization module.

[0067] In step 3.2, the pruning module performs geometry pruning based on rendering importance.

[0068] After calculating the rendering importance of all 3D Gaussian objects, we prune them. Based on the characteristics of different indoor scene datasets, we set an appropriate importance threshold τ. 3D Gaussian objects below this threshold are removed to obtain the pruned 3D scene rendering result.

[0069] Then continue to determine whether there are new key points added. If so, continue to step 1. Otherwise, end the current 3D scene rendering task.

[0070] In general, the various embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects can be implemented in hardware, while other aspects can be implemented in firmware or software executed by a controller, microprocessor or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controller or other computing device, or some combination thereof, as non-limiting examples.

[0071] Except for the technical features described in the specification, all other technical features are known to those skilled in the art. The present invention omits descriptions of well-known components and well-known technologies to avoid redundancy and unnecessary limitation of the present invention. The implementation methods described in the above embodiments do not represent all implementation methods consistent with the present application. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the present invention.

Claims

1. A 3D scene model pruning method based on rendering importance, characterized in that: include: Step 1: The initialization module receives key frames and key points, initializes each key point as a geometric body in the three-dimensional scene, and maintains the received key frames into a key frame set in chronological order; Step 2: The optimization module randomly selects a frame from the current keyframe set as the target frame, uses the camera pose of the target frame to render the scene, compares it with the real scene, calculates the rendering error loss, and optimizes the parameters of the geometry through backpropagation to improve the scene rendering effect; Step 3: After each backpropagation, the rendering importance-based pruning module calculates the rendering importance of all geometric objects in the current 3D scene. Based on the set threshold τ, geometric objects with rendering importance lower than the threshold τ are removed to obtain the pruned 3D scene rendering result. The module then determines whether new key points have been added. If so, the module proceeds to step 1; otherwise, the current 3D scene rendering task is terminated. Among them, the opacity factor α, the viewing angle factor V and the scaling factor S of the geometric body in the current three-dimensional scene are multiplied to obtain the rendering importance d = α·V·S, and then normalized to obtain the rendering importance score I = N(d) of the geometric body. The normalization function x max and x min The maximum value d of the rendering importance of all geometric bodies in the current 3D scene max and the minimum value d min .

2. The method according to claim 1, characterized in that In the step 1, in the 3D Gaussian rendering scene, the initialization module initializes the input key points into 3D Gaussian geometry, and the 3D Gaussian geometry is represented by Gaussian spheres. The parameters of each Gaussian sphere include position, covariance matrix, opacity, and size scaling value.

3. The method according to claim 1 or 2, characterized in that In step 3, the opacity factor α of the geometric body is calculated as follows: Among them, w a is the opacity weight, set w a =2.0; opacity is the opacity of the geometry; The view factor V of the geometry is calculated as follows: Among them, θ is the angle between the center of the geometry and the axis of the camera frustum, It is the vector from the geometry to the camera in the camera coordinate system of the current 3D scene. is the camera direction of the current 3D scene; V min is the minimum frustum weight, set V min =0.3; w v Is the frustum weight, set w v =1.0; The scaling factor S of the geometry is calculated as follows: Among them, w s is the scaling weight, set w s = 0.5; scale geometry size scaling value, N (scale) is the size scaling value using the normalization function N (x) processing, when processing x max and x min They correspond to the maximum and minimum values of the size scaling values of all geometries in the current 3D scene.

Citation Information

Patent Citations

  • Gaussian rendering and reconstruction method based on prior guidance of symbol distance radiation field

    CN118314268A

  • Large-scale three-dimensional scene real-time reconstruction method based on Gaussian expression

    CN118314280A

  • Three-dimensional reconstruction and rendering method based on non-parameter image

    CN118840490A