Three-dimensional scene-oriented real-time rendering method based on spatial perception
Through a three-dimensional scene rendering method based on space perception, scene geometric constraints are extracted using multi-view images and depth edge maps, point clouds are voxelized and neural Gaussians are derived, combined with joint loss function and multi-layer perceptron to optimize anchor point features and neural Gaussian properties, the problem of limited rendering efficiency and accuracy in complex scenes is solved, and efficient and accurate rendering effects are achieved.
Patent Information
- Application Number
- CN202411991972.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The rendering efficiency and accuracy of the three-dimensional scenes in complex scenes are limited by high-quality initial sparse point clouds, especially when the scene observation angle is limited or the SfM results are poor, it is difficult to achieve efficient and accurate rendering.
A real-time rendering method for three-dimensional scenes based on spatial perception is proposed. Through multi-view image acquisition and initial sparse point cloud generation, scene geometric constraints are extracted in combination with depth maps and edge maps, point clouds are voxelized and neural Gaussians are derived. Joint loss function and multi-layer perceptron are used to train and optimize anchor features and neural Gaussian properties, generate rendered images, and anchor distribution is optimized through cropping and adding new anchor mechanisms.
Reliance on high-quality initial sparse point clouds is reduced, rendering efficiency and accuracy of complex scenes is improved, and the under-reconstructed areas can be effectively covered, and the processing ability of complex geometric shapes and details is enhanced.
Smart Images

Figure CN119941950A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of three-dimensional reconstruction, and in particular to a real-time rendering method for three-dimensional scenes based on spatial perception. Background Art
[0002] With the development of 3D scene rendering technology, 3D Gaussian Splatting (3DGS) has shown excellent rendering fidelity and efficiency. However, the explicit 3DGS representation has very high requirements for computing and memory resources. The main reason is that each Gaussian needs to explicitly store its parameters (such as position, shape, color, transparency, etc.), and needs to be calculated and updated in real time during the rendering process. In addition, in order to accurately fit each perspective and present details, 3DGS often generates a large number of redundant Gaussian distributions. These redundant Gaussians repeatedly represent the same geometric features, greatly increasing memory usage and computational burden. This is especially significant when dealing with complex scenes or high-resolution rendering.
[0003] Scaffold-GS overcomes the above limitations by anchoring neural Gaussians. Scaffold-GS generates anchor points from sparse point clouds, constructs voxel grids, and trains neural Gaussians using multi-layer perceptrons (MLPs), significantly reducing memory consumption. Since the neural Gaussians on the anchor points only need to store a small number of parameters such as their mean and variance, and the calculation mainly depends on the data distribution within the local voxel grid, Scaffold-GS avoids global rendering calculations for the entire scene. Therefore, it can still run with low and stable memory overhead even when dealing with scenes such as complex buildings or large open spaces.
[0004] However, the effectiveness of Scaffold-GS relies heavily on the high-quality sparse point cloud obtained from Structure-from-Motion (SfM). If the SfM result is poor or the scene observation angle is limited, its anchor growth strategy will face challenges, resulting in the reconstruction of details in these areas still being difficult. Summary of the invention
[0005] The purpose of this application is to provide a real-time rendering method for three-dimensional scenes based on spatial perception, so as to reduce the dependence of Scaffold-GS on high-quality initial sparse point clouds, thereby achieving more efficient and accurate rendering in complex scenes.
[0006] To achieve the above purpose, the technical solution of this application is as follows: A real-time rendering method for three-dimensional scenes based on spatial perception, comprising: Step 1: Obtain multi-view images of the scene, estimate the camera pose and generate the initial sparse point cloud of the scene, and extract the depth map and edge map corresponding to the image; Step 2: De-noise the initial sparse point cloud and then extract the spatial boundary of the scene; Step 3: voxelize the denoised point cloud, regard the center of each voxel as an anchor point, derive a neural Gaussian for the anchor point according to the initialized anchor point offset and scaling factor, and use the initialized multi-layer perceptrons to obtain the anchor point features and neural Gaussian properties; Step 4: Randomly select a camera pose, generate a rendered image with the anchor feature and neural Gaussian attribute of the current anchor point, extract the depth map and edge map of the rendered image, calculate the loss and perform back propagation, and train to obtain the anchor offset, scaling factor and weights of each multi-layer perceptron; Step 5: Perform a cropping operation on the original anchor points, generate new anchor points based on spatial perception, merge the new anchor points with the cropped anchor points to obtain an anchor point set, and update the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set; Step 6: Determine whether the iteration termination condition is reached. If not, return to step 4 to continue the iteration. Otherwise, terminate the iteration and generate a rendered image using the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set.
[0007] Furthermore, the denoising of the initial sparse point cloud includes: For each point in the initial sparse point cloud, calculate the average distance between it and all points in the neighborhood; If the average distance exceeds the standard deviation threshold, the corresponding point is considered to be a noise point and is removed.
[0008] Furthermore, extracting the spatial boundary of the scene includes: Compute the convex hull of the point cloud; According to the set scale factor , expand each vertex in the convex hull outward, the expansion formula is as follows: in, are convex hull vertices, is the convex hull vertex after expansion, c is the centroid of the convex hull vertex, and the convex hull formed by the expanded vertices is used as the spatial boundary of the scene.
[0009] Furthermore, the joint loss function used in calculating the loss is as follows: in, is the joint loss function, is the depth loss, is the edge loss, is the original loss function of Scaffold-GS, , and are weight coefficients respectively.
[0010] Furthermore, performing a clipping operation on the original anchor point includes: Record the opacity and number of visits of each anchor point used when rendering the image; After a preset number of iterations, if the number of visits to an anchor point is greater than a minimum visit number threshold and its cumulative opacity is less than a minimum opacity threshold, the anchor point is removed.
[0011] Furthermore, the generating of the new anchor point based on the spatial perception includes: After a preset number of iterations, for each neural Gaussian, the Euclidean distance to each anchor point is calculated. If the minimum Euclidean distance is greater than the distance threshold, the center position of the voxel where the neural Gaussian is located is marked as a candidate anchor point; According to the spatial boundaries of the scene, candidate anchor points that exceed the spatial boundaries are eliminated; For the candidate anchor point set, select the candidate anchor point with the highest gradient value and put it into the newly added anchor point set, and remove the candidate anchor points whose distance to the candidate anchor point with the highest gradient value is less than the distance threshold from the candidate anchor point set; Repeat the previous step until the candidate anchor point set is empty, and obtain the final set of newly added anchor points; For each newly added anchor point in the newly added anchor point set, initialize their local context features, scaling factors, and anchor offsets, where the local context features are inherited from the candidate anchor points, initialize a tensor of all zeros as the anchor offset, and initialize a tensor of all 1s, multiply it by the voxel size of the current scene, and then take the logarithm to get the scaling factor.
[0012] Furthermore, each of the multilayer perceptrons includes a first multilayer perceptron for generating anchor point features and second to fourth multilayer perceptrons for predicting neural Gaussian attributes, and the updating of the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set includes: For the newly added anchor points, a feature library is created, and the view weights predicted by the first multi-layer perceptron after training are used to mix the feature library to obtain the anchor point features. Then, the scaling factor and offset are used to derive the neural Gaussian, and the trained second to fourth multi-layer perceptrons are used for prediction to obtain the neural Gaussian attributes.
[0013] Furthermore, each of the multilayer perceptrons includes a first multilayer perceptron for generating anchor point features and second to fourth multilayer perceptrons for predicting neural Gaussian attributes, and the updating of the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set includes: For the cropped anchor points, the first multi-layer perceptron after training is used to update the anchor point features; the trained scaling factor and offset are used to update the neural Gaussian position, and the trained second to fourth multi-layer perceptrons are used for prediction to update the neural Gaussian attributes.
[0014] The present invention proposes a real-time rendering method for three-dimensional scenes based on spatial perception, which has the following advantages compared with the prior art: The scene geometric constraints are extracted to ensure that anchor growth only occurs within a reasonable spatial range, thereby avoiding the generation of redundant anchors. A spatially aware anchor growth method is proposed, which enables anchors to effectively cover under-reconstructed areas, thereby significantly reducing the reliance of Scaffold-GS on high-quality initial point clouds. A joint loss function is designed, which adds depth loss and edge loss to the original Scaffold-GS loss function, further optimizing the prediction of neural Gaussian related attributes, while ensuring the quality of reconstruction and enhancing the ability to handle complex geometric shapes and details. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of the real-time rendering method for three-dimensional scenes based on spatial perception in this application.
[0016] Figure 2 This is the anchor point growth flow chart of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0018] An embodiment of the present application, such as Figure 1 As shown, a real-time rendering method for three-dimensional scenes based on spatial perception is proposed, including: Step S1, obtain multi-view images of the scene, estimate the camera pose and generate an initial sparse point cloud of the scene, and extract the depth map and edge map corresponding to the image.
[0019] Structure from Motion (SfM) is a technique for reconstructing three-dimensional scenes from image sequences.
[0020] In this embodiment, the scene refers to the scene that needs to be rendered using the method. The camera captures images of various perspectives in the scene in different postures as real images. Image RGB pixel values, depth maps and edge maps can be obtained from the real images.
[0021] This embodiment uses the SfM method to process the real image, estimate the camera pose, and generate an initial sparse point cloud of the scene. This method generates an initial sparse point cloud by matching feature points in multi-view images to extract preliminary geometric information of the scene.
[0022] This embodiment uses a monocular depth estimator (Depth Anything V2) and the kornia.filters.Canny algorithm in the Kornia library to extract the depth map and edge map corresponding to each image.
[0023] Step S2: denoise the initial sparse point cloud and then extract the spatial boundary of the scene.
[0024] This step extracts geometric constraints, performs geometric analysis on the generated initial sparse point cloud, and defines the spatial boundaries of the scene.
[0025] Specifically include: Step 2.1: Remove isolated noise points from the initial sparse point cloud.
[0026] The statistical outlier removal (SOR) method is used to identify and remove isolated noise points in the initial sparse point cloud. The specific steps are as follows:
[0027] For the initial sparse point cloud , for each point , calculate its relationship with the neighborhood The average distance between all points in : If the average distance Exceeds the set standard deviation threshold , then it is considered that point are noise points and remove them.
[0028] Step 2.2: Extract the spatial boundaries of the scene.
[0029] The ConvexHull algorithm is used to calculate the convex hull H of the point cloud. Specifically, the convex hull H is the smallest polyhedron that contains all the points of the point cloud. The following conditions are met: Among them, ConvexHull is the algorithm identifier, m represents the number of vertices, represents a vertex in the vertex set U, Represents a vertex Weights in the convex hull.
[0030] On this basis, in order to ensure that the subsequent anchor points can cover the edge of the scene and its near-edge area, the convex hull vertices need to be expanded. Let the centroid of the convex hull vertex be , then each vertex is transformed into Expand outward by a scaling factor :
[0031] in, It is the convex hull vertices after expansion. The convex hull formed by the expanded vertices is used as the spatial boundary of the scene to provide geometric constraints for the distribution of subsequent anchor points.
[0032] Step S3: voxelize the denoised point cloud, regard the center of each voxel as an anchor point, derive a neural Gaussian for the anchor point according to the initialized anchor point offset and scaling factor, and use the initialized multi-layer perceptrons to obtain the anchor point features and neural Gaussian properties.
[0033] This step takes the denoised point cloud as input and converts the point cloud into The scene voxelization is: in is the voxel center, is the voxel size, Represents rounding operation, Represents the removal of duplicate entries to reduce redundancy and irregularity in P.
[0034] The center of each voxel is regarded as an anchor point with a local context feature , a scaling factor , k anchor point offsets (Obtained by initialization). The local context feature is used to capture the geometric and appearance information of the environment around the anchor point, the scaling factor is used to adjust the size of the neural Gaussian associated with the anchor point, and the anchor point offset is used to describe the position of the neural Gaussian associated with the anchor point relative to the anchor point in space.
[0035] To further enhance the contextual features Dependency on multiple resolutions and views, for each anchor point v, create a feature library ,in express In the channel dimension Downsampling factor.
[0036] Then use the first multi-layer perceptron ( ) The predicted view weight is used to mix the feature library to obtain the anchor feature .
[0037] Specifically, given a position A camera and a Anchor points, calculate their relative distance and view direction: Then use Prediction weight Perform weighted summation with the feature library to obtain the anchor feature : .
[0038] Next, a neural Gaussian is derived for the anchor point, and the neural Gaussian properties are obtained through the initialized second to fourth multi-layer perceptrons.
[0039] Neural Gaussian is an efficient 3D representation unit that combines the mathematical properties of Gaussian distribution and the learning ability of neural networks to accurately represent complex 3D scenes. , Opacity , covariance-related quaternions and color to parameterize each neural Gaussian.
[0040] For each anchor point in the scene, k neural Gaussians are derived and their properties are obtained.
[0041] Specifically, given the The anchor point of the neural Gaussian is calculated as: in is the anchor point offset, is the scaling factor associated with this anchor point. Respectively represent the position of each neural Gaussian. In this step, the anchor offset is initialized. Specifically, for each anchor, the position of its neural Gaussian is calculated as the anchor position plus the offset multiplied by the scaling factor.
[0042] The opacity, color, and covariance related quaternions of the neural Gaussian are respectively initialized through the second to fourth multi-layer perceptrons MLP ( , and ) to make predictions.
[0043] It should be noted that the attribute is decoded once, opacity ,color , covariance-related quaternions It is predicted by the following formula.
[0044] It should be noted that the multi-layer perceptron MLP is used to predict the perspective weight and the properties of the neural Gaussian. This step uses the initialized multi-layer perceptron, which will be trained later to obtain a multi-layer perceptron that can accurately predict. The initialization of the multi-layer perceptron, anchor offset, and scaling factor is a relatively mature technology in this field and will not be repeated here. The weights of these MLPs are optimized during the training process to dynamically predict the properties of the neural Gaussian.
[0045] Step S4: randomly select a camera pose, generate a rendered image with the anchor point features and neural Gaussian properties of the current anchor point, extract the depth map and edge map of the rendered image, calculate the loss and perform back propagation, and train to obtain the anchor point offset, scaling factor, and weights of each multi-layer perceptron.
[0046] This step starts iterative training of each multi-layer perceptron, and first performs random view image rendering. That is, randomly select a camera pose obtained from step S1, generate a rendered image through the Scaffold-GS method based on the current anchor point features and neural Gaussian properties, and then obtain the depth map of the rendered image through the monocular depth estimator, and extract the edge map of the rendered image by calling the kornia.filters.Canny algorithm in the Kornia library.
[0047] Then, the loss and backpropagation are calculated to update the parameters of each multilayer perceptron.
[0048] This embodiment calculates the loss value by using the joint loss function. It can be expressed as: in, , and are weight coefficients used to balance the contribution of depth loss, edge loss and other loss terms. It is calculated by the real image depth map obtained in step S1 and the rendered image depth map. The distance measures the consistency between two depth maps. Edge loss The calculation method is the same as the depth loss, using The distance measures the consistency between two edge maps. is the original loss function of Scaffold-GS, including the rendered pixel color loss, structural similarity loss ( ) and the volume regularization loss ( ):
[0049] Next, we enter the back-propagation phase, where the gradients of each MLP weight, anchor offset, and scaling factor in the model are calculated based on the value of the loss function. Subsequently, the Adam optimizer for each parameter fine-tunes the parameters based on these gradient values and pre-set hyperparameters such as the learning rate. The update of the MLP weights can more accurately predict the relevant properties of the neural Gaussian. The adjustment of the anchor offset enables the neural Gaussian to more accurately locate the key feature points in the scene. The optimization of the scaling factor can more accurately control the diffusion range of the neural Gaussian, thereby capturing the scene details more meticulously. These updates work together to make the rendered scene details finer and the loss value lower in the next iteration, ultimately achieving a significant improvement in rendering quality.
[0050] Regarding backpropagation through loss, updating the weights of each MLP according to the loss value, and using the Adam optimizer to update the anchor offset and scaling factor, it is a relatively mature technology in this field and will not be repeated here.
[0051] Step S5: perform a cropping operation on the original anchor points, generate new anchor points based on spatial perception, merge the new anchor points with the cropped anchor points to obtain an anchor point set, and update the anchor point features and neural Gaussian properties of each anchor point in the anchor point set.
[0052] Specifically, firstly, a cropping operation is performed on the original anchor point, including: Step 5.1.1. Record the opacity and access count of each anchor point used when rendering the image.
[0053] For each anchor point involved in rendering, sum up their associated neural Gaussian opacities and add them to the cumulative opacity of these anchor points. At the same time, increase the number of visits to these anchor points by 1.
[0054] Step 5.1.2: After a preset number of iterations, if the number of visits to an anchor point is greater than a minimum visit number threshold and its cumulative opacity is less than a minimum opacity threshold, the anchor point is removed.
[0055] Anchor point cropping is performed after every N iterations. The cumulative opacity and access times of each anchor point in the previous N iterations are counted. If the access times of an anchor point are greater than the minimum access times threshold and its cumulative opacity is less than the minimum opacity threshold, the anchor point is removed to eliminate active anchor points that contribute little to the scene. Afterwards, the cumulative opacity and access times of the anchor points counted are reset to 0.
[0056] It should be noted that, in the iteration process before reaching N iterations, no cropping operation is performed, and all current anchor points are retained for the next iteration.
[0057] Then, if Figure 2 As shown in the figure, new anchor points are generated based on spatial perception, including: Step 5.2.1. After a preset number of iterations, for each neural Gaussian, calculate the Euclidean distance to each anchor point. If the minimum Euclidean distance is greater than the distance threshold, mark the center position of the voxel where the neural Gaussian is located as a candidate anchor point.
[0058] This step is used to generate candidate anchor points, and a distance threshold D is defined as the minimum distance between candidate anchor points. In each anchor point growth process, the anchor point set in the current scene is defined as , the neural Gaussian set is defined as Then iterate over the neural Gaussians in the scene , calculate the nearest anchor point in the current anchor point set A The Euclidean distance If it satisfies , then the neural Gaussian The center position of the voxel is marked as a candidate anchor point to ensure that the candidate anchor point can cover the area where the anchor points are insufficiently distributed.
[0059] Step 5.2.2: According to the spatial boundary of the scene, remove the candidate anchor points that exceed the spatial boundary.
[0060] The set of candidate anchor points is defined as , and then filter the candidate anchor points according to the scene space boundary, and remove the candidate anchor points that exceed the boundary range.
[0061] Step 5.2.3: For the candidate anchor point set, select the candidate anchor point with the highest gradient value and put it into the newly added anchor point set, and remove the candidate anchor points whose distance to the candidate anchor point with the highest gradient value is less than the distance threshold from the candidate anchor point set.
[0062] Then, a gradient-first greedy algorithm is used to control the density of candidate anchor points. Specifically, each candidate anchor point in the candidate anchor point set C , both have gradient values In the subsequent screening process, the candidate anchor point with the highest gradient value is selected , and move it into the new set S, namely:
[0063] Then, with the currently selected As the center, all the Candidate anchor points whose distance is less than the threshold D are removed to control the density of anchor points, which is specifically expressed as: .
[0064] Step 5.2.4: Repeat the previous step until the candidate anchor point set is empty, and obtain the final set of newly added anchor points.
[0065] Repeat the above steps until the set C is empty. The final set S is the final set of newly added anchor points after density control. For the newly added anchor point set S, for each newly added anchor point, initialize their local context features, scaling factors and anchor offsets, where the local context features Inherited from the candidate anchor point, initialize a tensor of all zeros as the anchor point offset, and initialize a tensor of all 1s, multiply it by the voxel size of the current scene, and then take the logarithm to get the scaling factor.
[0066] The cropped anchor points and the newly added anchor points are merged into an anchor point set, and the anchor point features and neural Gaussian properties of each anchor point in the anchor point set are updated for the next iteration.
[0067] Among them, for the newly added anchor points, a feature library is created to generate anchor point features, and a neural Gaussian is derived for the newly added anchor points. The trained multi-layer perceptron is used to predict the neural Gaussian properties, thus ensuring that the new anchor points can effectively express the geometric appearance characteristics of the scene.
[0068] Specifically, create a feature library , using the first multilayer perceptron after training ( ) The predicted view weight is used to mix the feature library to obtain the anchor feature , and then use the scaling factor and offset to derive the neural Gaussian, and use the trained second to fourth multi-layer perceptrons MLP ( , and ) to make predictions and obtain the neural Gaussian attributes.
[0069] For the existing anchor points after cropping, the first multi-layer perceptron after training ( ), update the anchor feature ; Use the trained scaling factor and offset to update the neural Gaussian position, and use the trained second to fourth multi-layer perceptrons MLP ( , and ) to make predictions and update the neural Gaussian attributes.
[0070] Step S6, determine whether the iteration termination condition is met, if not, return to step 4 to continue iteration, otherwise terminate the iteration, and generate a rendered image with the anchor point features and neural Gaussian properties of each anchor point in the anchor point set.
[0071] This step determines whether the iteration termination condition is reached. For example, the iteration termination condition is to take the maximum number of iterations as the iteration termination condition. The iteration ends after the maximum number of iterations is reached. If it is not reached, the process returns to step S4 to continue the iteration.
[0072] After the iteration, the scene is rendered. A random camera pose is input, and the anchor point within the camera coverage is determined. The spatial position of the neural Gaussian is determined based on the offset of the neural Gaussian relative to the anchor point. The relative distance and view direction between the camera and the anchor point are calculated. These and the anchor point features are input into the multi-layer perceptron to infer the color, opacity, and covariance of the neural Gaussian. Subsequently, the contribution of the neural Gaussian is weighted and accumulated in order of depth, and the color of each pixel is calculated to achieve real-time rendering of the three-dimensional scene.
[0073] It should be noted that how to use the Scaffold-GS method to generate a rendered image is a relatively mature technology in this field and will not be described in detail here.
[0074] In order to verify the effectiveness of the technical solution, training and testing were performed on three public datasets: Mip-NeRF360, Tanks&Temples, and Deep Blending. Similar to the 3DGS method, for the input images of each scene, one eighth of the total number of images is used as a test set, and the remaining images are used as training sets. In order to objectively evaluate the visual fidelity, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learning-perceptual image block similarity (LPIPS) are used as evaluation indicators. These indicators can compare the differences between images rendered by different methods and the corresponding real frames, thereby quantitatively analyzing the rendering effect. The specific experimental results are shown in Table 1:
[0075] Table 1
[0076] As can be seen from Table 1, the rendering quality of the present application in the above three datasets is better than other methods. As evaluation criteria, the larger the values of PSNR and SSIM, the higher the image similarity, while the smaller the value of LPIPS, the smaller the visual difference between the rendered image and the real image. Further analyzing the details of the rendering results, the present application is able to retain fine geometric details in complex scenes. For example: in the STUMP scene of the Mip-NeRF360 dataset, small branches are accurately rendered; in the PLAYROOM scene of the Deep Blending dataset, the switches on the wall are carefully displayed; in the TRAIN scene of the Tanks&Temples dataset, the details of the truck license plate number and windshield reflection are well reproduced.
[0077] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A real-time rendering method for three-dimensional scenes based on spatial perception, characterized in that: The real-time rendering method for three-dimensional scenes based on spatial perception includes: Step 1: Obtain multi-view images of the scene, estimate the camera pose and generate the initial sparse point cloud of the scene, and extract the depth map and edge map corresponding to the image; Step 2: De-noise the initial sparse point cloud and then extract the spatial boundary of the scene; Step 3: voxelize the denoised point cloud, regard the center of each voxel as an anchor point, derive a neural Gaussian for the anchor point according to the initialized anchor point offset and scaling factor, and use the initialized multi-layer perceptrons to obtain the anchor point features and neural Gaussian properties; Step 4: Randomly select a camera pose, generate a rendered image with the anchor feature and neural Gaussian attribute of the current anchor point, extract the depth map and edge map of the rendered image, calculate the loss and perform back propagation, and train to obtain the anchor offset, scaling factor and weights of each multi-layer perceptron; Step 5: Perform a cropping operation on the original anchor points, generate new anchor points based on spatial perception, merge the new anchor points with the cropped anchor points to obtain an anchor point set, and update the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set; Step 6: Determine whether the iteration termination condition is reached. If not, return to step 4 to continue the iteration. Otherwise, terminate the iteration and generate a rendered image using the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set.
2. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The denoising of the initial sparse point cloud comprises: For each point in the initial sparse point cloud, calculate the average distance between it and all points in the neighborhood; If the average distance exceeds the standard deviation threshold, the corresponding point is considered to be a noise point and is removed.
3. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The extracting the spatial boundary of the scene includes: Compute the convex hull of the point cloud; According to the set scale factor , expand each vertex in the convex hull outward, the expansion formula is as follows: in, are convex hull vertices, is the convex hull vertex after expansion, c is the centroid of the convex hull vertex, and the convex hull formed by the expanded vertices is used as the spatial boundary of the scene.
4. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The combined loss function used in the calculation of loss is as follows: in, is the joint loss function, is the depth loss, is the edge loss, is the original loss function of Scaffold-GS, , and are weight coefficients respectively.
5. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The performing a clipping operation on the original anchor point includes: Record the opacity and visit count of each anchor point used when rendering the image; After a preset number of iterations, if the number of visits to an anchor point is greater than a minimum visit number threshold and its cumulative opacity is less than a minimum opacity threshold, the anchor point is removed.
6. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The generating of the new anchor point based on the spatial perception includes: After a preset number of iterations, for each neural Gaussian, the Euclidean distance to each anchor point is calculated. If the minimum Euclidean distance is greater than the distance threshold, the center position of the voxel where the neural Gaussian is located is marked as a candidate anchor point; According to the spatial boundaries of the scene, candidate anchor points that exceed the spatial boundaries are eliminated; For the candidate anchor point set, select the candidate anchor point with the highest gradient value and put it into the newly added anchor point set, and remove the candidate anchor points whose distance to the candidate anchor point with the highest gradient value is less than the distance threshold from the candidate anchor point set; Repeat the previous step until the candidate anchor point set is empty, and obtain the final set of newly added anchor points; For each newly added anchor point in the newly added anchor point set, initialize their local context features, scaling factors, and anchor offsets, where the local context features are inherited from the candidate anchor points, initialize a tensor of all zeros as the anchor offset, and initialize a tensor of all 1s, multiply it by the voxel size of the current scene, and then take the logarithm to get the scaling factor.
7. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 6, characterized in that: The multi-layer perceptrons include a first multi-layer perceptron for generating anchor point features and second to fourth multi-layer perceptrons for predicting neural Gaussian attributes, and the updating of the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set includes: For the newly added anchor points, a feature library is created, and the view weights predicted by the first multi-layer perceptron after training are used to mix the feature library to obtain the anchor point features. Then, the scaling factor and offset are used to derive the neural Gaussian, and the trained second to fourth multi-layer perceptrons are used for prediction to obtain the neural Gaussian attributes.
8. The real-time rendering method for three-dimensional scenes based on spatial perception according to claim 1, characterized in that: The multi-layer perceptrons include a first multi-layer perceptron for generating anchor point features and second to fourth multi-layer perceptrons for predicting neural Gaussian attributes, and the updating of the anchor point features and neural Gaussian attributes of each anchor point in the anchor point set includes: For the cropped anchor points, the first multi-layer perceptron after training is used to update the anchor point features; the trained scaling factor and offset are used to update the neural Gaussian position, and the trained second to fourth multi-layer perceptrons are used for prediction to update the neural Gaussian attributes.
Citation Information
Patent Citations
A multi-core processor supporting real-time 3D image rendering on an autostereoscopic display
CN102835119A
Quick image matching method fusing point-line characteristics
CN109993747A
Nerve radiation field three-dimensional reconstruction method and device based on adaptive mask
CN117934710A
3D modeling reconstruction system, method and device based on point cloud information and Gaussian cloud cluster
CN118196306A
Three-dimensional volume video coding compression method
CN119052510A
Cited By
Underwater scene 3D characterization method based on three-dimensional Gaussian splashing
CN120219664A
Interactive three-dimensional teaching resource generation system and method based on 3D Gaussian splashing
CN121170164A
Rapid 3DGS three-dimensional reconstruction method based on graph neural network anchor point cooperation
CN121639922A