Static scene three-dimensional reconstruction method and system based on adaptive dynamic object elimination
Through adaptive dynamic threshold mechanism and residual statistical analysis, combined with segmentation large model to identify and eliminate dynamic objects, the problem of insufficient accuracy of static scene reconstruction in the existing technology is solved, and high-quality three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202510764654.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When handling dynamic object interference, the prior art cannot achieve accurate identification and removal in any scenario, resulting in insufficient three-dimensional reconstruction accuracy of static scenes.
Adaptive dynamic threshold mechanism is adopted, combined with segmentation large model and residual statistical analysis, and interfering objects are identified and eliminated through dynamic threshold judgment function, and a 3D Gaussian splattering radiation field loss function is constructed to achieve iterative optimization.
It significantly improves the recognition accuracy of dynamic interference objects in complex scenarios, reduces static scene reconstruction errors, and improves the quality of three-dimensional reconstruction.
Smart Images

Figure CN120279196A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual three-dimensional reconstruction, and particularly relates to a method and system for three-dimensional reconstruction of static scenes based on adaptive dynamic object removal. Background Art
[0002] Three-dimensional reconstruction is an important part of the field of computer vision and has important application values in fields such as robot perception and virtual reality technology. In recent years, three-dimensional reconstruction methods based on radiation fields have developed rapidly, and the three-dimensional reconstruction method based on 3D Gaussian splatting has shown great advantages in high-fidelity reconstruction and fast rendering. The current mainstream 3D Gaussian splatting method assumes that the scene is static and optimizes the parameters of Gaussian basis elements through multi-view RGB losses. However, when collecting data in a real environment, various dynamic objects such as pedestrians and vehicles will inevitably be captured. These dynamic interferences break the multi-view consistency assumed by 3D Gaussian splatting, causing the model to gradually deviate from the real static scene during training, and thus resulting in reconstruction errors.
[0003] Existing methods for solving the influence of dynamic objects on static scene reconstruction can be divided into two categories. The first category uses an instance segmentation model to obtain masks of common dynamic objects and sets the gradients of the masked regions to zero during gradient backpropagation. However, the predefined categories often cannot comprehensively cover dynamic objects in real complex scenes, resulting in reconstruction failures in some regions. The second category follows the idea that "dynamics are more difficult to fit than static", starting from the reconstruction photometric residuals between the rendered image and the real image, determines the regions where the residuals are higher than a certain threshold as the regions where dynamic objects are located, and sets the gradients of these regions to zero during backpropagation. However, the fixed threshold cannot accurately identify different interferences in different scenes, and misjudgments will occur, further leading to reconstruction failures in some regions. Therefore, designing an adaptive recognition method for dynamic interfering objects to achieve accurate recognition in any scene and thus high-quality three-dimensional reconstruction of static scenes has important practical application values.
[0004] For the Chinese patent with the patent application number CN202411248028.7 and the invention title "A Method for Three-Dimensional Reconstruction and Rendering Based on Outdoor Unconstrained Image Sets", this method first uses a pre-trained semantic segmentation model to obtain a semantic segmentation map and uses an MLP to construct an implicit function to train and obtain a visibility map representing static and transient visibility. Although this method can separate the static and transient phenomena of the image, it still cannot achieve robust recognition of dynamic objects in any scene. In addition, this method constructs an implicit function to obtain the transient visibility map, which will result in inaccurate prediction of the edges of dynamic objects, thus reducing the reconstruction quality. Summary of the Invention
[0005] The present invention provides a method and system for three-dimensional reconstruction of static scenes based on adaptive dynamic object removal, aiming to solve the problem of insufficient accuracy in static scene reconstruction caused by dynamic object interference in the prior art. The core of the present invention lies in adaptively identifying and removing interfering objects through a dynamic threshold mechanism, and combining a segmentation large model with a residual statistical analysis method to achieve robust three-dimensional reconstruction of static scenes.
[0006] A method for three-dimensional reconstruction of static scenes, comprising: The first step: segment all images in the scene image set to obtain a set of two-dimensional segmentation masks for each image; The second step: calculate the camera internal parameters and camera poses of all images in the scene image set, and at the same time obtain the initial sparse point cloud of the scene; initialize the 3D Gaussian splatter radiation field based on the initial sparse point cloud of the scene, and train the 3D Gaussian splatter radiation field in combination with the scene images and corresponding pose information, establish the image pixel-level normalization residual between the rendered image and the original image, and calculate the object-level normalization residual according to the image pixel-level normalization residual and the two-dimensional segmentation mask; The third step: calculate the residual mean and variance of the image according to the object-level normalization residual; The fourth step: construct a dynamic object threshold judgment function based on the residual mean and variance, and generate a set of dynamic object masks according to the threshold judgment function; The fifth step: define a 3D Gaussian splatter radiation field loss function in combination with the set of dynamic object masks, and iteratively optimize and train the 3D Gaussian splatter radiation field until a preset stop condition is reached to complete the three-dimensional reconstruction of the static scene.
[0007] Preferably, in the first step, the method for obtaining the set of two-dimensional segmentation masks for each image includes: For each image in the scene image set , use the segmentation large model to obtain the segmentation result of the image , and the specific expression form is:
[0008] where represents the segmentation mask of the th object in the segmentation result of the image , and represents the number of object segmentation masks in the segmentation result of the image .
[0009] Preferably, in the second step, the calculation formula for the image pixel-level normalization residual is:
[0010] Among them, is the normalized L1 residual between the 3D Gaussian splash radiation field rendering image and the original image, is the normalized D-SSIM residual between the 3D Gaussian splash radiation field rendering image and the original image, is the set weight coefficient.
[0011] Preferably, in the second step, the object-level normalized residual is:
[0012] Among them, represents the normalized residual of the th object in the image segmentation result, is the th object segmentation mask in the image segmentation result, is the pixel-level normalized residual of the image at training iteration represents is a pixel included in the segmentation mask , is the segmentation mask total number of pixels included.
[0013] Preferably, in the third step, the residual mean and variance of the image are respectively:
[0014]
[0015] Among them, represents the number of pixels included in the current training image, is the th normalized residual of the pixel in the image represents the th pixel probability, when all pixel weights are the same
[0016] Preferably, in the fourth step, the calculation formula of the dynamic object threshold judgment function is:
[0017] Among them, is the mean of the object-level normalized residuals of the image at training iteration , is the image The variance of the object-level normalized residual, is the maximum number of training iterations, and is the set weight coefficient.
[0018] Preferably, in the fourth step, the method for generating the dynamic object mask set includes: Taking the masks in the two-dimensional segmentation mask set whose object normalized residuals are greater than the dynamic object threshold judgment function to form the dynamic object mask set . Preferably, in the fifth step, the 3D Gaussian splash radiation field loss function is:
[0019] where, is the normalized L1 loss between the 3D Gaussian splash radiation field rendered image and the original image, is the normalized D-SSIM loss between the 3D Gaussian splash radiation field rendered image and the original image.
[0020] A static scene three-dimensional reconstruction system includes a first module, a second module, a third module, a fourth module, and a fifth module; The first module is used to execute the first step: segment all the images in the scene image set to obtain the two-dimensional segmentation mask set of each image; The second module is used to execute the second step: calculate the camera internal parameters and camera poses of all the images in the scene image set, and at the same time obtain the initial sparse point cloud of the scene; initialize the 3D Gaussian splash radiation field based on the initial sparse point cloud of the scene, train the 3D Gaussian splash radiation field in combination with the scene images and the corresponding pose information, establish the image pixel-level normalized residual between the rendered image and the original image, and calculate the object-level normalized residual according to the image pixel-level normalized residual and the two-dimensional segmentation mask; The third module is used to execute the third step: calculate the residual mean and variance of the image according to the object-level normalized residual; The fourth module is used to execute the fourth step: construct a dynamic object threshold judgment function based on the residual mean and variance, and generate a dynamic object mask set according to the threshold judgment function; The fifth module is used to execute the fifth step: define the 3D Gaussian splash radiation field loss function in combination with the dynamic object mask set, and complete the three-dimensional reconstruction of the static scene by iteratively optimizing and training the 3D Gaussian splash radiation field until the preset stop condition is reached.
[0021] The present invention has the following beneficial effects: By integrating the generalization ability of the segmentation large model and residual statistical analysis, the present invention constructs an adaptive dynamic threshold judgment function, breaking through the limitations of traditional fixed thresholds or predefined categories. The dynamic threshold judgment function combines the residual mean and variance, retaining potential static regions in the initial stage to avoid misjudgment, and gradually tightening the threshold in the later stage to accurately eliminate dynamic objects, significantly improving the recognition accuracy of dynamic interference objects in complex scenes and reducing the static scene reconstruction error. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart of an embodiment of the method of the present invention; FIG. 2(a) and FIG. 2(b) are respectively a scene image and a 3D Gaussian splash radiation field rendering image; FIG. 3(a) and FIG. 3(b) are respectively an image object-level normalized residual map and a dynamic object mask map. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.
[0024] A three-dimensional reconstruction method for static scenes based on adaptive dynamic object removal includes the following steps: First step: Use the segmentation large model to obtain the two-dimensional segmentation masks of all images in the current scene; Step 101, for each image in the scene image set , use the segmentation large model to obtain the segmentation result of the image , and the specific expression form is:
[0025] wherein, represents the segmentation mask of the th object in the segmentation result of the image , and represents that the segmentation result of the image contains a total of segmentation masks of objects.
[0026] Second step: Initialize the 3D Gaussian splash radiation field according to the initial sparse point cloud of the scene, train the 3D Gaussian splash radiation field according to the scene image and the corresponding pose information, establish a reconstruction residual function between the 3D Gaussian splash radiation field rendering image and the scene image, and calculate the object-level normalized residual of the training image in combination with the reconstruction residual function and the two-dimensional segmentation mask; Step 201, use the sparse point cloud obtained by the multi-view stereometry method according to the scene image to initialize the 3D Gaussian splash radiation field, denoted as ; Step 202: Train a 3D Gaussian splatter radiation field based on the scene image set and its corresponding pose information set, and establish a pixel-level normalization residual between the neural radiation field rendered image and the scene image . The calculation formula is as follows:
[0027] where is the normalized L1 residual between the 3D Gaussian splatter radiation field rendered image and the original image, is the normalized D-SSIM residual between the 3D Gaussian splatter radiation field rendered image and the original image, is the weight coefficient; Step 203: Calculate the object-level normalization residual under the training iteration by taking the reconstruction residual in Step 202 and the segmentation result in Step 101 . The calculation formula is as follows:
[0028] where is the segmentation mask of the th object in the image segmentation result, is the pixel-level normalization residual of the image under the training iteration , represents is a pixel included in the mask , is the mask the total number of pixels included.
[0029] Third step: Calculate the mean and variance of the object-level normalization residual of the image; Step 301: Take the object-level normalization residual of the image obtained in Step 203 and calculate the mean of the object-level normalization residual of the image under the training iteration . The calculation formula is as follows:
[0030] where represents the number of pixels included in the current training image, is the normalized residual of the th pixel in the image , represents the th pixel probability, when all pixel weights are the same .
[0031] Step 302: Take the mean of the image object-level normalized residuals obtained in Step 203 and the image object-level normalized residuals obtained in Step 301, and calculate the training iteration lower image Variance of object-level normalized residuals , and the calculation formula is:
[0032] where represents the number of pixels included in the current training image, is the training iteration lower image in the th pixel's normalized residual, is the training iteration lower image mean of object-level normalized residuals; Fourth step: Establish a dynamic object threshold judgment function based on the mean and variance of the image object-level normalized residuals , and establish a dynamic object mask set based on the image object-level normalized residuals and the dynamic object threshold judgment function ; Step 401: Take the mean of the image object-level normalized residuals obtained in Step 301 and the variance of the image object-level normalized residuals obtained in Step 302, and calculate the dynamic object threshold judgment function for the training iteration lower image , and the calculation formula is:
[0033] where is the training iteration lower image mean of object-level normalized residuals, is the training iteration lower image variance of object-level normalized residuals, is the maximum training iteration, is the weight coefficient; Step 402: Based on the object-level normalized residuals of the training iteration lower image and the dynamic object threshold judgment function, establish the dynamic object mask set of the training iteration lower image , and the calculation formula is:
[0034] where is the image The segmentation mask of the th object in the segmentation result, for the image at training iteration is the normalized residual of the th object in the image. Equation (7) represents taking the masks in the set of two-dimensional segmentation masks of all objects whose normalized residuals are greater than the dynamic object threshold judgment function to form the dynamic object mask set . , For the image at training iteration , is the dynamic object threshold judgment function, represents the number of object masks contained in the current training image.
[0035] Step 5: Combine the dynamic object mask set to establish the 3D Gaussian splash radiation field loss function and set the training stop condition; In step 501, take the dynamic object mask set obtained in step 402 and establish the 3D Gaussian splash radiation field loss function for backpropagation gradient , and the calculation formula is:
[0036] where is the normalized L1 loss between the 3D Gaussian splash radiation field rendered image and the original image, is the normalized D-SSIM loss between the 3D Gaussian splash radiation field rendered image and the original image, is the weight coefficient, is the dynamic object mask set of the image corresponding to the current iteration at this time; In step 502, return to the second step and repeat all the previous steps to train the 3D Gaussian splash radiation field, set the maximum number of training iterations . When the training iteration reaches the maximum number of training iterations , stop training; obtain the 3D Gaussian splash radiation field and realize three-dimensional reconstruction.
[0037] The present invention also provides a static scene three-dimensional reconstruction system based on adaptive dynamic object elimination, including a first module, a second module, a third module, a fourth module, and a fifth module.
[0038] The first module is used to execute step 1: use a segmentation large model to segment all images in the scene image set to obtain the set of two-dimensional segmentation masks of each image; The second module is used to perform the second step: calculating the camera intrinsics and camera poses of all images in the scene image set, and simultaneously obtaining the initial sparse point cloud of the scene; initializing the 3D Gaussian splatting radiance field based on the initial sparse point cloud of the scene, training the 3D Gaussian splatting radiance field by combining the scene images and the corresponding pose information, establishing the pixel-level normalization residuals between the rendered images and the original images, and calculating the object-level normalization residuals according to the pixel-level normalization residuals and the two-dimensional segmentation masks; The third module is used to perform the third step: calculating the residual mean and variance of each training image according to the object-level normalization residuals; The fourth module is used to perform the fourth step: constructing a dynamic object threshold judgment function based on the residual mean and variance, and generating a set of dynamic object masks according to the judgment function; The fifth module is used to perform the fifth step: defining the 3D Gaussian splatting radiance field loss function by combining the set of dynamic object masks, and completing the three-dimensional reconstruction of the static scene through iterative optimization training until a preset stopping condition is reached.
[0039] In summary, the above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A three-dimensional reconstruction method for static scenes, characterized in that, Including: The first step: Segment all the images in the scene image set to obtain a set of two-dimensional segmentation masks for each image; The second step: Calculate the camera intrinsics, camera poses of all the images in the scene image set, and at the same time obtain the initial sparse point cloud of the scene; Initialize the 3D Gaussian splash radiation field based on the initial sparse point cloud of the scene, train the 3D Gaussian splash radiation field by combining the scene images and the corresponding pose information, establish the image pixel-level normalized residual between the rendered image and the original image, and calculate the object-level normalized residual according to the image pixel-level normalized residual and the two-dimensional segmentation mask; The third step: Calculate the residual mean and variance of the image according to the object-level normalized residual; The fourth step: Based on the residual mean and variance, construct a dynamic object threshold judgment function, and generate a set of dynamic object masks according to the threshold judgment function; The fifth step: Define the 3D Gaussian splash radiation field loss function by combining the set of dynamic object masks, and complete the three-dimensional reconstruction of the static scene by iteratively optimizing and training the 3D Gaussian splash radiation field until a preset stop condition is reached.
2. The three-dimensional reconstruction method for a static scene according to claim 1, wherein, In the first step, the method for obtaining a set of two-dimensional segmentation masks for each image includes: For the scene image set For each image in it, use the segmentation large model to obtain the segmentation result of the image The specific expression form is as follows: Among them, represents the segmentation mask of the n-th object in the image segmentation result, and represents the number of object segmentation masks in the image segmentation result.
3. The three-dimensional reconstruction method for a static scene according to claim 2, wherein In the second step, the calculation formula for the image pixel-level normalized residual is: Among them, is the normalized L1 residual between the 3D Gaussian splash radiation field rendering image and the original image, is the normalized D-SSIM residual between the 3D Gaussian splash radiation field rendering image and the original image, is the set weight coefficient.
4. The three-dimensional reconstruction method for a static scene according to claim 3, wherein In the second step, the object-level normalized residual is: Among them, represents the normalized residual of the th object in the image segmentation result, is the segmentation mask of the th object in the image segmentation result, is the pixel-level normalized residual of the image under the training iteration , represents is a pixel included in the segmentation mask , is the total number of pixels included in the segmentation mask .
5. The three-dimensional reconstruction method for a static scene according to claim 4, characterized in that In the third step described above, the residual mean of the image and variance are respectively: Among them, represents the number of pixels included in the current training image, is the image in the normalized residual of the -th pixel, represents the probability of the -th pixel, when all pixel weights are the same . .
6. The three-dimensional reconstruction method for a static scene according to claim 5, wherein, In the fourth step, the calculation formula for the dynamic object threshold judgment function is: Among them, is the training iteration of the following image the mean of the object-level normalized residuals, is the training iteration of the following image the variance of the object-level normalized residuals, is the maximum number of training iterations, is the set weight coefficient.
7. A three-dimensional reconstruction method for a static scene according to claim 6, characterized in that, In the fourth step, the method for generating a set of dynamic object masks includes: Take the masks in the two-dimensional segmentation mask set where the normalized residuals of all objects are greater than the dynamic object threshold judgment function to form a dynamic object mask set .
8. The three-dimensional reconstruction method for a static scene according to claim 7, characterized in that In the fifth step, the 3D Gaussian splash radiation field loss function is: Among them, is the normalized L1 loss between the 3D Gaussian splash radiation field rendered image and the original image, is the normalized D-SSIM loss between the 3D Gaussian splash radiation field rendered image and the original image.
9. A three-dimensional reconstruction system for static scenes, characterized in that, Including the first module, the second module, the third module, the fourth module, and the fifth module; The first module is used to execute the first step: Segment all the images in the scene image set to obtain a set of two-dimensional segmentation masks for each image; The second module is used to execute the second step: Calculate the camera intrinsics, camera poses of all the images in the scene image set, and at the same time obtain the initial sparse point cloud of the scene; Initialize the 3D Gaussian splash radiation field based on the initial sparse point cloud of the scene, train the 3D Gaussian splash radiation field by combining the scene images and the corresponding pose information, establish the image pixel-level normalized residual between the rendered image and the original image, and calculate the object-level normalized residual according to the image pixel-level normalized residual and the two-dimensional segmentation mask; The third module is used to execute the third step: Calculate the residual mean and variance of the image according to the object-level normalized residual; The fourth module is used to execute the fourth step: Based on the residual mean and variance, construct a dynamic object threshold judgment function, and generate a set of dynamic object masks according to the threshold judgment function; The fifth module is used to execute the fifth step: Define the 3D Gaussian splash radiation field loss function by combining the set of dynamic object masks, and complete the three-dimensional reconstruction of the static scene by iteratively optimizing and training the 3D Gaussian splash radiation field until a preset stop condition is reached.
Citation Information
Patent Citations
Image processing method, electronic equipment and computer readable storage medium
CN116012339A
Sparse view three-dimensional reconstruction method based on generative scene image
CN118262050A
Plant image processing method, device and equipment based on three-dimensional phenotypic modeling
CN118447047A
Cited By
Static scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing
CN121414978A