Three-dimensional Gaussian sputtering method for sparse visual angle semantic priori

Through the three-dimensional Gaussian sputtering method of sparse perspective semantic priors, combined with the SAM2 segmentation model and interactive graphical interface, the problems of semantic 3D reconstruction efficiency and quality under sparse perspective in the existing technology are solved, and high-quality and controllable semantic 3D reconstruction effect is achieved.

CN120107434APending Publication Date: 2025-06-06XIDIAN UNIV

Patent Information

Application Number
CN202510170331.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently reconstruct a high-quality semantic three-dimensional Gaussian sputtering model based on sparse perspective semantic priors, and the reconstruction quality is low and difficult to control in real time.

Method used

A three-dimensional Gaussian sputtering method with a sparse viewing angle semantic prior is adopted. By collecting the target scene images, a real image collection is constructed, and the object of interest is manually selected for semantic three-dimensional reconstruction, the initial semantic image of the sparse viewing angle is obtained. Using the pre-trained SAM2 segmentation model as the interactive image sequence segmentation model, three-dimensional Gaussian primitives are initialized, and the reconstruction results are corrected in real time through the interactive graphical interface to realize parameter training and reconstruction of semantic three-dimensional Gaussian sputtering.

Benefits of technology

It realizes efficient reconstruction of high-quality semantic three-dimensional models under sparse perspective conditions, and the reconstruction results can be corrected in real time through an interactive graphical interface, improving reconstruction quality and controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107434A_ABST
    Figure CN120107434A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional Gaussian sputtering method for sparse visual angle semantic priori. The method comprises the following steps of: 1, acquiring a target scene image, constructing a real image set as a data set, and manually selecting an interested object in a target scene to perform semantic three-dimensional reconstruction to obtain an initial semantic image of a sparse view angle; a multi-view image is collected, camera external parameters and scene sparse point clouds are obtained, a plurality of views are selected and input into the SAM2 segmentation model, and a view set Is with semantic images is obtained; 2, using a pre-trained SAM2 segmentation model as an interactive image sequence segmentation model, and initializing a three-dimensional Gaussian primitive according to the scene sparse point cloud; and step 3, training parameters of semantic three-dimensional Gaussian sputtering based on the trained three-dimensional Gaussian sputtering model and the interactive image sequence segmentation model, and reconstructing a target scene. According to the method, the high-quality semantic model is efficiently reconstructed. And the reconstruction result can be easily corrected through the interactive graphical interface in the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of semantic three-dimensional reconstruction, and in particular relates to a three-dimensional Gaussian sputtering method with sparse perspective semantics prior. Background Art

[0002] In current semantic 3D reconstruction systems, methods based on 3D Gaussian sputtering reconstruction processes often rely on multi-view high-quality semantic images, which makes data acquisition difficult; it is difficult to control the reconstruction process in real time, resulting in low reconstruction quality.

[0003] S. Zhou et al. proposed a method based on distilling high-dimensional features output by a two-dimensional basic model in their paper "Feature 3dgs: Supercharging 3d gaussiansplatting to enabledistilled feature fields" (Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition. 2024: 21676-21685.); this method requires rendering high-dimensional feature maps, which will greatly affect the training and rendering efficiency. In order to distill semantic features, this method needs to use a pre-trained image model to extract the feature map of the input image in advance and save it to the storage medium, and repeatedly load the feature maps of each view during the training process. Due to the high dimensionality of the feature map, the storage space occupied by each feature map is quite large, which causes the training efficiency to be greatly affected by the bandwidth rate of the storage medium.

[0004] Publication number CN118657874A is a semantic 3D reconstruction method based on 3D Gaussian sputtering. This method uses the SAM segmentation model to process each training view image separately, uses a multimodal model to encode the semantic features of each instance, and also needs to combine the SIFT descriptor to construct color-semantic features, and finally uses the K nearest neighbor algorithm to match instances under different viewpoints. This method does not have the multi-view consistency of the segmentation results, which limits the quality of its final semantic 3D model.

[0005] In summary, the defect of the existing technology is that it is impossible to efficiently reconstruct a semantic three-dimensional Gaussian sputtering model based on sparse perspective semantic priors. Summary of the invention

[0006] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a 3D Gaussian sputtering method with sparse perspective semantic prior, which can efficiently reconstruct a high-quality semantic model under the condition of inputting a sparse perspective semantic image. And the reconstruction result can be easily corrected through an interactive graphical interface during the training process.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is:

[0008] A three-dimensional Gaussian sputtering method with sparse perspective semantic prior includes the following steps:

[0009] Step 1: Collect the target scene image, build a real image set as the data set, and manually select the objects of interest in the target scene for semantic 3D reconstruction to obtain the initial semantic image of the sparse perspective;

[0010] In step 1, a camera with calibrated intrinsic parameters is used to collect multi-view images in the scene to be reconstructed, and the camera extrinsic parameters and the scene sparse point cloud are obtained using the motion recovery structure algorithm. Several viewpoints are selected and input into the SAM2 (SegmentAnything 2) segmentation model to establish a viewpoint set I with semantic images. s ;

[0011] Step 2: Initialize the training process, use the pre-trained SAM2 segmentation model as the interactive image sequence segmentation model, and initialize the 3D Gaussian primitives according to the scene sparse point cloud.

[0012] Step 3: Based on the 3D Gaussian sputtering model and the interactive image sequence segmentation model in training, continuously expand the view set I with semantic images s , train the parameters of semantic 3D Gaussian sputtering to reconstruct the target scene.

[0013] The step 1 specifically comprises the following steps:

[0014] Step 1.1: Uniformly collect multi-view RGB images in the scene to be reconstructed to establish a real image set I;

[0015] Step 1.2: Use colmap 3D reconstruction software to perform motion recovery structure reconstruction on all images in the real image set I to obtain the intrinsic parameters of all images, camera pose parameters and the sparse point cloud P of the scene. init ;

[0016] Step 1.3, randomly select no less than 3 view images from the collected multi-view RGB images, the shooting view angles of the selected view images have a certain overlapping area, manually mark the objects of interest in the form of points or target boxes, input the marking results into the SAM2 segmentation model, obtain the semantic images of the initially selected view angles, and establish the view angle set I with semantic images from these semantic images s .

[0017] In the step 2, a training program of three-dimensional Gaussian sputtering and an interactive image sequence segmentation model backend program are run simultaneously;

[0018] To ensure the efficiency of the algorithm, the 3D Gaussian sputtering training and interactive image sequence segmentation model backend are run in multi-threaded mode on two independent computing cards.

[0019] The interactive image sequence segmentation model uses a pre-trained SAM2 segmentation model, and the video segmentation based on the SAM2 segmentation model can perform consistent multi-view image segmentation. The step 2 specifically includes the following steps:

[0020] Step 2.1: Initialize the backend of the interactive image sequence segmentation model. Use the pre-trained SAM2 model as the backend of the interactive image sequence segmentation model. Input all images in the real image set I into the interactive image sequence segmentation model for encoding.

[0021] Step 2.2: Initialization of the 3D Gaussian primitives, using the sparse point cloud P obtained in step 1.2 init Initialize the three-dimensional Gaussian primitives and the sparse point cloud P init Every point p in i Initialized as a three-dimensional Gaussian basis element g i , each three-dimensional Gaussian basis element g i Has the following properties: spatial position x i , scale component s i , Opacity α i , rotation quaternion q i , spherical harmonics c i And the class probability distribution p i ;

[0022] where x i ,s i ,α i ,q i ,c i is the basic parameter of the three-dimensional Gaussian sputtering process, which is determined by the number of scene categories N C Initialize the class probability distribution of each 3D Gaussian primitive The step 3 specifically includes the following steps:

[0023] Step 3.1: RGB image rendering and loss calculation; The rendering and loss calculation of the RGB image follow the standard 3D Gaussian sputtering training process, and the image of the training view is rendered using the image intrinsic parameters and camera pose obtained in step 1.2;

[0024] Select training view I from real image set I i , the rendered image of the training view is obtained through the differentiable standard 3D Gaussian sputtering rendering pipeline , and the loss is calculated with the real RGB image in is the real image I from the training perspectivei and render the image The L1 loss, L SSIM is the real image I from the training perspective i and render the image and D-SSIM loss, λ is a hyperparameter that controls the weight of the two, and back propagation is performed to optimize the three-dimensional Gaussian primitive parameter x i ,s i ,α i ,q i ,c i ;

[0025] Step 3.2: semantic image rendering and loss calculation;

[0026] When executing step 3.1, if the training perspective I selected in the current iteration i With semantic image, that is, I i ∈I S , then execute the training process of semantic image;

[0027] The category probability distribution p of all three-dimensional Gaussian primitives in the current training view cone is used as the rendering channel of the three-dimensional Gaussian sputtering rendering pipeline to obtain a dimension of H×W×N C The class probability distribution image P i , by formula Get the semantic image of the training perspective The semantic image loss is calculated using the cross entropy loss function, and backpropagation is performed to update the p of the 3D Gaussian basis i property;

[0028] Step 3.3: Depth regularization term based on semantic consistency constraint: The back propagation process of standard 3D Gaussian sputtering cannot effectively constrain the depth of 3D Gaussian primitives, which affects the reconstruction effect. This problem can be improved by using a depth regularization term based on semantic consistency constraint.

[0029] When performing semantic image rendering, the category probability distribution p on each rendered pixel k Similar 3D Gaussian primitives should often be at the same depth, and the depth regularization term is designed

[0030] where d i d j is the projection depth of the three-dimensional Gaussian primitive, p i 、p j is the class probability distribution of the three-dimensional Gaussian basis.

[0031] Step 3.4: During the training of 3D Gaussian sputtering, the interactive image sequence segmentation model is used to update the semantic images of each view in real time. updateAfter iterations, the algorithm will execute I s The final result is that all images in the real image set I complete the segmentation operation, I s Contains all images in the real image set I.

[0032] The step 3.4 includes the following steps.

[0033] Step 3.4.1, first the algorithm will select anchor points from the current 3D Gaussian primitives. The specific selection criteria are as follows:

[0034] (1) In front of K update The cumulative sum of rendering weights in iterations w k >thr w ;

[0035] (2) Class probability P k =argmax(softmax(p k ))>thr p ;

[0036] (3) Each time w is selected k P k The first M largest three-dimensional Gaussian primitives are used as anchor points, thr w ,thr p is a hyperparameter; thr w ,thr p It is a hyperparameter, which is set dynamically during training according to the number and quality of anchor points;

[0037] Step 3.4.2, the algorithm will then select the perspective that needs to update the semantic image; the specific selection strategy is as follows:

[0038] (1) Prioritize the selection of viewpoints that have not been segmented;

[0039] (2) Select N for each update update A perspective.

[0040] Step 3.4.3, projecting the anchor point to the selected perspective, organizing the current category of the anchor point and its coordinates in each perspective as prompts of the interactive image sequence segmentation model to perform image segmentation, and obtaining a new semantic image of the selected perspective.

[0041] In step 3.4.3, the user previews the viewpoints and anchor points automatically selected by the algorithm in real time through an interactive graphical interface or manually modifies the prompt input of the interactive image sequence segmentation model to control the semantic modeling process in real time.

[0042] Beneficial effects of the present invention:

[0043] First, the method proposed in the present invention can manually input prompts during the training process through a real-time graphical interface interaction, and control the reconstruction results in real time, so as to obtain high-quality and controllable reconstruction results.

[0044] Second, the method proposed in the present invention is based on a training process rather than a post-processing process, which can achieve end-to-end global optimization and more efficient global modeling. Therefore, the scene reconstructed by the method of the present invention is of higher quality.

[0045] Third, the method proposed in the present invention is based on the SAM2 segmentation model, and the obtained semantic image has better multi-perspective consistency. Combined with the semantic three-dimensional Gaussian sputtering process, high-quality semantic three-dimensional reconstruction can be achieved, and there are more application scenarios.

[0046] Fourth, the algorithm flow of updating the semantic image set based on the selection of anchor point three-dimensional Gaussian primitives and the real-time graphical interface interaction to provide prompts for the interactive image sequence segmentation model described in step 3.4 of the present invention. Figure 1 The consistent semantic image can improve the reconstruction result of semantic 3D Gaussian sputtering.

[0047] Fifth, the depth regularization term based on semantic consistency constraint described in step 3.3 of the present invention can constrain the depth of the three-dimensional Gaussian primitives based on semantic information, reduce the depth distortion in the reconstruction process, and bring about better reconstruction quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the algorithm flow of the present invention.

[0049] Figure 2 This is a schematic diagram of the reconstruction results of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0051] like Figure 1 As shown, a three-dimensional Gaussian sputtering method with sparse perspective semantic prior includes the following steps;

[0052] Step 1: Data collection and preprocessing: collect target scene images, build a dataset, and manually select objects of interest in the target scene for semantic 3D reconstruction.

[0053] Use a camera with calibrated intrinsic parameters to collect multi-view image sequences in the scene to be reconstructed, use the motion recovery structure algorithm to obtain the camera extrinsic parameters and the scene sparse point cloud, select several perspectives to input into the SAM2 (Segment Anything2) segmentation model, and establish a perspective set I with semantic images s . .

[0054] The step 1 specifically comprises the following steps:

[0055] Step 1.1: Evenly collect at least 50 multi-view images in the scene to be reconstructed and establish Figure 1 The real image set I is shown.

[0056] Step 1.2: Use 3D reconstruction software such as colmap to reconstruct the structure from motion for all images in the real image set I, and obtain the intrinsic parameters of all images, camera pose parameters, and the sparse point cloud P of the scene. init .

[0057] Step 1.3, randomly select no less than 3 sparse view images from the collected images, and ensure that the shooting view angles of the selected images have a certain overlap area, manually mark the objects of interest in the form of points or target boxes, and input the marking results into the SAM2 segmentation model to obtain the initial sparse view semantic image. The initial sparse view semantic image is established Figure 1 The view set I with semantic images is shown s .

[0058] Step 2: Initialize the training process to initialize the model parameters and prepare for the subsequent training process.

[0059] Using an interactive image sequence segmentation model to continuously expand and optimize the view set of semantic images during the training process of 3D Gaussian sputtering I S To do this, it is necessary to run the 3D Gaussian sputtering training program and the interactive image sequence segmentation model backend program at the same time.

[0060] To ensure the efficiency of the algorithm, the 3D Gaussian sputtering training and interactive image sequence segmentation model backend are run in multi-threaded mode on two independent computing cards.

[0061] The interactive image sequence segmentation model uses the pre-trained SAM2 segmentation model and performs consistent multi-view image segmentation based on the video segmentation capability of the SAM2 segmentation model.

[0062] The step 2 specifically includes the following steps.

[0063] Step 2.1: Initialization of the interactive image sequence segmentation model backend. Use the pre-trained SAM2 segmentation model as the interactive image sequence segmentation model backend and input all views in the real image set I into the interactive image sequence segmentation model for encoding.

[0064] Step 2.2: Initialization of the 3D Gaussian primitives, using the sparse point cloud P obtained in step 1.2 init Initialize the three-dimensional Gaussian primitive, Pinit Every point p in i Initialized as a three-dimensional Gaussian basis element g i , each three-dimensional Gaussian basis element g i With Figure 1 Properties shown: Spatial position x i , scale component s i , Opacity α i , rotation quaternion q i , spherical harmonics c i And the class probability distribution p i ;

[0065] where x i ,s i ,α i ,q i ,c i is the basic parameter of the three-dimensional Gaussian sputtering process, which is determined by the number of scene categories N C Initialize the class probability distribution of each 3D Gaussian primitive

[0066] Step 3: Semantic 3D Gaussian sputtering algorithm process based on interactive image sequence segmentation model. The purpose and function of this step is to continuously expand the view set I with semantic images. s , train the parameters of the 3D Gaussian primitives and reconstruct the target scene.

[0067] Figure 1 A 3D reconstruction semantic modeling method based on an interactive image sequence segmentation model and 3D Gaussian sputtering is demonstrated. The present invention refers to this process as an interactive semantic 3D Gaussian sputtering training process.

[0068] The step 3 specifically includes the following steps:

[0069] Step 3.1: RGB image rendering and loss calculation; The rendering and loss calculation of the RGB image follow the standard 3D Gaussian sputtering training process, and the image of the training view is rendered using the image intrinsic parameters and camera pose obtained in step 1.2;

[0070] Select training view I from real image set I i , the rendered image of the training view is obtained through the differentiable standard 3D Gaussian sputtering rendering pipeline Calculate the loss with the real RGB image in is the real image I from the training perspective i and render the image The L1 loss, L SSIM is the real image I from the training perspective i and render the image and D-SSIM loss, λ is a hyperparameter that controls the weight of the two, with a value range of [0,1), and back propagation is performed to optimize the three-dimensional Gaussian primitive parameter x i ,s i ,α i ,q i ,c i ;

[0071] Step 3.2: Semantic image rendering and loss calculation; When executing step 3.1, if the training perspective I selected in the current iteration i With semantic image, that is, I i ∈I S , then execute the training process of semantic image;

[0072] The category probability distribution p of all three-dimensional Gaussian primitives in the current training view cone is used as the rendering channel of the three-dimensional Gaussian sputtering rendering pipeline to obtain a dimension of H×W×N C The class probability distribution image P i , by formula Get the semantic image of the training perspective The semantic image loss is calculated using the cross entropy loss function, and backpropagation is performed to update the p of the 3D Gaussian basis i property;

[0073] Step 3.3: Depth regularization term based on semantic consistency constraint: The back propagation process of standard 3D Gaussian sputtering cannot effectively constrain the depth of 3D Gaussian primitives, which affects the reconstruction effect. This problem can be improved by using a depth regularization term based on semantic consistency constraint.

[0074] When performing semantic image rendering, the category probability distribution p on each rendered pixel k Similar 3D Gaussian primitives should often be at the same depth, and the depth regularization term is designed where d i d j is the projection depth of the three-dimensional Gaussian primitive, p i 、p j is the class probability distribution of the three-dimensional Gaussian basis.

[0075] Step 3.4: In order to use the interactive image sequence segmentation model to update the semantic images of each view in real time during the training process of 3D Gaussian sputtering, update After iterations, the algorithm will execute I s The final result is that all images in the real image set I complete the segmentation operation, I s Contains all images in the real image set I.

[0076] Steps 3.1-3.3 are performed continuously during training. Each rendering will render the RGB image and the semantic image at the same time, and calculate the deep regularization term of the semantic consistency constraint. Step 3.4 is performed at intervals during training. update It will be performed only once per iteration.

[0077] The step 3.4 includes the following steps.

[0078] Step 3.4.1, first the algorithm will select anchor points from the current 3D Gaussian primitives. The specific selection criteria are as follows:

[0079] (1) In front of K update The cumulative sum of rendering weights in iterations w k >thr w ;

[0080] (2) Class probability P k =argmax(softmax(p k ))>thr p ;

[0081] (3) Each time w is selected k P k The first M largest three-dimensional Gaussian primitives are used as anchor points, thr w ,thr p is a hyperparameter.

[0082] Step 3.4.2, the algorithm will then select the viewpoint that needs to update the semantic image. The specific selection strategy is as follows:

[0083] (1) Prioritize the selection of perspectives without semantic images;

[0084] (2) Prioritize the view angles of 3D Gaussian primitives with more clear semantic categories within the viewing cone;

[0085] (3) Select N for each update update A perspective.

[0086] Step 3.4.3, after obtaining the anchor point and viewpoint, project the anchor point to the selected viewpoint, organize the current category of the anchor point and its coordinates in each viewpoint into prompts of the interactive image sequence segmentation model for image segmentation, and obtain a new semantic image of the selected viewpoint. In this process, the user can preview the viewpoint and anchor point automatically selected by the algorithm in real time through the interactive graphical interface, or manually modify the prompt input of the interactive image sequence segmentation model to control the semantic modeling process in real time.

[0087] Simulation experiment:

[0088] The semantic 3D Gaussian sputtering method proposed in the present invention is used to conduct experiments using scenes in the indoor environment dataset. Figure 2 The reconstruction result shown in FIG. Among them, the upper left image is the original RGB image, the lower left image is the semantic image of the corresponding perspective, the upper right image is the RGB image rendered at the same perspective after reconstruction, and the lower right image is the semantic image rendered at the same perspective after reconstruction. It can be seen that the method proposed in the present invention has a high reconstruction quality.

Claims

1. A three-dimensional Gaussian sputtering method with sparse perspective semantic prior, characterized in that: The steps include: Step 1: Collect the target scene image, build a real image set as the data set, and manually select the objects of interest in the target scene for semantic 3D reconstruction to obtain the initial semantic image of the sparse perspective; Use a camera with calibrated intrinsic parameters to collect multi-view images in the scene to be reconstructed, use the motion recovery structure algorithm to obtain the camera extrinsic parameters and the scene sparse point cloud, select several perspectives to input into the SAM2 segmentation model, and obtain the perspective set I with semantic images s ; Step 2: Initialize the training process, use the pre-trained SAM2 segmentation model as the interactive image sequence segmentation model, and initialize the 3D Gaussian primitives based on the scene sparse point cloud; Step 3: Based on the 3D Gaussian sputtering model and the interactive image sequence segmentation model in training, continuously expand the view set I with semantic images s , train the parameters of semantic 3D Gaussian sputtering to reconstruct the target scene.

2. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 1, characterized in that: The step 1 specifically comprises the following steps: Step 1.1: Uniformly collect multi-view RGB images in the scene to be reconstructed to establish a real image set I; Step 1.2: Use colmap 3D reconstruction software to perform motion recovery structure reconstruction on all images in the real image set I to obtain the intrinsic parameters of all images, camera pose parameters and the sparse point cloud P of the scene. init ; Step 1.3, randomly select no less than 3 view images from the collected multi-view RGB images, the shooting view angles of the selected view images have a certain overlapping area, manually mark the objects of interest in the form of points or target boxes, input the marking results into the SAM2 segmentation model, obtain the semantic image of the initial selected view, and establish the view set I with semantic images from the semantic image s .

3. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 1, characterized in that: In the step 2, the training process of the three-dimensional Gaussian sputtering and the interactive image sequence segmentation model backend are run simultaneously; the three-dimensional Gaussian sputtering training and the interactive image sequence segmentation model backend are run in multi-threading on two independent computing cards; The interactive image sequence segmentation model adopts a pre-trained SAM2 segmentation model, and the video segmentation based on the SAM2 segmentation model can perform consistent multi-view image segmentation.

4. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 3, characterized in that: The step 2 specifically includes the following steps: Step 2.1: Initialize the backend of the interactive image sequence segmentation model. Use the pre-trained SAM2 model as the backend of the interactive image sequence segmentation model and input all images in the real image set I into the interactive image sequence segmentation model. Step 2.2: Initialization of the 3D Gaussian primitives, using the sparse point cloud P obtained in step 1.2 init Initialize the three-dimensional Gaussian primitives and the sparse point cloud P init Every point p in i Initialized as a three-dimensional Gaussian basis element g i , each three-dimensional Gaussian basis element g i Has the following properties: spatial position x i , scale component s i , Opacity α i , rotation quaternion q i , spherical harmonics c i And the class probability distribution p i ; where x i ,s i , α i ,q i , c i is the basic parameter of the three-dimensional Gaussian sputtering process, which is determined by the number of scene categories N C Initialize the class probability distribution of each 3D Gaussian primitive 5. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 1, characterized in that: The step 3 specifically includes the following steps: Step 3.1: Select training view I from real image set I i , the rendered image of the training view is obtained through the differentiable standard 3D Gaussian sputtering rendering pipeline Calculate the loss with the real RGB image in is the real image I from the training perspective i and render the image The L1 loss, L SSIM is the real image I from the training perspective i and render the image and D-SSIM loss, λ is a hyperparameter that controls the weight of the two, and back propagation is performed to optimize the three-dimensional Gaussian primitive parameter x i ,s i , α i ,q i , c i ; Step 3.2: When executing step 3.1, if the training view angle I selected in the current iteration i With semantic image, that is, I i ∈I S , then execute the training process of semantic image; The class probability distribution of all 3D Gaussian primitives in the current training view frustum is used as the rendering channel of the 3D Gaussian sputtering rendering pipeline, and the dimension is H×W×N. C The class probability distribution image P i , by formula Get the semantic image of the training perspective The semantic image loss is calculated using the cross entropy loss function, and backpropagation is performed to update the p of the 3D Gaussian basis i property; Step 3.3: When performing semantic image rendering, the category probability distribution p for each pixel rendered k Similar 3D Gaussian primitives should often be at the same depth, and the depth regularization term is designed where d i ,d j is the projection depth of the three-dimensional Gaussian primitive, p i 、p j is the class probability distribution of the three-dimensional Gaussian basis element; Step 3.4: During the training of 3D Gaussian sputtering, the interactive image sequence segmentation model is used to update the semantic images of each view in real time. update After iterations, the algorithm will execute I s The final result is that all images in the real image set I complete the segmentation operation, I s Contains all images in the real image set I.

6. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 5, characterized in that: The step 3.4 comprises the following steps: Step 3.4.1, first the algorithm will select anchor points from the current 3D Gaussian primitives; Step 3.4.2, the algorithm will then select the viewpoint that needs to update the semantic image; Step 3.4.3, projecting the anchor point to the selected perspective, organizing the current category of the anchor point and its coordinates in each perspective as prompts of the interactive image sequence segmentation model to perform image segmentation, and obtaining a new semantic image of the selected perspective.

7. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 6, characterized in that: In step 3.4.1, the specific screening criteria are as follows: (1) In front of U pdate The cumulative sum W of rendering weights in iterations k >thr w ; (2) Class probability P k =argmax(softmax(p k ))>thr p ; (3) Each time w is selected k , P k The first M largest three-dimensional Gaussian primitives are used as anchor points, thr w , thr p is a hyperparameter, thr w , thr p The hyperparameter is dynamically set during the training process according to the number and quality of anchor points.

8. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 6, characterized in that: In step 3.4.2, the specific selection strategy is as follows: (1) Prioritize the selection of viewpoints that have not been segmented; (2) Select N for each update update A perspective.

9. The three-dimensional Gaussian sputtering method with sparse perspective semantic prior according to claim 6, characterized in that: In step 3.4.3, the user previews the viewpoints and anchor points automatically selected by the algorithm in real time through an interactive graphical interface or manually modifies the prompt input of the interactive image sequence segmentation model to control the semantic modeling process in real time.

Citation Information

Patent Citations

  • Efficient semantic three-dimensional reconstruction method based on feature grid mapping

    CN118657874A

Cited By

  • Three-dimensional Gaussian model construction method based on semantic segmentation

    CN120451360A

  • Visual three-dimensional reconstruction method and system of structure prior, equipment and medium

    CN120510306A

  • A structural prior visual three-dimensional reconstruction method, system, device and medium

    CN120510306B

  • Three-dimensional target positioning method, system and related equipment

    CN120807634A

  • Automatic shield tunnel leakage inspection system and method based on adaptive training optimization

    CN121053565A