Semantic feature embedding open segmentation method based on space 3D Gaussian model

By introducing an instance-consistent CLIP feature extraction module and feature dimensionality and sparse sampling module in the 3D Gaussian splattering model, the existing 3D scene open segmentation method has solved the problem of large computing resource consumption and cumbersome training process, and efficient 3D scene appearance reconstruction and semantic embedding are achieved.

CN120198449APending Publication Date: 2025-06-24TIANJIN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510187353.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The open segmentation method in the existing 3D scenes has problems such as huge computing resource consumption, large training errors, and cumbersome multi-step training process. It is difficult to realize the appearance reconstruction and semantic embedding of 3D scenes under the premise of efficient and low computing costs.

Method used

Using the semantic feature embedding open segmentation method based on 3D Gaussian splattering model, an instance-consistent CLIP feature extraction module and feature dimensionality and sparse sampling module are constructed. The spatial Gaussian modeling and semantic information injection are performed by inputting multi-view pictures, simplifying the training process and realizing open segmentation.

Benefits of technology

Under the premise of high efficiency and low computational cost, the appearance reconstruction and semantic embedding of 3D scenes are realized, which simplifies the training process and improves the accuracy and efficiency of open segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198449A_ABST
    Figure CN120198449A_ABST
Patent Text Reader

Abstract

The invention provides a semantic feature embedding open segmentation method based on a space 3D Gaussian model, and the method comprises the following steps: preparing an input image set needed by training, which comprises the steps: inputting a multi-view image and corresponding camera pose parameter information; carrying out a pre-experiment to obtain a semantic-independent space 3D Gaussian; step 2, performing semantic feature training, adding semantic feature fields for the space 3D Gaussian on the basis of the model stored in the step 2, and performing feature dimension improvement on the added semantic fields through a feature dimension improvement and sparse sampling module; meanwhile, a training true value is obtained from a training input image through a CLIP feature extraction module with consistent instances and is used for monitoring semantic training; performing optimization training on the semantic feature field of the space 3D Gaussian by adopting a back propagation algorithm; and performing visual angle rendering on the semantic feature field of the space 3D Gaussian obtained by training through given test visual angle camera pose parameters to obtain a high-dimensional semantic feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a 3D space open segmentation method in the field of computer vision such as autonomous driving, virtual reality AR / VR, intelligent robots, etc., and particularly relates to a method for 3D space open segmentation based on deep learning. Background Art

[0002] With the rapid development of artificial intelligence, since the 3D Gaussian model in space was proposed, the research on it has become a hot topic in the field of computer vision. By performing Gaussian modeling on objects in three-dimensional space, this model can not only effectively capture the geometric information of spatial objects, but also generate 2D images from corresponding perspectives through a more stable, efficient and almost real-time rendering process. Various technologies improved on this model have also emerged in an endless stream. Recently, adding semantic information to the 3D Gaussian model in space to achieve downstream semantic-related tasks is also one of the popular research directions.

[0003] In recent years, deep learning technology has made semantic segmentation methods based on neural networks have broad application prospects in many fields such as autonomous driving, medical image analysis, image editing, remote sensing image analysis, etc. As Figure 1 (a) shown, semantic segmentation technology can be used for remote sensing satellite images to achieve precise positioning. Figure 1 (b) Using semantic segmentation technology to accurately segment various types of images in the medical field has become a powerful auxiliary tool for doctors to analyze the condition. Semantic segmentation is a key technology in the field of computer vision, aiming to assign corresponding semantic category labels to each pixel in an image. The developed open semantic segmentation task is more in line with real-world scenario applications, and the labels assigned to pixels can cover unseen object categories.

[0004] Open semantic segmentation is different from traditional semantic segmentation and needs to break through the limitation of fixed vocabulary. The early model proposed by Desai et al. [1] based on image-text matching, through introducing multimodal learning, enables the visual model to automatically recognize new categories according to text descriptions. With the emergence of image-text pair models such as CLIP [2], it has been widely applied to open segmentation tasks. Li et al. [3] proposed LSeg to introduce language descriptions into the segmentation task, providing detailed language guidance for image regions during the semantic segmentation process, significantly improving the flexibility and accuracy of segmentation. Xu et al. [4] proposed SAN, an attention mechanism-based vision-language model, which more precisely captures complex semantic information in images by introducing lateral adapters and multi-scale attention mechanisms. OVSeg proposed by Liang et al. [5] and CAT-Seg proposed by Cho et al. [6] further expand this task and achieve better results. Even so, these studies are all explorations at the two-dimensional image level.

[0005] Existing open segmentation methods in 3D scenes can be mainly divided into methods based on NeRF [8] and methods based on Gaussian splatting [9]. The previous work LERF [7] incorporated language information on the basis of NeRF [8], but it has a fatal flaw of huge computational resource consumption. The recently proposed Gaussian splatting technology [9] well solves this problem, and the improved methods based on Gaussian splatting technology can simultaneously perform spatial Gaussian modeling and semantic information injection of 3D scenes, so as to complete relevant reconstruction and segmentation tasks. Even so, they still have problems such as training errors caused by inconsistent multi-view of 3D space input and cumbersome multi-step training processes.

[0006] References:

[0007] [1] Desai K, Johnson J. Virtex: Learning visual representations from textual annotations[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021:11162 - 11173.

[0008] [2] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C] / / International conference on machine learning. PMLR, 2021:8748 - 8763.

[0009] [3] Li B, Weinberger K Q, Belongie S, et al. Language-driven semantic segmentation[J]. arXiv preprint arXiv:2201.03546, 2022.

[0010] [4] Xu M, Zhang Z, Wei F, et al. Side adapter network for open-vocabulary semantic segmentation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023:2945-2954.

[0011] [5] Liang F, Wu B, Dai X, et al. Open-vocabulary semantic segmentation with mask-adapted clip[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023:7061-7070.

[0012] [6] Cho S, Shin H, Hong S, et al. Cat-seg: Cost aggregation for open-vocabulary semantic segmentation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:4113-4123.

[0013] [7] Kerr J, Kim C M, Goldberg K, et al. Lerf: Language embedded radiance fields[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023:19729-19739.

[0014] [8]Mildenhall B,Srinivasan P P,Tancik M,et al.Nerf:Representing scenes as neural radiance fields for view synthesis[J].Communications of the ACM,2021,65(1):99-106.

[0015] [9]Kerbl B,Kopanas G,Leimkühler T,et al.3d gaussian splatting for real-time radiance field rendering[J].ACM Trans.Graph.,2023,42(4):139:1-139:14.

[0016]

[10] Kirillov A,Mintun E,Ravi N,et al.Segment anything[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2023:4015-4026.

[0017]

[11] Lyu W,Li X,Kundu A,et al.Gaga:Group Any Gaussians via 3D-aware Memory Bank[J].arXiv preprint arXiv:2404.07977,2024. Summary of the Invention

[0018] The technical problem to be solved by the present invention is how to achieve the appearance reconstruction and semantic embedding of 3D scenes under the premise of high efficiency and low computational cost, so as to complete the open segmentation of 3D scenes. To this end, this patent proposes an open segmentation method for semantic feature embedding based on the 3D Gaussian splatting model, and constructs an instance-consistent CLIP[2] feature extraction module and a feature dimension elevation and sparse sampling module. During training, the method completes the spatial Gaussian modeling and semantic information injection of the scene by inputting multi-view images, and then performs related scene reconstruction and open segmentation tasks during inference. The technical solutions are as follows:

[0019] An open segmentation method for semantic feature embedding based on a spatial 3D Gaussian model, comprising the following steps:

[0020] Step 1: Prepare the input image set required for training, including input multi-view images and their corresponding camera pose parameter information; semantic annotation information of the input images does not need to be provided;

[0021] Step 2: Conduct a pre-experiment, randomly initialize the parameters of the 3D Gaussian in space, and use the parameter settings of the 3D Gaussian splashing method for training to obtain a semantically irrelevant 3D Gaussian in space. After the training is completed, save the relevant parameters for subsequent training;

[0022] Step 3: Conduct semantic feature training. Based on the model saved in Step 2, add a semantic feature field to the 3D Gaussian in space. The added semantic field is passed through a feature dimension elevation and sparse sampling module to complete the feature dimension elevation. At the same time, the training input images are passed through an instance-consistent CLIP feature extraction module to obtain the training ground truth for semantic training supervision. The two parallel processes are as follows:

[0023] (1) Instance-consistent CLIP feature extraction module, the specific operations include:

[0024] A. Segment the input multi-view images to obtain the corresponding segmentation masks Segmentation mask After mask matching, complete the instance matching process of the masks at the spatial level to obtain instance-consistent segmentation masks; finally, through feature extraction and feature selection, obtain the instance-consistent CLIP features of each instance;

[0025] B. Fill the instance-consistent CLIP features of each obtained instance to obtain a pixel-level CLIP feature map For the supervision of the subsequent training process;

[0026] (2) Feature dimension elevation and sparse sampling module, the specific operations include:

[0027] After adding a learnable language feature field to each 3D Gaussian in space, the learnable language feature field of the 3D Gaussian in space can be transformed into a low-dimensional feature map under a specific perspective through rendering Subsequently, use a learnable MLP model to elevate the dimension to 512 dimensions, which is the same as F clip for alignment;

[0028] Step 4: Use the backpropagation algorithm to optimize and train the semantic feature field of the 3D Gaussian in space, and save the weight information of the relevant MLP model during the training process for subsequent testing and inference;

[0029] Step 5: Render the semantic feature field of the spatially 3D Gaussian obtained through training from the perspective of the given test perspective camera pose parameters to generate a low-dimensional semantic feature map under this perspective; use the MLP loaded with the relevant MLP model weight information saved in Step 4 to perform dimensionality elevation on the low-dimensional semantic feature map to obtain a high-dimensional semantic feature map;

[0030] Step 6: The predefined test text is used to obtain text features through the CLIP text encoder, calculate the similarity between the high-dimensional semantic feature map and the text features, obtain the final semantic segmentation map, and complete the open semantic segmentation task.

[0031] Further, in Step 3, for the i-th instance, the instance-consistent CLIP feature extraction is specifically as follows: Crop the original input image regions corresponding to multiple segmentation masks of the same instance from multiple perspectives and send them into the image encoder of CLIP to obtain multiple CLIP features of this instance where n i represents that the i-th instance exists in n i input perspectives. In the subsequent feature selection, calculate the cosine similarity between the n i features and the other n i -1 features Select the feature with the maximum similarity as the instance-consistent CLIP feature of the i-th instance.

[0032] Further, during the training process of Step 3, the high-dimensional feature map obtained through the language feature field is supervised by the already obtained pixel-level CLIP feature map ; Sparse sampling is performed on and and to achieve a reduction in computing resources and an improvement in training efficiency through sparse sampling.

[0033] Further, during the training process of Step 3, the specific implementation of achieving a reduction in computing resources and an improvement in training efficiency through sparse sampling is to randomly generate a binary mask to achieve sparse sampling at a specified ratio.

[0034] The present invention is based on the Gaussian splash technology, which ensures the input multi-perspective consistency of 3D scene objects and simplifies the training process to complete the open segmentation of the target. Description of the Drawings

[0035] Figure 1 Application Example of Semantic Segmentation

[0036] Figure 2 Basic Architecture of the Semantic Feature Training Strategy

[0037] Figure 3Specific implementation of the method proposed in this patent

[0038] Figure 4 Schematic diagram of the segmentation result of the "sofa" scene

[0039] Figure 5 Quantitative comparison (mIoU) of 3D semantic segmentation on the 3D-OVS dataset Specific implementation

[0040] The present invention will be described below in conjunction with the accompanying drawings and embodiments.

[0041] We first introduce the implementation details of the proposed instance-consistent CLIP [2] feature extraction module and the feature dimension elevation and sparse sampling module, and then introduce how to apply the trained spatial Gaussian model of this method to the open segmentation task. Figure 2 The overall process architecture of the training of the method of this patent is given.

[0042] (1) Instance-consistent CLIP feature extraction module

[0043] The instance-consistent CLIP feature extraction module first uses the frozen SAM (Segment Angthing Model)

[10] model to segment the input multi-view images to obtain the corresponding segmentation masks where N represents the number of mask segmentations of all input multi-view images, and H and W represent the height and width of the images. Subsequently, these masks pass through the mask matching module, and the GAGA (Group AnyGaussians)

[11] in this module completes the instance matching process of the masks at the spatial level, obtaining instance-consistent segmentation masks That is, each mask is assigned an ID value, and the masks belonging to the same instance are assigned the same ID value. The instance-consistent CLIP (Contrastive Language–Image Pretraining) feature extraction module constructed in the embodiment of the present invention refers to the model in reference [2] and incorporates newer algorithms, such as references

[10] and

[11] . If the above-mentioned frozen SAM and GAGA algorithms are not adopted, there are various algorithms in the prior art that can achieve this function, and the technical effects may be different.

[0044] We crop out the original input image regions corresponding to multiple masks of the same instance under multiple views and send them into the image encoder of CLIP, thereby obtaining multiple CLIP features of this instance where n i represents that the i-th instance exists in n iAmong the input perspectives, after that, we need to select the most distinguishable CLIP feature as the CLIP feature of the i-th instance. In the subsequent feature selection module, the cosine similarities between n i features and the other n i - 1 features are calculated respectively Finally, the feature with the maximum similarity is selected as the instance-consistent CLIP feature of the i-th instance. This operation is performed on all instances in turn, and the instance-consistent CLIP features of all instances are obtained. The specific feature selection process is as follows:

[0045]

[0046] After obtaining the instance-consistent CLIP features of each instance, we fill these features back into the mask to obtain the pixel-level CLIP feature map for the supervision of the subsequent training process.

[0047] (3) Feature Dimension Elevation and Sparse Sampling Module

[0048] We add a learnable language feature field to each spatial 3D Gaussian. The feature dimension elevation and sparse sampling module first uses an MLP (Multilayer Perceptron) to elevate the low-dimensional language features to high-dimensional language features, and then performs a sparse sampling strategy on the high-dimensional features to improve the training efficiency.

[0049] The learnable language feature field of the spatial 3D Gaussian is first rendered into a low-dimensional feature map under a specific perspective The whole process is similar to the rendering process of the spatial 3D Gaussian color, as follows:

[0050]

[0051] where l i is the language feature field assigned to the i-th spatial 3D Gaussian, and α i represents the opacity of the i-th 3D Gaussian, that is, the weight determining the rendering contribution. Subsequently, a learnable MLP elevates the dimension to 512 dimensions aligned with F clip as follows:

[0052]

[0053] During the training process, we use the pixel-level CLIP feature map already obtained to Supervise. However, this method is time-consuming and memory-intensive. To avoid this problem, we introduce a sparse sampling strategy that randomly selects some features for training. Specifically, by randomly generating a binary mask ensuring that a certain proportion of its elements are 1 and the rest are 0 to achieve sparse sampling at a specified ratio, we choose L1 loss supervision, which is calculated as follows:

[0054]

[0055] (4) Apply the proposed semantic feature embedding open segmentation method to the open segmentation in 3D space

[0056] To apply the proposed two-stage spatial Gaussian training to 3D space open semantic segmentation, we need to go through two steps: the training process and the testing process. The training process aims to learn the learnable semantic feature fields of each spatial 3D Gaussian to complete the injection of semantic information. In the testing stage, the trained semantic feature fields are rendered into low-dimensional feature maps, and the open semantic segmentation is completed through the corresponding text features of CLIP. The training process and testing process of the proposed method are introduced in detail below.

[0057] First, we introduce the specific training process:

[0058] Step 1: Prepare the input image set for training (such as the LERF-Mask, 3D-OVS datasets), including the input pictures and their corresponding camera pose parameter information. Here, we do not need the semantic annotation-related information of the input images.

[0059] Step 2: Complete the pre-experiment. The pre-experiment randomly initializes the spatial 3D Gaussian parameters and trains using the parameter settings of the 3D Gaussian splashing method to obtain semantically irrelevant spatial 3D Gaussians, and saves the relevant parameters for subsequent training.

[0060] Step 3: Based on the model saved in Step 2, add semantic feature fields to the spatial 3D Gaussians. The spatial 3D Gaussians go through the feature upsampling and sparse sampling module, while the input images go through the instance-consistent CLIP feature extraction module to complete the semantic training of the spatial 3D Gaussians.

[0061] Step 4: Based on the backpropagation algorithm, train the semantic feature fields of the spatial 3D Gaussians and save the weight information of the relevant MLP during the training process.

[0062] Then, we introduce the specific testing process:

[0063] Step 1: Given the camera pose parameters of the test perspective, we render the semantic feature field of the spatially 3D Gaussian obtained through training from the given perspective to obtain the corresponding low-dimensional semantic feature map, and use the MLP loaded with the saved relevant weight information to increase the dimension of the low-dimensional semantic feature map to obtain the high-dimensional semantic feature map. Then, the pre-defined test text in the dataset is passed through the CLIP text encoder to obtain text features, and the final segmentation map corresponding to these texts is obtained by calculating the similarity between these text features and the high-dimensional semantic feature map, thus completing the open segmentation task in the 3D scene.

[0064] The semantic training method based on spatially 3D Gaussian proposed by the present invention is applicable to the open segmentation task in the 3D scene. Figure 3 The specific implementation method of the method proposed by the present invention is given, and the specific implementation steps are as follows:

[0065] Step 1: First, prepare the input image set required for training, such as the LERF-Mask and 3D-OVS datasets, including the input images and their corresponding camera pose parameter information. It should be noted that in this step, semantic annotation information of the input images does not need to be provided, only the images themselves and their camera poses are required. Taking the 3D-OVS dataset as an example, there are approximately 30 input pictures for each scene.

[0066] Step 2: Conduct a pre-experiment, randomly initialize the parameters of the spatially 3D Gaussian, and use the parameter settings of the 3D Gaussian splash method for training to obtain a spatially 3D Gaussian that is semantically irrelevant. After the training is completed, save the relevant parameters for subsequent training use.

[0067] Step 3: Conduct semantic feature training. Based on the model saved in Step 2, add a semantic feature field to the spatially 3D Gaussian, and complete the semantic training in the process of the feature dimension increase and sparse sampling module and the instance-consistent CLIP feature extraction module.

[0068] Step 4: Use the backpropagation algorithm to optimize the training of the semantic feature field of the spatially 3D Gaussian, and save the weight information of the relevant MLP model during the training process for subsequent testing and inference.

[0069] Step 5: Given the camera pose parameters of the test perspective, perform perspective rendering on the semantic feature field of the spatially 3D Gaussian obtained through training to generate the corresponding low-dimensional semantic feature map. Use the MLP loaded with the saved relevant weight information to increase the dimension of the low-dimensional semantic feature map to obtain the high-dimensional semantic feature map.

[0070] Step 6: The pre-defined test text passes through the CLIP text encoder to obtain text features. Calculate the similarity between the high-dimensional semantic feature map of the image and the text features to obtain the final semantic segmentation map and complete the open semantic segmentation task.

Claims

1. The semantic feature embedding open segmentation method based on the spatial 3D Gaussian model includes the following steps: Step 1: Prepare the input image set required for training, including the input multi-view images and their corresponding camera pose parameter information; Step 2: Conduct a preliminary experiment, randomly initialize the parameters of the spatial 3D Gaussian, and use the parameter settings of the 3D Gaussian splashing method for training to obtain semantically irrelevant spatial 3D Gaussian. After the training is completed, save the relevant parameters for subsequent training; Step 3: Perform semantic feature training. Based on the model saved in step 2, add semantic feature fields to the spatial 3D Gaussian. The added semantic fields are upgraded through the feature dimension upgrade and sparse sampling modules. At the same time, the training input image is extracted through the instance-consistent CLIP feature extraction module to obtain the training true value for semantic training supervision. The two parallel processes are as follows: (1) CLIP feature extraction module with consistent instance, the specific operations include: A. Segment the input multi-view image to obtain the corresponding segmentation mask Segmentation Mask After mask matching, the instance matching process of the mask is completed at the spatial level to obtain the instance-consistent segmentation mask; finally, after feature extraction and feature selection, the instance-consistent CLIP features of each instance are obtained; B. Fill the CLIP features consistent with each instance to obtain the pixel-level CLIP feature map Used for supervision of the subsequent training process; (2) Feature dimension upgrading and sparse sampling module, the specific operations include: After adding a learnable language feature field to each spatial 3D Gaussian, the learnable language feature field of the spatial 3D Gaussian can be transformed into a low-dimensional feature map at a specific perspective after rendering. Then a learnable MLP model is used to transform The dimension is raised to the same level as F clip The same 512 dimensions are used for alignment; Step 4: Use the back propagation algorithm to optimize the semantic feature field of the spatial 3D Gaussian, and save the weight information of the relevant MLP model during the training process for subsequent testing and reasoning; Step 5: Use the given test view camera pose parameters to render the semantic feature field of the trained spatial 3D Gaussian to generate a low-dimensional semantic feature map under this view; use the MLP loaded with the relevant MLP model weight information saved in step 4 to increase the dimension of the low-dimensional semantic feature map to obtain a high-dimensional semantic feature map; Step 6: The predefined test text is passed through the CLIP text encoder to obtain text features, and the similarity between the high-dimensional semantic feature map and the text features is calculated to obtain the final semantic segmentation map, completing the open semantic segmentation task.

2. The semantic feature embedding open segmentation method according to claim 1, characterized in that: In step 3, for the i-th instance, the instance consistent CLIP feature extraction is specifically as follows: crop the original input image area corresponding to multiple segmentation masks of the same instance under multiple views, and send it to the CLIP image encoder to obtain multiple CLIP features of this instance where n i Represents that the i-th instance exists in n i In the subsequent feature selection, n i features and other n i - Cosine similarity of 1 feature The feature with the largest similarity is selected as the instance-consistent CLIP feature for the i-th instance.

3. The semantic feature embedding open segmentation method according to claim 1, characterized in that: During the training process of step 3, the pixel-level CLIP feature map is obtained. The high-dimensional feature graph obtained by the language feature field Through sparse sampling strategy and Sparse sampling is performed to reduce computing resources and improve training efficiency.

4. The semantic feature embedding open segmentation method according to claim 3, characterized in that: In the training process of step 3, sparse sampling is performed to reduce computing resources and improve training efficiency by randomly generating a binary mask. Implements sparse sampling with a specified ratio.

Citation Information

Cited By

  • Scene perception model training method and device, robot control method and robot

    CN120635678A