Multi-granularity open vocabulary query method based on object-level lossless Gaussian field

Through the combination of object-level lossless Gaussian field and multi-layer perceptron, the problem of insufficient object-level understanding in traditional methods is solved, the retention of high-dimensional semantic features and accurate understanding of component-levels is achieved, and the accuracy and efficiency of robot scene understanding is improved.

CN120580432APending Publication Date: 2025-09-02BEIJING INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510679561.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Traditional semantic graphing methods lack accurate object-level understanding in robot scene understanding, and the method based on neural radiation field takes time, feature compression methods lose semantic fidelity, and methods that retain complete semantic features lack object-level geometric constraints, resulting in blurred boundaries and fragmented object understanding.

Method used

A multi-grained open vocabulary query method based on object-level lossless Gaussian field is adopted to construct object-level Gaussian fields through object-level feature codebook model and Gaussian rendering model, and combine multi-layer perceptron to learn component-level features to achieve object-level understanding.

Benefits of technology

While retaining high-dimensional semantic features, it realizes accurate understanding of object level and component level, supports multi-grained scene editing, and improves the accuracy and efficiency of robot scene understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580432A_ABST
    Figure CN120580432A_ABST
Patent Text Reader

Abstract

The invention provides a multi-granularity open vocabulary query method based on an object-level lossless Gaussian field, and the method comprises the steps: introducing an object-level Gaussian field with a global consistency codebook, rendering a learnable semantic tag vector in the Gaussian field back to a corresponding object tag; the direct mapping between the label and the corresponding uncompressed high-dimensional feature is established through the code book, so that the semantic feature of any dimension is supported, additional compression is not needed, and the understanding capability on an object is remarkably improved; according to the method, wide quantitative and qualitative evaluation is carried out in a plurality of scenes, excellent performance in the aspects of object level zero sample segmentation and open vocabulary understanding is shown, the highest precision is particularly achieved in object-component hierarchical retrieval, and meanwhile multi-granularity scene editing is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robotics technology, and in particular relates to a multi-granularity open vocabulary query method based on object-level lossless Gaussian field. Background Art

[0002] Traditional semantic mapping methods usually use deep learning to perform object detection and unsupervised segmentation, which are seamlessly integrated into scene models. With the breakthrough progress of visual language models in the two-dimensional field, fine-grained segmentation methods that integrate visual language features have opened up new possibilities for understanding three-dimensional open vocabulary. A common strategy is to intuitively project the two-dimensional segmentation and understanding results into three dimensions, and achieve visual comprehension through iterative fusion of each frame’s features based on geometric and feature similarity. Figure 1 Another class of methods is based on neural radiance fields (NeRfs), which learn semantic features through neural implicit fields to obtain a continuous multi-scale feature space, significantly improving the spatial continuity of semantic features and providing richer scene semantic information. Another class of methods is based on three-dimensional Gaussian splatting (3DGS). Although tile-based rasterization accelerates the rendering and learning of high-dimensional open-word features through dense regularization, direct explicit embedding leads to high memory consumption, hindering efficiency and granularity. To reduce the memory burden of millions of raw features from three-dimensional Gaussian distributions, methods such as Langsplat compress features using dimensionality reduction, such as variational autoencoding, and then embed low-dimensional semantic features into Gaussians, thereby efficiently learning semantic features. Additionally, methods such as Opengaussian and OccamLGS tend to preserve the original underlying information. The Opengaussian method discretizes instance features using a codebook, while the Occam LGS method introduces a global optimization method that does not require training, both of which maintain the integrity of semantic features.

[0003] Traditional semantic mapping methods either rely on segmentation models pre-trained with a limited set of categories or are limited by coarse semantic understanding, making it difficult to provide robots with accurate scene understanding. Neural radiance field-based methods lack a clear and consistent understanding of objects in a scene, and their time-consuming ray sampling strategies also limit the usability of maps.

[0004] 3D Gaussian splatting methods based on feature compression inevitably reduce semantic fidelity and lose some semantic information regardless of the form of dimensionality reduction employed. Methods that retain complete semantic features, while enabling the retrieval of specific objects, rely on point-level representations and lack explicit object-level geometric constraints, leading to blurred boundaries and fragmented object understanding. Furthermore, these methods lack structured, multi-granular representations for scene understanding, limiting their performance in fine-grained downstream robotics tasks. Summary of the Invention

[0005] To solve the above problems, the present invention provides a multi-granularity open vocabulary query method based on object-level lossless Gaussian field, which can achieve object-level and component-level understanding while retaining high-dimensional semantic features.

[0006] A multi-granularity open vocabulary query method based on object-level lossless Gaussian field, comprising the following steps:

[0007] Input multiple RGB images of the current scene into the object-level feature codebook model to obtain a lossless triplet codebook corresponding to all objects in the current scene, where any triplet consists of a label, a description text, and a feature vector.

[0008] Input multiple RGB images of the current scene into the Gaussian rendering model to obtain the object-level Gaussian field corresponding to the current scene. The attributes of each Gaussian sphere that constitutes the Gaussian field include the center coordinates (x, y, z) of the Gaussian sphere and the object label label corresponding to the Gaussian sphere;

[0009] The object-level vocabulary of interest set by the user is converted into an object vocabulary feature vector. The label corresponding to the object vocabulary feature vector is obtained according to the triple lossless code book. Then, the Gaussian sphere corresponding to the object-level vocabulary of interest is filtered out from the Gaussian field based on the label. Finally, the object Gaussian formed by all the filtered Gaussian spheres is used as the object-level query result;

[0010] The center coordinates (x, y, z) of all the filtered Gaussian balls are input into the multi-layer perceptron to obtain the component-level features corresponding to each Gaussian ball;

[0011] Each Gaussian sphere is colored based on the cosine similarity between the component vocabulary feature vector converted from the user-set component-level vocabulary of interest and the component-level features corresponding to each Gaussian sphere, and the Gaussian formed by the darkest Gaussian sphere is used as the component-level query result.

[0012] Furthermore, the multi-frame RGB images of the current scene are input into the object-level feature codebook model to obtain the triplet lossless codebook corresponding to all objects contained in the current scene:

[0013] The CropFormer model is used to identify and segment instances in multiple RGB images and generate object segmentation masks corresponding to each RGB image frame. Where t represents the t-th frame, i represents the i-th object in the t-th frame RGB image;

[0014] Segment the object mask corresponding to each frame of RGB image The visual embedding branch uses the CLIP model to perform contrast alignment in the visual-language space to obtain the visual language features of the object segmentation mask. Extract visual-linguistic information and generate visual embeddings corresponding to each frame of RGB image At the same time, the language reasoning branch uses the frozen image encoder of the TAP model to generate descriptive text captions containing color, material, attributes, and spatial relationships. t,i , and then use the SBERT model to cap the description text corresponding to each frame RGB image t,i Encoded as language feature vector

[0015] Segment the object mask Visual Embedding Description text cap t,i , language feature vector The object-level high-dimensional features corresponding to each frame of RGB image;

[0016] Perform mask clustering on the object-level high-dimensional features corresponding to each frame of RGB image to obtain the initial triplet corresponding to each frame of RGB image, where the initial triplet is label-description text-feature vector;

[0017] The description texts in the initial triples corresponding to the RGB images of all frames are clustered among the frames to obtain the final triplet lossless codebook corresponding to all objects in the current scene.

[0018] Furthermore, the loss function L used in training the Gaussian rendering model is as follows:

[0019]

[0020] in, is the L1 loss function of RGB image, λ1 is the weight of the L1 loss function of RGB image, is the RGB image structure similarity loss function, λ2 is the weight of the RGB image structure similarity loss function, is the object label loss function, λ3 is the weight of the object label loss function, is the depth map loss function, and λ4 is the weight of the depth map loss function.

[0021] Furthermore, the depth map loss function The calculation method is as follows:

[0022]

[0023] in, is the true depth map corresponding to the t-th frame RGB image of the current scene, is the depth image of the tth frame after Gaussian model tile rasterization rendering, and β is the preset depth map difference threshold.

[0024] Furthermore, the training method of the multilayer perceptron is:

[0025] Use the SAM model to analyze the RGB images of each frame of the current scene Perform preprocessing to obtain the dense segmentation mask corresponding to each frame of RGB image Where p represents the component level, t represents the t-th frame, and i represents the i-th component in the t-th frame RGB image;

[0026] The CLIP model is used to align the dense segmentation masks of each frame of RGB images in the visual-language space. Then from the aligned dense segmentation mask Extract visual-language information from the image and obtain component-level features corresponding to each frame of RGB image

[0027] Based on the mask confidence score, the component level features corresponding to all frame RGB images are fused The overlapping component-level features in the final result are the dense component-level features corresponding to each pixel position of each frame RGB image.

[0028] The depth value Z at each pixel coordinate (u, v) of the rendered depth map obtained after tile rasterization based on the Gaussian field of the current scene projects the pixel (u, v) to the world coordinate system (x, y, Z), and then all dense component-level features of the RGB images from different frames corresponding to the pixel coordinate (u, v) are calculated. Assign values ​​to the coordinates (x, y, Z) in the world coordinate system to obtain the high-dimensional features f corresponding to each coordinate in the world coordinate system p , to form a four-tuple (x, y, Z, f p ) form of training set;

[0029] Each three-dimensional coordinate (x, y, Z) and its corresponding four-tuple (x, y, Z, f p ) as the input of the multi-layer perceptron, which outputs the predicted component-level features corresponding to each Gaussian ball in the current scene

[0030] According to the high-dimensional feature f p and predicting part-level features The loss function is constructed by the difference between them, and the parameters of the multilayer perceptron are optimized according to the backpropagation of the loss function until the training reaches the set round and the model loss converges, and a trained multilayer perceptron is obtained.

[0031] Furthermore, we can obtain the four-tuple (x, y, Z, f) corresponding to each three-dimensional coordinate (x, y, Z) in the world coordinate system. p ), the four-tuple (x,y,Z,f p ) performing voxel-based feature extraction to obtain a high-dimensional quadruple corresponding to each voxel, and then using the high-dimensional quadruple corresponding to each voxel as the input of the multi-layer perceptron;

[0032] The voxel feature extraction operation includes the following steps:

[0033] Divide the three-dimensional space in the world coordinate system corresponding to the current scene into multiple voxels;

[0034] The three-dimensional coordinates corresponding to each pixel take the voxel closest to itself as its own voxel;

[0035] The four-tuples (x, y, Z, f) corresponding to all three-dimensional coordinates in each voxel are respectively p ) are combined into an overall feature, and each overall feature and the center coordinate of each voxel are combined into a high-dimensional quadruple corresponding to each voxel.

[0036] Furthermore, the method for obtaining the label corresponding to the object vocabulary feature vector according to the triple lossless codebook is:

[0037] The argmax function is used to calculate the cosine similarity between the object vocabulary feature vector and each feature vector in the triple lossless codebook, and the label with the largest cosine similarity is used as the label corresponding to the object vocabulary feature vector;

[0038] At the same time, the Gaussian sphere corresponding to the object-level query result is highlighted in red; the component-level query result is displayed in the Gaussian field in the form of a heat map.

[0039] Furthermore, the CLIP model and the SBERT model are used to convert the object-level vocabulary of interest set by the user into object vocabulary feature vectors, and also to convert the component-level vocabulary of interest set by the user into component vocabulary feature vectors.

[0040] Furthermore, when assigning colors to the Gaussian spheres, the Gaussian spheres with greater cosine similarity to the component vocabulary feature vectors have darker corresponding colors.

[0041] Beneficial effects:

[0042] 1. The present invention provides a multi-granularity open vocabulary query method based on object-level lossless Gaussian fields. It introduces an object-level Gaussian field with a globally consistent codebook. After the learnable semantic label vector in the Gaussian field is rendered back to the corresponding object label, a direct mapping between the label and the corresponding uncompressed high-dimensional feature is established through the codebook, thereby supporting semantic features of arbitrary dimensions without the need for additional compression, significantly improving the ability to understand objects. The present invention has conducted extensive quantitative and qualitative evaluations in multiple scenarios, demonstrating excellent performance in zero-shot segmentation and open vocabulary understanding at the object level, especially achieving the highest accuracy in object-part hierarchical retrieval, while also supporting multi-granularity scene editing.

[0043] 2. The present invention provides a multi-granularity open vocabulary query method based on object-level lossless Gaussian fields, and develops an implicit component-level feature field. It directly learns high-dimensional features through a multi-layer perceptron fully connected network, seamlessly integrates fine-grained understanding capabilities into object-level Gaussian fields, and promotes the application of multi-granularity downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 The multi-granularity open vocabulary query framework based on object-level lossless Gaussian field provided by the present invention;

[0045] Figure 2 The voxel feature extraction process provided by the present invention;

[0046] Figure 3 Comparison of semantic segmentation effects provided by the present invention;

[0047] Figure 4 Comparison of the segmentation effects of the examples provided by the present invention;

[0048] Figure 5 Comparison of object-level query effects provided by the present invention;

[0049] Figure 6 Comparison of component-level query effects provided by the present invention;

[0050] Figure 7 This is a qualitative comparison of the results of the multi-granularity 3D scene editing provided by the present invention. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0052] In order to solve the problem of retaining the semantic information of the original basic features in the Gaussian field at the object level without compression, and to solve the problem of extracting high-dimensional features from the Gaussian feature field at the object level, thereby supporting multi-level cognition from object-level semantics to fine-grained part-level details; the present invention processes the input color image I c , Depth Map I d and pose P, constructing an object-level Gaussian field. Each Gaussian distribution contains basic attributes and object labels associated with high-dimensional information-preserving features, which are mapped through a consistent codebook. Subsequently, the object-level Gaussian coordinates x, y, z are combined with the component-level features f through an implicit feature field based on a compact MLP. p This architecture enables the system to achieve accurate color information, geometric representation, and multi-level understanding at the object and component level.

[0053] The overall framework of the present invention is as follows Figure 1 As shown in the figure. First, visual-linguistic features are extracted from the color image, and geometric and photometric attributes are aligned to generate pixel-level object label pseudo-truth maps and feature codebooks. Next, the object-level Gaussian field is trained using the RGB-D image and the object label pseudo-truth maps. Then, component-level features are extracted using SAM (Segment Anything Model) and VLM (Visual Language Model), and combined with the object-level Gaussian coordinates to train an MLP to obtain detailed component information. Finally, the object-level codebook and component-level feature field are used to calculate similarity with the query text, thereby achieving accurate multi-level retrieval.

[0054] Specifically, a multi-granularity open vocabulary query method based on object-level lossless Gaussian field includes the following steps:

[0055] S1: Input multiple RGB images of the current scene into the object-level feature codebook model to obtain a lossless codebook of triples corresponding to all objects in the current scene, where any triplet consists of a label, a description text, and a feature vector.

[0056] Specifically, the multi-frame RGB images of the current scene are input into the object-level feature codebook model to obtain the triplet lossless codebook corresponding to all objects contained in the current scene:

[0057] S11: For each RGB input frame Use the pre-trained category-independent model CropFormer to identify and segment instances in multiple frames of RGB images, and generate object segmentation masks corresponding to each frame of RGB image Among them, c is the abbreviation of color, which is obtained by RGB image processing, t represents the t-th frame, and i represents the i-th object in the t-th frame RGB image;

[0058] S12: In order to bridge the gap between low-level visual information and high-level semantics, the present invention segments the object corresponding to each frame of RGB image into The visual embedding branch uses the CLIP model to perform contrast alignment in the visual-language space, and extracts the object segmentation mask from the object segmentation mask. Extract visual-linguistic information and generate visual embeddings corresponding to each frame of RGB image The clip represents the semantic feature obtained using the CLIP model, and the semantic feature is 512-dimensional. At the same time, the language reasoning branch uses the frozen image encoder of the TAP model to generate a descriptive text cap containing color, material, attributes and spatial relationships. t,i , and then use the SBERT model to cap the description text corresponding to each frame RGB image t,i Encoded as language feature vector Among them, cap represents the semantic feature obtained by encoding the caption of the object, and the semantic feature is 384-dimensional. This feature can capture deep information from the object description and enhance target understanding through language common sense knowledge;

[0059] S13: Segment the object mask Visual Embedding Description text cap t,i , language feature vector The object-level high-dimensional features corresponding to each frame of RGB image;

[0060] S14: performing mask clustering on the object-level high-dimensional features corresponding to each frame of RGB image to obtain an initial triplet corresponding to each frame of RGB image, wherein the initial triplet is a label-description text-feature vector;

[0061] It should be noted that occlusion and viewpoint changes can lead to inconsistencies in object segmentation across frames, which can affect scene reasoning and downstream tasks. To ensure global consistency of object understanding trained with Gaussian, we use OpenObj’s two-stage clustering method. Coarse clustering builds an association graph, which uses the object mask generated by the intra-frame understanding. As nodes, the weights of edges are determined by the geometric, color, and semantic similarity of the mask. Fine clustering further refines these results to resolve object fragmentation issues at the edges of the image. This process generates the label index idx for each object in each frame. i (randomly assigned), semantic features f iclip ,f i cap and descriptive phrases cap i and dense object label maps Where o is the abbreviation for object, and t represents the tth frame. While globally consistent object understanding is taking shape, the same object in different frames, while having the same label index, may have different semantic features and description phrases due to the different perspectives of each frame, making it far from completely unified and accurate.

[0062] S15: Perform the maximum clustering between frames on the description texts in the initial triples corresponding to the RGB images of all frames to obtain the final triplet lossless codebook corresponding to all objects contained in the current scene.

[0063] It should be noted that to achieve globally consistent segmentation and understanding, this paper systematically integrates cross-frame labels, descriptions, and high-dimensional feature representations to construct an object-level semantic codebook. Due to the inherent variability of per-frame descriptions and feature embeddings, direct aggregation introduces redundancy and inconsistency, while false detections can degrade the quality of the global representation. To address this issue, this paper adopts a two-stage fusion strategy: first, objects detected across all frames are integrated, and then their semantic representations are optimized.

[0064] Specifically, the present invention aggregates all object descriptions and their occurrence frequencies and applies the Maximum Clustering Clusters (MCC) method to select the most representative descriptions. This prevents outlier detection from corrupting the global feature space and ensures semantic consistency. Subsequently, high-dimensional features associated with the selected descriptions are clustered and averaged to produce a robust and compact semantic representation of each object.

[0065] In the final codebook, each object is indexed by label idx i , representative description cap i and high-dimensional semantic feature embedding f i clip ,f i cap Representation, where label indexing and clustering across frames remain consistent. This structured representation allows direct retrieval of semantic information based on object labels decoded from Gaussian representations, ensuring efficient and information-preserving integration into downstream tasks.

[0066] As can be seen, step S1 generally designs a two-branch feature extraction framework, combined with globally consistent mask clustering within and across frames to obtain an information-preserving object-level feature codebook for Gaussian training. To achieve consistent object understanding across frames, a coarse-fine two-stage mask clustering is used to assign consistent labels to objects in the scene across all frames (randomly assigned during the clustering process). The label, description, and high-dimensional features of the object in each frame are then integrated through maximum clustering to construct a lossless codebook containing "label-description-high-dimensional feature" triplets.

[0067] S2: Input multiple RGB images of the current scene into the Gaussian rendering model to obtain the object-level Gaussian field corresponding to the current scene. The attributes of each Gaussian sphere that constitutes the Gaussian field include the center coordinates (x, y, z) of the Gaussian sphere and the object label label corresponding to the Gaussian sphere;

[0068] It should be noted that in order to represent the object information in the scene, the present invention uses a low-dimensional learnable vector v∈R 16 To represent the object label of each Gaussian sphere, we use a vector that is consistent for all objects under different views and is directly linked to the codebook to retrieve the corresponding high-dimensional semantic features. Ultimately, we can use the same differentiable renderer as the spherical harmonic coefficients to render 2D object label maps from various views.

[0069] (1) Scene rendering:

[0070] Specifically, 3D Gaussian sputtering explicitly represents a 3D scene as a set of anisotropic 3D Gaussian distributions. Each Gaussian distribution G(x) is represented by its mean μ∈R 3 And the covariance matrix Σ represents:

[0071]

[0072] G 2D (x) = JWG(x)W T J T (2)

[0073] Under the viewing pose W and projection matrix J, the influence weight of each Gaussian distribution The total influence of each pixel is contributed by depth-based sorting and blending overlapping points, denoted as T i :

[0074]

[0075] Among them, the final rendered 2D label feature V for each pixel id is the label vector v representing the contribution Gaussian distribution i Then, the feature extraction model M is used to extract the label feature V id Mapping to dense object label map Indicates the object label of each pixel in the current view.

[0076] (2) Loss function:

[0077] The present invention uses RGB images Pseudo-real object label map and depth map (where d represents depth, i.e., depth, and t represents the tth frame) to supervise the training process. Similar to the original 3DGS, the present invention adopts L1 loss and SSIM loss To calculate and Photometric regularization (image loss) between and using cross entropy loss Calculate predicted labels and target label The classification loss between (object label loss).

[0078] For deep supervision, we use smooth L1 loss It replaces the L1 loss to reduce training instability caused by changes in depth distribution due to random initialization of Gaussian parameters. Unlike the L1 loss, the smoothed L1 loss provides a quadratic response to small errors, smoothing gradients and stabilizing training, while approximating the L1 loss to ensure efficient optimization for larger errors. This adjustment accelerates learning of consistent depth constraints.

[0079] In summary, the loss function L used in training the Gaussian rendering model in this invention is as follows:

[0080]

[0081] in, is the L1 loss function of RGB image, λ1 is the weight of the L1 loss function of RGB image, is the RGB image structure similarity loss function, λ2 is the weight of the RGB image structure similarity loss function, is the object label loss function, λ3 is the weight of the object label loss function, is the depth map loss function, and λ4 is the weight of the depth map loss function.

[0082] Depth map loss function The calculation method is as follows:

[0083]

[0084] in, is the true depth map corresponding to the t-th frame RGB image of the current scene, is the depth image of the tth frame after Gaussian model tile rasterization rendering, and β is the preset depth map difference threshold.

[0085] That is, in the Gaussian rendering and training phase, the Gaussian field is initialized based on the RGB image. In addition to the conventional attributes (such as coordinates x, y, z, color c, opacity α and covariance Σ), each Gaussian ball also embeds semantic attributes (such as object label label and component-level features f p Through tile-based rasterization, the RGB images, depth maps, and object label maps of each perspective are rendered. Multiple losses are calculated together with the first mask clustering results generated by the mask clustering process in the previous module. The rendering parameters are optimized through backpropagation, and the object-level Gaussian field is constructed through multiple rounds of iterations.

[0086] S3: The CLIP model and SBERT model are used to convert the object-level vocabulary of interest set by the user into an object vocabulary feature vector. The label corresponding to the object vocabulary feature vector is obtained based on the triple lossless codebook. The Gaussian spheres corresponding to the object-level vocabulary of interest are then filtered from the Gaussian field based on the label. Finally, the object Gaussian formed by all the filtered Gaussian spheres is used as the object-level query result.

[0087] Among them, the method for obtaining the label corresponding to the object vocabulary feature vector based on the triple lossless code book is:

[0088] The argmax function is used to calculate the cosine similarity between the object vocabulary feature vector and each feature vector in the triple lossless codebook, and the label with the largest cosine similarity is used as the label corresponding to the object vocabulary feature vector; at the same time, the Gaussian sphere corresponding to the object-level query result is highlighted in red.

[0089] S4: Input the center coordinates (x, y, z) of all the filtered Gaussian spheres into the multi-layer perceptron to obtain the component-level features corresponding to each Gaussian sphere;

[0090] It should be noted that although the previous modules provide object-level semantic scene understanding, they still lack the fine-grainedness required for accurate scene manipulation. To bridge this gap, step S4 is implemented as follows: first, dense feature maps are extracted from the input color image through fine-grained feature encoding. These high-dimensional feature maps are then fused with the trained geometrically clear object-level Gaussian distribution through implicit feature fields (MLPs), thereby achieving multi-granular scene interpretation by combining local component-level details with global object-level structures, see Figure 1 Component-level feature field building module;

[0091] Before using the MLP network to learn dense high-dimensional component-level semantic features, the present invention first implements a position encoding module. This module maps the input 3D coordinates to a high-dimensional space and generates embedded_cords = Encoding(x, y, z)∈R 297 This process helps preserve rich geometric features and enhances the network's ability to model complex spatial structures.

[0092] This paper uses a four-layer multilayer perceptron (MLP) to learn component-level features, combined with position encoding to optimize local information. By randomly sampling 50% of voxel-level features for MLP training and combining high-dimensional features with Gaussian coordinates, this method integrates component-level semantics into the representation, enabling efficient and information-preserving feature learning in 3D scenes.

[0093] Among them, the training method of the multi-layer perceptron MLP is:

[0094] S41: Use the SAM model to analyze the RGB images of each frame of the current scene Perform preprocessing to obtain the dense segmentation mask corresponding to each frame of RGB image Where p represents the component level, t represents the t-th frame, and i represents the i-th component in the t-th frame RGB image;

[0095] S42: Using the CLIP model to align the dense segmentation masks of each frame of RGB image in the visual-language space Then from the aligned dense segmentation mask Extract visual-language information from the image and obtain component-level features corresponding to each frame of RGB image

[0096] S43: Based on the mask confidence score, fuse the component level features corresponding to all frame RGB images The overlapping component-level features in the final result are the dense component-level features corresponding to each pixel position of each frame RGB image.

[0097] S44: The depth value Z at each pixel coordinate (u, v) of the rendered depth map obtained after tile rasterization based on the Gaussian field of the current scene is used to project the pixel (u, v) to the world coordinate system (x, y, Z), and then all dense component-level features of the RGB images from different frames corresponding to the pixel coordinate (u, v) are calculated. Assign values ​​to the coordinates (x, y, Z) in the world coordinate system to obtain the high-dimensional features f corresponding to each coordinate in the world coordinate system p , to form a four-tuple (x, y, Z, f p ) form of training set;

[0098] S45: Each three-dimensional coordinate (x, y, Z) and its corresponding four-tuple (x, y, Z, f p ) as the input of the multi-layer perceptron, which outputs the predicted component-level features corresponding to each Gaussian ball in the current scene

[0099] S46: According to the high-dimensional feature f p and predicting part-level features The loss function is constructed by the difference between them, and the parameters of the multilayer perceptron are optimized according to the backpropagation of the loss function until the training reaches the set round and the model loss converges, and a trained multilayer perceptron is obtained.

[0100] It can be seen that in the training process of MLP, unlike object-level segmentation, recognizing fine details in a scene requires a finer-grained level. Therefore, the present invention uses the SAM model, which is famous for its multi-scale feature sensitivity, to train the input Perform preprocessing to obtain dense segmentation masks Next, the present invention extracts the visual-linguistic CLIP features of each mask region to generate compact component-level features Based on the mask confidence score, the features of overlapping parts are fused to finally generate dense component-level features per pixel.

[0101] In other words, to achieve more fine-grained scene understanding, a component-level feature field module is introduced. The SAM model is used to segment the RGB image to generate a component-level mask. The CLIP model is used to extract a dense pixel-level component feature map, which corresponds one-to-one with the depth map pixels rendered by the Gaussian field. The pixel (u, v) is projected to the world coordinate system (x, y, Z) through the depth value Z, and the corresponding component features are matched to generate a training set containing coordinates and lossless features. The training set is input into the multi-layer perceptron (MLP). The MLP predicts the component feature f through the coordinates. p , calculate the loss with the pseudo-truth component feature map extracted by CLIP, back-propagate the optimized parameters, and obtain the accurate component features.

[0102] Furthermore, in order to obtain the data required for training the implicit feature field, the present invention designs a voxel feature extraction process, namely Figure 1 The process of extracting pixel-level component features and constructing a training set to be used as the input of the implicit feature field. The specific implementation details are as follows Figure 2 As shown in Figure 2, voxelized representation effectively combines 2D semantic information with 3D geometric structure, providing richer feature representation for downstream tasks in subsequent training.

[0103] Based on this, we can get the four tuples (x, y, Z, f) corresponding to each three-dimensional coordinate (x, y, Z) in the world coordinate system. p), the four-tuple (x,y,Z,f p ) performing voxel-based feature extraction to obtain a high-dimensional quadruple corresponding to each voxel, and then using the high-dimensional quadruple corresponding to each voxel as the input of the multi-layer perceptron;

[0104] The voxel feature extraction operation includes the following steps:

[0105] Divide the three-dimensional space in the world coordinate system corresponding to the current scene into multiple voxels;

[0106] The three-dimensional coordinates corresponding to each pixel take the voxel closest to itself as its own voxel;

[0107] The four-tuples (x, y, Z, f) corresponding to all three-dimensional coordinates in each voxel are respectively p ) are combined into an overall feature, and each overall feature and the center coordinate of each voxel are combined into a high-dimensional quadruple corresponding to each voxel.

[0108] S5: The CLIP model and the SBERT model are used to convert the component-level vocabulary of interest set by the user into a component vocabulary feature vector. Each Gaussian sphere is colored according to the cosine similarity between the component vocabulary feature vector converted from the component-level vocabulary of interest set by the user and the component-level features corresponding to each Gaussian sphere. The Gaussian formed by the darkest Gaussian sphere is used as the component-level query result. The component-level query result is displayed in the Gaussian field in the form of a heat map. At the same time, when coloring each Gaussian sphere, the Gaussian sphere with a greater cosine similarity with the component vocabulary feature vector has a darker corresponding color.

[0109] For example, for object-level queries (such as "a vase"), CLIP and SBERT are used to embed text into a high-dimensional space. Combining the object-level codebook constructed by the first module and the label attribute of the Gaussian sphere constructed by the second module, the argmax algorithm is used to identify the Gaussian sphere with the best semantic matching and highlight it in red in the scene. For component-level queries (such as "pebbles"), the text is embedded into a high-dimensional space, and the converged MLP trained in the third module is called to regress the component features f corresponding to the Gaussian sphere coordinates within the object. p ,Through cosine similarity, the Gaussian of the most semantically matched component is identified,and highlighted in the form of a heat map, thus achieving multi-granularity and ,accurate open vocabulary query at the object level and the component level.

[0110] This paper introduces a hierarchical retrieval framework that supports object- and component-level queries. For object-level queries, text is encoded into a high-dimensional embedding using CLIP and SBERT. The label attribute in the Gaussian sphere and the object-level codebook constructed in the first module are then used to retrieve the object-level semantic features in the corresponding Gaussian. The argmax algorithm is then used to obtain the high-confidence Gaussian for the object, which is highlighted in red on the map.

[0111] On this basis, to achieve component-level positioning, we first use the same text encoding method to form a high-dimensional embedding, and then use the trained implicit feature field, namely MLP, to obtain the coordinates (x, y, z) of the Gaussian corresponding to the high-confidence object obtained above, and use MLP to regress the high-dimensional component-level features f p , followed by fine-grained similarity computation with text embeddings to identify the semantically closest component Gaussians. Furthermore, semantic segmentation follows the same retrieval paradigm as object-level queries, where semantic labels replace text input for feature similarity computation, thus achieving comprehensive scene understanding by assigning labels to each object Gaussian.

[0112] The advantages of the present invention over the prior art are described in detail below.

[0113] 1. Comparison of 2D & 3D zero-sample segmentation effects

[0114] Baseline Methods: For 2D semantic segmentation, our method is compared with NeRF-based methods (LERF, 3D-OVS) and 3DGS-based methods (LangSplat, Graspsplat). For 3D semantic segmentation, our method uses ConceptGraph (CG), LangSplat, and Graspsplat as benchmarks. For 3D object segmentation, ConceptGraph and GaussianGrouping (GG) are used as baseline methods. The latter is limited to geometric grouping and lacks semantic capabilities.

[0115] Datasets and Evaluation Metrics: We selected two commonly used indoor datasets: eight scenes from Replica and six scenes from ScanNet. Due to the lack of detailed 2D annotations in ScanNet, we performed 3D validation only on the ScanNet dataset. In our experiments, we used the semantic labels provided by the dataset as query text. For semantic evaluation metrics, we used mean intersection over union (mIoU) and mean accuracy (mAcc), while for segmentation tasks, we reported mean average precision (mAP).

[0116] Table 1 Comparison of quantitative results of 2D semantic segmentation

[0117]

[0118] Comparison results: The qualitative and quantitative results of 2D zero-shot semantic segmentation are as follows: Figure 3 and as shown in Table 1. LERF and 3D-OVS use multi-scale or sliding window techniques for feature extraction and rely on a single MLP to regress the feature field of the entire scene, which often leads to blurry segmentation results. Although LangSplat integrates SAM and embeds low-dimensional semantic features into a Gaussian distribution, training the autoencoder and then restoring the high-dimensional features severely affects the effectiveness and completeness of the features, making it difficult to accurately understand objects and capture clear object boundaries. Graspsplat relies on CLIP / DINO features and similarly performs poorly in terms of boundary clarity.

[0119] In contrast, the present invention achieves a scene representation with complete information preservation through its integrated codebook framework, maintaining uncompressed high-dimensional semantic features. Combining the dual-branch feature extraction architecture of CLIP embedding and object description attributes, it can accurately understand objects and fully interpret scenes. Figure 3 It can be seen that the proposed method has achieved significant improvements in zero-shot semantic segmentation compared with the baseline.

[0120] The quantitative results of 3D zero-shot segmentation are shown in Tables 2 and 3. ConceptGraph often leads to insufficient scene segmentation by widely merging objects based on the similarity of different point cloud segments, e.g. Figure 3 and Figure 4 As shown. LangSplat and GraspSplat both use different degrees of feature compression, diluting semantic information, and it is difficult to avoid blurred boundaries even after dimensionality recovery. In contrast, the present invention provides an optimized global object segmentation framework with consistent object label attributes, and achieves accurate 3D semantic segmentation through information-preserving representation. Figure 4 It can be seen that the present invention achieves accurate segmentation with its object-level information preserving understanding, which is a significant improvement compared to the baseline.

[0121] Table 2 Comparison of quantitative results of 3D semantic segmentation

[0122]

[0123] Table 3 Comparison of quantitative results of 3D instance segmentation

[0124]

[0125]

[0126] 2. Comparison of Object-Level Open Vocabulary Retrieval Results

[0127] Baseline Methods: For retrieval evaluation, we use ConceptGraph as the baseline method due to its global object-level representation capabilities. Unlike other 3DGS-based methods, ConceptGraph is able to accurately evaluate recall of the top three search results. We also evaluate its LLM enhancement, ConceptGraph-LLM (CG-L.), which leverages ChatGPT to identify objects with semantically related descriptions, thereby improving retrieval accuracy.

[0128] Dataset and Evaluation Metrics: We conducted experiments on four scenarios in the Replica dataset. The texts used for retrieval are divided into three types:

[0129] 1) Ontology description: directly describe the characteristics of the object, such as an owl sculpture.

[0130] 2) Relevance description: Discuss elements related to the object, such as garbage (i.e., trash can).

[0131] 3) Functional description: emphasizes the function of the item, such as a gift for a loved one (i.e., flowers).

[0132] For each description type, we selected 20 samples and measured the recall of the model when predicting the top 1, top 2, or top 3 results.

[0133] Comparison Results: The results in Table 4 show that our method outperforms ConceptGraph and its variants in all retrieval tasks. Although ConceptGraph-LLM performs better in relevance and functionality queries by leveraging the superior reasoning capabilities of LLM, its performance in ontology queries is limited by the inherent hallucination problem of LLM. Our method addresses this limitation through its dual-branch feature extraction framework, enhancing the model's ability to handle diverse retrieval types, such as Figure 5 As shown. Figure 5 It can be seen that when the results are visualized in 3D point cloud format, under three different types of query texts, the present invention is able to accurately retrieve the most semantically relevant objects.

[0134] Table 4 Comparison of recall results of object-level retrieval

[0135]

[0136] 3. Comparison of component-level fine segmentation and hierarchical retrieval

[0137] To evaluate the effectiveness of our method in component-level understanding and its hierarchical structure, which combines component-level features with object-level Gaussian fields, we designed a comparative experiment focusing on the component-level retrieval task.

[0138] Dataset and Evaluation Metrics: We conducted experiments on four scenes from the Replica dataset. The retrieved text primarily describes the internal components of an object. In hierarchical retrieval, the text consists of two parts: the object description and its component descriptions (e.g., flowerpot - green leaves). In contrast, in direct retrieval, the text is linked by the character "of" to form a single query (e.g., flowerpot - green leaves).

[0139] Implementation details: The hierarchical retrieval process is divided into two stages, which are elaborated in Section 4. In contrast, direct retrieval methods extract part-level features from all Gaussian distributions in the scene, calculate their similarity with a single query text, and visualize the results as a heatmap.

[0140] Comparison and analysis: Quantitative accuracy indicators and visualization results such as Figure 6 As shown in Table 5. The present invention combines object-level Gaussian distribution with component-level feature fields to achieve multi-granularity representation and demonstrate excellent component segmentation and recognition capabilities. In contrast, direct retrieval methods extract dense component-level features and lack clear object boundary perception, which limits the effect of fine-grained recognition. Figure 6 As shown, when the results are also visualized using 3D point clouds, the hierarchical retrieval results based on the present invention have clearer boundaries and more comprehensive semantic understanding.

[0141] Table 5 Performance comparison of different types of component-level retrieval

[0142]

[0143] 4. Multi-granularity scene editing

[0144] This section demonstrates the multi-granularity 3D scene editing capabilities of the present invention. Figure 7 . Each Gaussian embeds the object label attribute, which can be associated with the object-level semantic features through the object-level code book, and the fine-grained component features are regressed using the coordinates of the Gaussian inside the object, thereby achieving accurate positioning and selection of object-level and component-level Gaussians based on user instructions. After determining the target Gaussian sphere, this method supports two editing operations based on the characteristics of the Gaussian sphere: (1) spatial existence editing by removing or retaining the Gaussian sphere to control the addition or deletion of objects; (2) color transformation by adjusting the spherical harmonic function parameters to change the visual appearance of objects or components. Figure 7 The results of locating object-level and component-level Gaussian spheres through instructions and performing the above-mentioned editing operations are demonstrated, verifying the efficiency and accuracy of this method in multi-granularity scene editing. Figure 7 It is shown that the present invention can perform removal and color changes at the object and component level.

[0145] In summary, the present invention introduces an object-level Gaussian field with a globally consistent codebook. After the learnable semantic label vector in the Gaussian field is rendered back to the corresponding object label, a direct mapping between the label and the corresponding uncompressed high-dimensional feature is established through the codebook, thereby supporting semantic features of arbitrary dimensions without the need for additional compression, significantly improving the ability to understand objects.

[0146] The present invention develops an implicit component-level feature field, which directly learns high-dimensional features through a multi-layer perceptron fully connected network, seamlessly integrates fine-grained understanding capabilities into the object-level Gaussian field, and promotes the application of multi-granularity downstream tasks.

[0147] We conduct extensive quantitative and qualitative evaluations across multiple scenarios, demonstrating superior performance in object-level zero-shot segmentation and open vocabulary understanding. Notably, our approach achieves state-of-the-art accuracy in object-part hierarchical retrieval while also supporting multi-granular scene editing.

[0148] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may of course make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A multi-granularity open vocabulary query method based on object-level lossless Gaussian field, characterized by: The following steps are involved: Input multiple RGB images of the current scene into the object-level feature codebook model to obtain a lossless triplet codebook corresponding to all objects in the current scene, where any triplet consists of a label, a description text, and a feature vector. Input multiple RGB images of the current scene into the Gaussian rendering model to obtain the object-level Gaussian field corresponding to the current scene. The attributes of each Gaussian sphere that constitutes the Gaussian field include the center coordinates (x, y, z) of the Gaussian sphere and the object label label corresponding to the Gaussian sphere; The object-level vocabulary of interest set by the user is converted into an object vocabulary feature vector. The label corresponding to the object vocabulary feature vector is obtained according to the triple lossless code book. Then, the Gaussian sphere corresponding to the object-level vocabulary of interest is filtered out from the Gaussian field based on the label. Finally, the object Gaussian formed by all the filtered Gaussian spheres is used as the object-level query result; The center coordinates (x, y, z) of all the filtered Gaussian balls are input into the multi-layer perceptron to obtain the component-level features corresponding to each Gaussian ball; Each Gaussian sphere is colored based on the cosine similarity between the component vocabulary feature vector converted from the user-set component-level vocabulary of interest and the component-level features corresponding to each Gaussian sphere, and the Gaussian formed by the darkest Gaussian sphere is used as the component-level query result.

2. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 1, characterized in that: Input the multi-frame RGB images of the current scene into the object-level feature codebook model to obtain the triplet lossless codebook corresponding to all objects contained in the current scene: The CropFormer model is used to identify and segment instances in multiple RGB images and generate object segmentation masks corresponding to each RGB image frame. Where t represents the t-th frame, i represents the i-th object in the t-th frame RGB image; Segment the object mask corresponding to each frame of RGB image The visual embedding branch uses the CLIP model to perform contrast alignment in the visual-language space to obtain the visual language features of the object segmentation mask. Extract visual-linguistic information and generate visual embeddings corresponding to each frame of RGB image At the same time, the language reasoning branch uses the frozen image encoder of the TAP model to generate descriptive text captions containing color, material, attributes, and spatial relationships. t,i , and then use the SBERT model to cap the description text corresponding to each frame RGB image t,i Encoded as language feature vector Segment the object mask Visual Embedding Description text cap t,i , language feature vector The object-level high-dimensional features corresponding to each frame of RGB image; Perform mask clustering on the object-level high-dimensional features corresponding to each frame of RGB image to obtain the initial triplet corresponding to each frame of RGB image, where the initial triplet is label-description text-feature vector; The description texts in the initial triples corresponding to the RGB images of all frames are clustered among the frames to obtain the final triplet lossless codebook corresponding to all objects in the current scene.

3. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 2, characterized in that: The loss function L used when training the Gaussian rendering model is as follows: in, is the RGB image L1 loss function, λ1 is the weight of the RGB image L1 loss function, is the RGB image structure similarity loss function, λ2 is the weight of the RGB image structure similarity loss function, is the object label loss function, λ3 is the weight of the object label loss function, is the depth map loss function, and λ4 is the weight of the depth map loss function.

4. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 3, characterized in that: Depth map loss function The calculation method is as follows: in, is the true depth map corresponding to the t-th frame RGB image of the current scene, is the depth image of the tth frame after Gaussian model tile rasterization rendering, and β is the preset depth map difference threshold.

5. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 1, characterized in that: The training method of the multilayer perceptron is: Use the SAM model to analyze the RGB images of each frame of the current scene Perform preprocessing to obtain the dense segmentation mask corresponding to each frame of RGB image Where p represents the component level, t represents the t-th frame, and i represents the i-th component in the t-th frame RGB image; The CLIP model is used to align the dense segmentation masks of each frame of RGB images in the visual-language space. Then from the aligned dense segmentation mask Extract visual-language information from the image and obtain component-level features corresponding to each frame of RGB image Based on the mask confidence score, the component level features corresponding to all frame RGB images are fused The overlapping component-level features in the final result are the dense component-level features corresponding to each pixel position of each frame RGB image. The depth value Z at each pixel coordinate (u, v) of the rendered depth map obtained after tile rasterization based on the Gaussian field of the current scene projects the pixel (u, v) to the world coordinate system (x, y, Z), and then all dense component-level features of the RGB images from different frames corresponding to the pixel coordinate (u, v) are calculated. Assign values ​​to the coordinates (x, y, Z) in the world coordinate system to obtain the high-dimensional features f corresponding to each coordinate in the world coordinate system p , to form a four-tuple (x,y,Z,f p ) form of training set; Each three-dimensional coordinate (x, y, Z) and its corresponding four-tuple (x, y, Z, f p ) as the input of the multi-layer perceptron, which outputs the predicted component-level features corresponding to each Gaussian ball in the current scene According to the high-dimensional feature f p and predicting part-level features The loss function is constructed by the difference between them, and the parameters of the multilayer perceptron are optimized according to the backpropagation of the loss function until the training reaches the set round and the model loss converges, and a trained multilayer perceptron is obtained.

6. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 5, characterized in that: Get the four tuples (x, y, Z, f) corresponding to each three-dimensional coordinate (x, y, Z) in the world coordinate system p ), the four-tuple (x,y,Z,f p ) performing voxel-based feature extraction to obtain a high-dimensional quadruple corresponding to each voxel, and then using the high-dimensional quadruple corresponding to each voxel as the input of the multi-layer perceptron; The voxel feature extraction operation includes the following steps: Divide the three-dimensional space in the world coordinate system corresponding to the current scene into multiple voxels; The three-dimensional coordinates corresponding to each pixel take the voxel closest to itself as its own voxel; The four-tuples (x, y, Z, f) corresponding to all three-dimensional coordinates in each voxel are respectively p ) are combined into an overall feature, and each overall feature and the center coordinate of each voxel are combined into a high-dimensional quadruple corresponding to each voxel.

7. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 1, characterized in that: The method for obtaining the label corresponding to the object vocabulary feature vector based on the triple lossless code book is: The argmax function is used to calculate the cosine similarity between the object vocabulary feature vector and each feature vector in the triple lossless codebook, and the label with the largest cosine similarity is used as the label corresponding to the object vocabulary feature vector; At the same time, the Gaussian sphere corresponding to the object-level query result is highlighted in red; the component-level query result is displayed in the Gaussian field in the form of a heat map.

8. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 1, characterized in that: The CLIP model and SBERT model are used to convert the object-level vocabulary of interest set by the user into object vocabulary feature vectors, and also to convert the component-level vocabulary of interest set by the user into component vocabulary feature vectors.

9. The multi-granularity open vocabulary query method based on object-level lossless Gaussian field according to claim 1, characterized in that: When assigning colors to the Gaussian balls, the Gaussian balls with greater cosine similarity to the component vocabulary feature vectors have darker corresponding colors.

Citation Information

Cited By

  • Open vocabulary 3D reasoning Gaussian sputtering method based on implicit text query

    CN120932241A

  • Open vocabulary 3d object query method based on semantic guided geometric reasoning

    CN122492833A

  • An Open-Lexicographic 3D Object Query Method Based on Semantic-Guided Geometric Reasoning

    CN122492833B