Training method and device for three-dimensional open vocabulary semantic segmentation model
By combining a multi-stage visual language model and a cross-frame consistency module, a local-to-global 3D semantic segmentation framework is constructed, which solves the problem of insufficient generalization ability of traditional 3D semantic segmentation in open-world environments and improves the stability and accuracy of semantic segmentation.
Patent Information
- Application Number
- CN202510880031.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional 3D semantic segmentation tasks heavily rely on predefined dataset labeling systems, resulting in a lack of generalization ability in open-world environments. Furthermore, existing methods ignore image semantic information and inconsistencies in multi-view features, affecting stability and accuracy.
A multi-stage visual language model is used to generate a list of target words and point-by-point text labels. The model is pre-trained through a sparse encoder-decoder structure, and a cross-frame consistency module is introduced to construct a three-dimensional semantic segmentation framework from local to global. The LLaVA-NeXT and CLIP models are used for visual-language alignment and feature matching.
It significantly improves the stability and accuracy of semantic labels in real-world scenarios for 3D semantic segmentation models, enhances the understanding of open vocabulary, solves the problem of inconsistent features from multiple perspectives, and achieves higher generalization ability and robustness.
Smart Images

Figure CN120997837A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure belongs to the technical field of three-dimensional scene understanding, and particularly relates to a training method and device of a three-dimensional open-vocabulary semantic segmentation model. BACKGROUND
[0002] Three-dimensional semantic segmentation is a key technology in three-dimensional scene understanding because it can assign semantic labels to each point in a three-dimensional space, and has wide application in the technical fields of autonomous driving, robot navigation, augmented reality, etc.
[0003] Traditional three-dimensional semantic segmentation tasks heavily rely on pre-defined and fixed data set label systems during training. This "closed-vocabulary" setting makes the model lack effective generalization ability when facing unseen classes or open-world environments, which seriously restricts its practicality and expansibility.
[0004] Current three-dimensional semantic segmentation methods for open vocabulary usually rely on full-scene multi-view images as intermediates to transfer image-level or pixel-level semantic information to three-dimensional point clouds. Although this "global training" paradigm has made some progress, it still has the following problems, which affect the stability and accuracy performance in real scenes:
[0005] On the one hand, this kind of method ignores the rich semantic information carried by the image itself, and does not fully exploit the representation potential of the image itself; on the other hand, due to the complex factors such as illumination difference, occlusion relationship and view offset between different views, the multi-view image features corresponding to the same three-dimensional point often lack consistency in content and expression, thereby introducing ambiguity and noise in the feature fusion stage. SUMMARY
[0006] Embodiments of the present disclosure propose a three-dimensional open-vocabulary semantic segmentation scheme to solve the problem that current schemes have poor stability and accuracy performance in real scenes due to following the "global training" paradigm.
[0007] A first aspect of embodiments of the present disclosure provides a training method of a three-dimensional open-vocabulary semantic segmentation model, comprising:
[0008] obtaining multi-view RGB-D images of a target region, for each of the images, performing multi-stage inference through a visual language model, generating a target vocabulary list and prompting a two-dimensional segmentation model to establish a pixel-level text label, generating a first point cloud by depth mapping the image, and generating a point-by-point text label by mapping the text label to the first point cloud;
[0009] Pre-training a neural network model with a sparse encoder-decoder structure using the point-wise text label as a supervision signal to generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a text semantic space, wherein the text semantic space is generated by a multimodal encoder text branch encoding the text label;
[0010] For a second point cloud of a complete scene of a target region, a global scene vocabulary and a text embedding are generated by aggregating the target vocabulary list, point feature embeddings of the second point cloud are extracted using the three-dimensional segmentation model, the point feature embeddings are matched with the text embedding with the highest similarity in a shared visual-linguistic feature space to generate a trusted point-text label pair, and the three-dimensional segmentation model is fine-tuned based on the trusted point-text label pair.
[0011] In some embodiments of the present disclosure, the target vocabulary list is generated by performing multi-stage inference on the image of each view by a visual-linguistic model, including:
[0012] The visual-linguistic model first observes the image and describes the image content;
[0013] Then, the target classes in the image are identified to generate the target vocabulary list, wherein the visual-linguistic model is LLaVA-NeXT.
[0014] In some embodiments of the present disclosure, the prompting of the two-dimensional segmentation model to establish pixel-level text labels includes:
[0015] The visual-linguistic model uses the target vocabulary list to prompt the two-dimensional segmentation model for text description;
[0016] In response to the prompt, the two-dimensional segmentation model locates the object classes in the target classes and associates the successfully located object classes with the corresponding pixel-level semantic segmentation mask to establish a pixel-class correspondence, wherein the two-dimensional segmentation model is GroundedSAM.
[0017] In some embodiments of the present disclosure, the learning of visual features aligned with a text semantic space includes:
[0018] Selecting CLIP as a text encoder to obtain a text embedding corresponding to each point in the first point cloud
[0019] Extracting point features from the first point cloud point by point by the three-dimensional segmentation model
[0020] aligning visual features of the first point cloud with clip text space by minimizing an alignment loss, wherein the alignment loss is as follows:
[0021]
[0022] wherein, is the alignment loss, is a point feature, is a text embedding.
[0023] In some embodiments of the present disclosure, the pre-trained total loss function further comprises a consistency loss, which is based on a cosine similarity measure of the directional difference between feature vectors from two first point clouds of consecutive frames, wherein the cosine similarity is defined as follows:
[0024]
[0025] wherein, is the cosine similarity of corresponding points in point clouds P1 and P2, and represent the feature embeddings of corresponding points in P1 and P2, respectively, P1 and P2 are partial point clouds of the first point cloud generated from the RGB-D images of two consecutive frames;
[0026] The consistency loss function is as follows:
[0027]
[0028] wherein, is the consistency loss, is the cosine similarity of corresponding points in point clouds P1 and P2, and represent the feature embeddings of corresponding points in P1 and P2, respectively, P1 and P2 are partial point clouds of the first point cloud generated from the RGB-D images of two consecutive frames, A represents a set of consecutive points matching physical points in the first point cloud.
[0029] In some embodiments of the present disclosure, the formula for calculating the total loss function is as follows:
[0030]
[0031] wherein, is the total loss function, is the alignment loss, is the consistency loss, and λ is a balance hyperparameter for controlling the trade-off between inter-frame consistency and semantic alignment, wherein λ = 0.2.
[0032] In some embodiments of the present disclosure, the generating a trusted point-text label pair by matching the point feature embedding with the text embedding with the highest similarity in the shared visual-linguistic feature space comprises:
[0033] In the shared visual-linguistic feature space, matching the point feature embedding with the text embedding with the smallest feature distance to generate the trusted point-text label pair.
[0034] In some embodiments of the present disclosure, the method further comprises:
[0035] voxelizing the second point cloud and randomly sampling one point from each voxel to form a subset of samples;
[0036] repeating the above steps until each point of the target region is sampled at least once;
[0037] using the three-dimensional segmentation model to estimate the class probability distribution of each point on all the subset of samples, and then averaging the prediction probability of each point;
[0038] assigning a corresponding text label according to the final average distribution of the class probability of each point.
[0039] In some embodiments of the present disclosure, the fine-tuning the three-dimensional segmentation model based on the trusted point-text label pair comprises:
[0040] fine-tuning the three-dimensional segmentation model by optimizing the consistency loss based on the trusted point-text label pair.
[0041] A second aspect of the embodiments of the present disclosure provides a training device of a three-dimensional open vocabulary semantic segmentation model, comprising:
[0042] a generating module configured to obtain multi-view RGB-D images of a target region, for each of the images, perform multi-stage inference by a visual-linguistic model, generate a target vocabulary list and prompt a two-dimensional segmentation model to establish a pixel-level text label, generate a first point cloud by depth mapping the image, and map the text label to the first point cloud to generate a point-by-point text label;
[0043] a pre-training module configured to pre-train a neural network model with a sparse encoder-decoder structure using the point-by-point text label as a supervision signal to generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a text semantic space, and the text semantic space is generated by a multi-modal encoder text branch encoding the text label;
[0044] a fine-tuning module configured to fine-tune the three-dimensional segmentation model by aggregating the target vocabulary list to generate a global scene vocabulary and a text embedding, extracting point feature embeddings of a second point cloud of a complete scene of the target region using the three-dimensional segmentation model, matching the point feature embeddings with the text embedding having the highest similarity in a shared visual-linguistic feature space to generate a trusted point-text label pair, and fine-tuning the three-dimensional segmentation model based on the trusted point-text label pair.
[0045] In summary, the training method of a three-dimensional open vocabulary semantic segmentation model and the training device of a three-dimensional open vocabulary semantic segmentation model provided by the embodiments of the present disclosure effectively improve the stability and accuracy of generating semantic labels in real scenes by constructing a two-stage open vocabulary three-dimensional semantic segmentation framework from local to global. Specifically, the present disclosure significantly improves the model's understanding of open vocabulary by using multi-modal large models such as LLaVA-NeXT and CLIP to perform chain-of-thought understanding and semantic label generation on multi-view images, thereby introducing semantic information carried by the multi-view images themselves in the three-dimensional semantic segmentation task. Secondly, through the process of "image-text description -> target vocabulary list -> GroundedSAM pseudo label", a more fine-grained and more semantically consistent pseudo-supervised signal is obtained, greatly improving the accuracy of point cloud semantic learning. In addition, in order to solve the problem of inconsistent feature expression between multi-views, the present disclosure introduces a cross-frame consistency module to constrain the point features in the overlapping area with cosine similarity, enhancing the stability and robustness of the features. The present disclosure also guides the model to learn local semantics through local dense supervision, and then migrates high-confidence pseudo labels to the global sparse scene for weakly supervised training, thereby realizing the ability transfer from "small semantic density" to "large structure sparsity" and improving the expressiveness in complex scenes. BRIEF DESCRIPTION OF DRAWINGS
[0046] The features and advantages of the present disclosure will be more clearly understood through reference to the following drawings, which are presented as illustrative and should not be construed as limiting the present disclosure, in which:
[0047] Figure 1 is a local-to-global open vocabulary three-dimensional semantic segmentation framework shown by the present disclosure;
[0048] Figure 2 is a flowchart of a training method of a three-dimensional open vocabulary semantic segmentation model according to some embodiments of the present disclosure;
[0049] Figure 3 is a difference example of a partial point cloud generated from a partial scene and a complete point cloud corresponding to a complete scene;
[0050] Figure 4 is an RGB-D image used to pre-train a three-dimensional segmentation model in an embodiment of the present disclosure;
[0051] Figure 5 is Figure 4 a scene point cloud used for fine-tuning the three-dimensional segmentation model in the embodiments;
[0052] Figure 6 is Figure 4 pixel-wise pseudo labels generated by the embodiments in the pre-training stage;
[0053] Figure 7 is Figure 4 point-wise pseudo labels generated by the embodiments in the fine-tuning stage;
[0054] Figure 8 is Figure 4 the open semantic segmentation results of the scene predicted by the trained three-dimensional segmentation model in the embodiments;
[0055] Figure 9 is a schematic diagram of a training device of a three-dimensional open vocabulary semantic segmentation model according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0056] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the relevant disclosure. However, it will be apparent to one skilled in the art that the present disclosure can be practiced without these details. It will be understood that the use of "system", "device", "unit", and / or "module" terminology in this disclosure is used as a method of distinguishing different components, elements, parts, or assemblies in sequential arrangement. However, these terms can be replaced by other expressions as long as the same purpose is achieved.
[0057] It should be understood that when an apparatus, unit, or module is referred to as being "on", "connected to", or "coupled to" another apparatus, unit, or module, it can be directly on, connected, or coupled to the other apparatus, unit, or module, or intervening apparatus, units, or modules can be present, unless the context clearly indicates otherwise. For example, the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0058] The terminology used by the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the description of the present disclosure and the claims, the following terms are intended to have the following definitions unless the context indicates otherwise: "includes" means including but not limited to, "including" means comprising;
[0059] These and other features and characteristics of the present disclosure, the manner of obtaining and producing such a structure and the functions thereof, the related elements of the parts and the combination and interworking of these parts are more fully understood in connection with the following description and the appended drawings, where:
[0060] Various structural diagrams are used in the present disclosure to illustrate various modifications according to embodiments of the present disclosure. It should be understood that the foregoing or the following structures are not intended to limit the present disclosure. The scope of protection of the present disclosure is subject to the claims.
[0061] With the development of three-dimensional sensing technology, three-dimensional scene understanding has gradually become a key component in intelligent systems, which mainly includes three-dimensional target detection, semantic segmentation, and instance segmentation. These tasks rely on the continuous evolution of deep neural network structures and the gradual improvement of large-scale artificial annotation datasets, and have achieved significant performance improvement in recent years. Among these tasks, point cloud semantic segmentation can assign semantic labels to each point in three-dimensional space, thus having dense and fine-grained understanding ability, and has high practical value for downstream applications such as autonomous driving, robot navigation, and augmented reality. However, traditional three-dimensional semantic segmentation tasks are heavily dependent on pre-defined and fixed dataset label systems during training. This "closed vocabulary" setting makes the model lack effective generalization ability when facing unseen classes or open-world environments, which seriously restricts its practicality and expandability.
[0062] In the research of open-vocabulary semantic segmentation in the 2D field, typical methods are usually based on cross-modal alignment mechanisms between images and texts, and can perform semantic recognition on any text category at the pixel level without re-labeling data for specific tasks. The success of this "zero-shot" ability has inspired researchers to transfer it to the three-dimensional field. However, due to the high cost of three-dimensional data annotation and the sparsity of semantic information, the direct application of the above method in the three-dimensional end-to-end faces significant challenges.
[0063] To address this issue, various 3D semantic understanding methods for open vocabularies have emerged in recent years, exploring the feasibility of this emerging direction. Current mainstream methods can be broadly categorized into two paradigms: one extracts image-text aligned semantic features from multi-view images and maps these features onto a 3D point cloud using geometric projection, thereby generating semantic supervision information for each point; the other directly utilizes pre-trained vision-language models to generate text labels from images, and then transfers the labels to the point cloud for segmentation through the geometric correspondence between the view and 3D space. These methods, to some extent, compensate for the scarcity of 3D labels and improve the model's ability to recognize open vocabularies.
[0064] Despite significant progress, current methods generally follow a "global training" paradigm, relying on multi-view images of the entire scene as a semantic bridge to transmit image-level or pixel-level semantic information to 3D point clouds. While structurally sound, this approach suffers from several drawbacks. Firstly, it neglects the inherent contextual semantic information of the images themselves, failing to fully exploit their representational potential. Secondly, due to complex factors such as lighting differences, occlusion relationships, and viewpoint shifts between different perspectives, the multi-view image features corresponding to the same 3D point often lack consistency in content and expression, introducing ambiguity and noise during feature fusion. These issues, to some extent, limit the stability and accuracy of existing methods in real-world scenarios.
[0065] In view of this, this disclosure provides an open lexical three-dimensional semantic segmentation framework from local to global perspectives. For example... Figure 1 As shown. Figure 1 In this context, PGOV3D is the framework described. The framework comprises two stages: pre-training and fine-tuning. In the pre-training stage, local scenes generated from multi-view RGB-D images (depth-enhanced red, green, and blue three-channel color images) through pixel-level depth projection are used as input to the 3D segmentation network. To supervise the training process, pixel-wise pseudo-labels automatically generated from the multi-view RGB-D images are identified using LLaVA-NeXT and a 2D base model as alignment supervision signals. Furthermore, an auxiliary inter-frame consistency module is introduced in this stage to enhance feature consistency. In the fine-tuning stage, the pre-trained 3D segmentation model generates high-confidence point-level pseudo-labels as signals for further training of the 3D model under the same architecture.
[0066] Figure 2 This is a flowchart illustrating a training method for a three-dimensional open-vocabulary semantic segmentation model according to some embodiments of the present disclosure. In some embodiments, the training method for the three-dimensional open-vocabulary semantic segmentation model is executed by a training server, which is used to train the three-dimensional open-vocabulary semantic segmentation model. The method includes the following steps:
[0067] S210, acquire multi-view RGB-D images of the target region, for each of the images, perform multi-stage inference through a visual language model, generate a target vocabulary list and prompt a two-dimensional segmentation model to establish a pixel-level text label, generate a first point cloud by depth mapping the image, and map the text label to the first point cloud to generate a point-by-point text label.
[0068] Specifically, in some embodiments of the present disclosure, first, a pixel-by-pixel pseudo label is generated:
[0069] The present disclosure employs LLaVA-NeXT to identify the target class C′ in each image i . Instead of directly prompting the model with a list of target names, the present disclosure designs a Chain-of-Thought (CoT) strategy to guide the model to generate more accurate and reliable target classes. Before generating the object list, LLaVA-NeXT is first required to carefully observe and describe the image content, thereby having a more detailed understanding of the image. This phased inference process significantly improves the accuracy and completeness of target recognition. Subsequently, the extracted target list is used to prompt a two-dimensional segmentation base model (such as GroundedSAM) to locate the object classes in C′ i , and the successfully located classes C i are associated with the corresponding pixel-level semantic segmentation mask. Thus, an accurate pixel-class correspondence (i.e., pixel-by-pixel pseudo label) is established for supervising the training of the three-dimensional segmentation network.
[0070] Please note that LLaVA-NeXT and GroundedSAM in the above embodiments are only for example. Any visual language model based on large language model technology and having image and text fusion understanding capability is suitable for the present disclosure. Similarly, any two-dimensional segmentation base model is suitable for the present disclosure.
[0071] Re-generate point-wise pseudo labels :
[0072] First, the image is converted into a parsed three-dimensional point cloud P i . Specifically, for each RGB-D image I i , given the camera intrinsic matrix and pose , each pixel (u, v) with depth d(u, v) is mapped to a three-dimensional point p(u, v) in the world coordinate through homogeneous transformation:
[0073] p(u, v) = T i [d (u,v) · K -1 [u, v] T ].
[0074] By applying the formula to all pixels in the RGB-D image, a pixel-level partial point cloud scene is obtained. Then, the pixel-wise pseudo labels are converted into point-wise pseudo labels by the correspondence between the pixel pairs and the three-dimensional points, which are used to supervise the three-dimensional segmentation network.
[0075] S220, pre-training a neural network model with a sparse encoder-decoder structure using the point-wise pseudo labels as supervision signals to generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a text semantic space, and wherein the text semantic space is generated by a multi-modal encoder text branch encoding the text labels.
[0076] The present disclosure aligns the visual-linguistic space on the local point cloud to pre-train the neural network model Pre-training Specifically:
[0077] In the pre-training phase, the point-wise pseudo labels are used as supervision signals to guide the three-dimensional segmentation network with a sparse encoder-decoder structure to learn visual features aligned with an open text space. Specifically, the present disclosure extracts point-wise features For each corresponding text label, since CLIP has an open vocabulary capability and a semantically rich embedding, the present disclosure selects CLIP as the text encoder to obtain the pseudo label embedding corresponding to each point To align the point features with the CLIP text space, the present disclosure uses an alignment loss that minimizes the cosine distance between the point features and the text embedding The formula is as follows:
[0078]
[0079] Through the formula, the point features and the open feature space of CLIP can be aligned, so that the three-dimensional segmentation model has the generalization ability to recognize unseen objects.
[0080] CLIP is used only as an example. In principle, any encoder with open vocabulary understanding ability and capable of mapping image-text cross-modal contrastive learning to the same semantic space can be applicable.
[0081] Some embodiments of the present disclosure also include a cross-frame consistency module to solve the problem of lack of consistency of multi-view image features corresponding to the same three-dimensional point.
[0082] Cross-frame consistency module :
[0083] Since partial RGB-D point clouds are projected from continuous scenes, the same physical points often appear in adjacent frames. Ideally, the feature representation of the same physical points should remain consistent across frames. However, due to factors such as lighting changes, occlusions, and viewpoint changes, the model can learn inconsistent features, which can affect its performance on the complete scene point cloud and lead to unstable semantic segmentation results. To solve this problem, the present disclosure introduces an auxiliary task called the "consistency module" that explicitly enforces feature consistency in overlapping regions under different viewpoints. By encouraging the three-dimensional segmentation network to produce consistent features for the same physical points in multiple views, this module reduces semantic noise caused by conflicting predictions and helps to reconstruct a more coherent three-dimensional spatial representation from fragmented observations. Specifically, let P1 and P2 represent partial point clouds from two consecutive frames, and let A represent the coherent point set containing the matching physical points between these frames. The present invention uses cosine similarity to measure the directional similarity between feature vectors, with a range of [-1, 1], where a higher value indicates stronger consistency. The cosine similarity between two points is defined as follows:
[0084]
[0085] where and represent the feature embeddings of the corresponding points in P1 and P2, respectively. Based on cosine similarity, the present invention defines a three-dimensional consistency loss L consistency to align the features of matching physical points in different frames:
[0086]
[0087] Therefore, the overall training objective of the present invention combines the consistency loss and the alignment loss, encouraging both stable point-level features and aligning them with open-set textual semantics:
[0088]
[0089] where λ = 0.2 is a balance hyperparameter controlling the trade-off between temporal consistency and semantic alignment.
[0090] S230, for the target region complete scene second point cloud, aggregate the target vocabulary list to generate a global scene vocabulary table and a text embedding, use the three-dimensional segmentation model to extract the point feature embedding of the second point cloud, match the point feature embedding with the highest similarity in the text embedding in the shared visual-linguistic feature space, generate a trusted point-text label pair, and fine-tune the three-dimensional segmentation model based on the trusted point-text label pair.
[0091] Although the RGB-D images based on the target region partial scene provide relatively dense local geometric information during pre-training, they are still inherently incomplete due to two major limiting factors, as shown in: Figure 3 (a) inaccurate depth measurements due to surface material properties or color interference with infrared reflection; (b) object occlusions due to camera viewpoint changes. These issues introduce domain differences between the partial scene-level point clouds used in pre-training and the complete scene-level point clouds used in the final target. Therefore, directly applying the pre-trained network to the full-scene segmentation results in suboptimal performance.
[0092] To reduce this difference, the present disclosure introduces a fine-tuning stage that automatically generates high-confidence point-wise text labels from the complete scene using the pre-trained model. The network is then fine-tuned using these pseudo-labels. This process constructs a seamless and unified visual-linguistic embedding space that is aligned with the complete three-dimensional scene geometry. Notably, the fine-tuning stage only requires unannotated scene-level point clouds and does not rely on any predefined class list. Specifically:
[0093] First, generate point-wise pseudo labels composed of reliable point-text label pairs :
[0094] To achieve supervision without human annotation, the present disclosure first extracts a vocabulary C i from all RGB-D frames and aggregates them to form a global scene-level vocabulary c. Each word in the vocabulary c is encoded into an open-text embedding by a CLIP text encoder. At the same time, the present disclosure extracts three-dimensional point embeddings from the complete scene point cloud using a pre-trained three-dimensional segmentation model. Each three-dimensional point in the scene is matched with a word based on its feature distance in the shared visual-linguistic feature space. To improve the reliability of the point-wise pseudo-labels, the present disclosure employs a probability smoothing strategy based on repeated grid sampling. Specifically, first, the point cloud is voxelized, and a point is randomly sampled from each voxel to form a subset. This sampling process is repeated multiple times to ensure that all points in the scene are sampled. The pre-trained three-dimensional segmentation model estimates the class probability distribution of each point on multiple sampling subsets. Then the prediction probabilities of each point are averaged, and the corresponding text label is assigned according to the final average distribution. This process provides a scalable weak supervision mechanism while mitigating the noise introduced by single prediction or low-confidence associations.
[0095] Then, perform visual-linguistic alignment on the panoramic point cloud :
[0096] Using the selected reliable point-text label pairs, the present disclosure optimizes Fine-tune the three-dimensional segmentation model. Unlike the pre-training stage, which operates on dense but partial RGB-D views, this fine-tuning stage utilizes sparse but complete scene-level geometry information. This shift in supervision granularity enables the model to better capture global spatial context and strengthen the alignment between three-dimensional point features and open-vocabulary text labels. As a result, the model exhibits better generalization capability in the open-vocabulary three-dimensional semantic segmentation task.
[0097] One embodiment of the present disclosure is trained on the ScanNet dataset Figure 2 The method described in S210-S230 (hereinafter referred to as the method or PGOV3D) is verified. ScanNet is a large-scale three-dimensional dataset of indoor scenes, focusing on RGB-D data, and is widely used in three-dimensional scene understanding tasks.
[0098] The RGB-D image data used for pre-training in the embodiment is as shown in Figure 4 The scene point cloud data used for fine-tuning is as shown in Figure 5 Figure 4 The image in the middle is obtained by the multi-modal large model LLaVA-NeXT and the visual base model GroundedSAM to obtain pixel-by-pixel pseudo labels, as shown in Figure 6 The pre-trained three-dimensional segmentation model is used to obtain open spatial features of each point in the scene point cloud in the fine-tuning stage, and the CLIP text encoder is used to encode all scene vocabularies to obtain open spatial features of the text. By similarity matching, the nearest text feature of the point feature is obtained to obtain reliable point-by-point pseudo labels, as shown in Figure 7
[0099] The trained three-dimensional segmentation model is used to predict the open semantic segmentation result of the scene, as shown in Figure 8 As can be seen from Figure 8 The method can accurately segment objects in the scene, with clear edge contours and accurate semantics.
[0100] Figure 9 is a schematic diagram of a three-dimensional open-vocabulary semantic segmentation model training device according to some embodiments of the present disclosure. As shown in Figure 9 The three-dimensional open-vocabulary semantic segmentation model training device 900 includes a generation module 910, a pre-training module 920, and a fine-tuning module 930. In the present disclosure, the training function of the three-dimensional open-vocabulary semantic segmentation model is performed by a training server, which is used to train the three-dimensional open-vocabulary semantic segmentation model, wherein:
[0101] The generating module 910 is configured to obtain multi-view RGB-D images of a target region, perform multi-stage inference on each of the images by using a visual language model, generate a target vocabulary list and prompt a two-dimensional segmentation model to establish a pixel-level text label, generate a first point cloud by mapping the images, and map the text label to the first point cloud to generate a point-by-point text label.
[0102] The pre-training module 920 is configured to pre-train a neural network model with a sparse encoder-decoder structure using the point-by-point text label as a supervision signal, and generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a text semantic space, and the text semantic space is generated by a multi-modal encoder text branch encoding the text label.
[0103] The fine-tuning module 930 is configured to aggregate the target vocabulary list to generate a global scene vocabulary table and a text embedding for a second point cloud of a complete scene of the target region, extract point feature embeddings of the second point cloud by using the three-dimensional segmentation model, match the point feature embeddings with the text embedding with the highest similarity in a shared visual-linguistic feature space to generate a reliable point-text label pair, and fine-tune the three-dimensional segmentation model based on the reliable point-text label pair.
[0104] In summary, the training method of the three-dimensional open vocabulary semantic segmentation model and the training device of the three-dimensional open vocabulary semantic segmentation model provided by the embodiments of the present disclosure effectively improve the stability and accuracy of generating semantic labels in real scenes by constructing a two-stage open vocabulary three-dimensional semantic segmentation framework from local to global. Specifically, the present disclosure uses multi-modal large models such as LLaVA-NeXT and CLIP to perform chain-like thinking and semantic label generation on multi-view images, which significantly improves the model's understanding of open vocabulary, thereby introducing the semantic information carried by the multi-view images themselves in the three-dimensional semantic segmentation task. Secondly, through the process of "image-text description -> target vocabulary list -> GroundedSAM pseudo label", a more fine-grained and more semantically consistent pseudo supervision signal is obtained, which greatly improves the accuracy of point cloud semantic learning. In addition, in order to solve the problem of inconsistent feature expression between multi-view images, the present disclosure introduces a cross-frame consistency module to constrain the point features in the overlapping area by cosine similarity, thereby enhancing the stability and robustness of the features. The present disclosure also guides the model to learn local semantics through local dense supervision, and then migrates high-confidence pseudo labels to the global sparse scene for weakly supervised training, thereby realizing the transfer of "small semantic density" to "large structural sparsity" and improving the expressiveness in complex scenes.
[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the devices and modules described above can refer to the corresponding description in the foregoing device embodiments, and will not be repeated here.
[0106] Although the subject matter described herein is provided in the general context of computer-executable instructions of a program module being executed by a computer system on a computing device, those skilled in the art will recognize that other implementations can be performed in combination with other types of program modules. Generally, program modules include routines, programs, components, data structures, and other types of structures that perform particular tasks or implement particular abstract data types. Those skilled in the art will also recognize that the subject matter described herein can be practiced with other computer system configurations, including hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. The described subject matter can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.
[0107] Those ordinarily skilled in the art can appreciate that the units and method steps of the examples described in connection with the embodiments disclosed herein can be realized in electronic hardware, or in a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0108] It should be understood that the foregoing detailed description of the specific implementation of the present disclosure is merely for illustrative or explanatory purposes, and does not constitute a limitation on the present disclosure. Therefore, any modification, equivalent replacement, improvement, etc. made without departing from the spirit and scope of the present disclosure shall be included in the protection scope of the present disclosure. In addition, the appended claims of the present disclosure are intended to cover all variations and modifications falling within the scope and boundary of the appended claims, or the equivalent forms of such scope and boundary.
Claims
1. A method for training a three-dimensional open vocabulary semantic segmentation model, the method comprising: The method comprises: acquiring multi-view RGB-D images of a target region, for each of the images, performing multi-stage inference by a visual language model, generating a target vocabulary list, prompting a two-dimensional segmentation model to establish pixel-level text labels, generating a first point cloud by depth mapping the images, and mapping the text labels to the first point cloud to generate point-by-point text labels; using the point-by-point text labels as a supervision signal, pre-training a neural network model with a sparse encoder-decoder structure to generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a text semantic space, and wherein the text semantic space is generated by a multimodal encoder text branch encoding the text labels; for a second point cloud of a complete scene of the target region, aggregating the target vocabulary list to generate a global scene vocabulary table and text embeddings, extracting point feature embeddings of the second point cloud using the three-dimensional segmentation model, matching the point feature embeddings with the text embeddings with the highest similarity in a shared visual-linguistic feature space to generate a trusted point-text label pair, and fine-tuning the three-dimensional segmentation model based on the trusted point-text label pair.
2. The method of claim 1, wherein, The method further comprises: the multi-stage inference by the visual language model for each of the images to generate the target vocabulary list comprises: the visual language model first observes the image and describes the image content; 3. The method of claim 2, wherein, then identifies target classes in the image to generate the target vocabulary list, wherein the visual language model is LLaVA-NeXT. The method further comprises: the visual language model uses the target vocabulary list to prompt the two-dimensional segmentation model for text description; 4. The method of claim 1, wherein, in response to the prompt, the two-dimensional segmentation model locates object classes in the target classes and associates successfully located object classes with corresponding pixel-level semantic segmentation masks to establish a pixel-class correspondence, wherein the two-dimensional segmentation model is GroundedSAM. selecting CLIP as a text encoder to obtain a text embedding f corresponding to each point in the first point cloud i t ; extracting point features from the first point cloud point by point through the three-dimensional segmentation model The method further comprises: wherein, is an alignment loss, is a point feature, f i t is a text embedding. aligning visual features of the first point cloud with a clip text space by minimizing an alignment loss, wherein the alignment loss is as follows:
5. The method of claim 4, wherein: wherein, is a cosine similarity of corresponding points in point clouds P1 and P2, and denote feature embeddings of corresponding points in P1 and P2, respectively, P1 and P2 are partial point clouds of the first point cloud generated from the RGB-D images of two consecutive frames; the total loss function of the pre-training further comprises a consistency loss, and the consistency loss measures directional differences between feature vectors from two of the first point clouds of consecutive frames based on a cosine similarity, wherein the cosine similarity is defined as follows: wherein, is a consistency loss, is a cosine similarity of corresponding points in point clouds P1 and P2, and denote feature embeddings of corresponding points in P1 and P2, respectively, P1 and P2 are partial point clouds of a first point cloud generated from the RGB-D images of two consecutive frames, and A denotes a set of consecutive points that match a physical point in the first point cloud. the consistency loss function is as follows:
6. The method of claim 5, wherein: wherein, is the total loss function, is the alignment loss, is the consistency loss, and λ is a balancing hyper-parameter to control the trade-off between inter-frame consistency and semantic alignment, where λ = 0.
2.
7. The method of claim 1, wherein, the calculation formula of the total loss function is as follows: The method further comprises:
8. The method of claim 1, wherein, matching the point feature embeddings with the text embeddings with the highest similarity in the shared visual-linguistic feature space to generate the trusted point-text label pair comprises: in the shared visual-linguistic feature space, matching the point feature embeddings with the text embeddings with the smallest feature distance to generate the trusted point-text label pair. The method further comprises: voxelizing the second point cloud and randomly sampling one point from each voxel to form a subset of samples; repeating the above steps until each point of the target region is sampled at least once; estimating a class probability distribution for each point on all the subset of samples using the three-dimensional segmentation model and averaging the predicted probabilities for each point; assigning a corresponding textual label to each point according to the final averaged distribution of class probabilities.
9. The method of claim 6, wherein, the fine-tuning of the three-dimensional segmentation model based on the trusted point-textual label pairs comprises: fine-tuning the three-dimensional segmentation model by optimizing the consistency loss based on the trusted point-textual label pairs. 10.A device for training a three-dimensional open vocabulary semantic segmentation model, characterized in that, comprises: a generation module configured to obtain multi-view RGB-D images of a target region, for each of the images, perform multi-stage inference by a visual language model, generate a target vocabulary list and prompt a two-dimensional segmentation model to establish a pixel-level textual label, generate a first point cloud by depth mapping the image, and map the textual label to the first point cloud to generate a point-by-point textual label; a pre-training module configured to pre-train a neural network model with a sparse encoder-decoder structure using the point-by-point textual label as a supervision signal to generate a three-dimensional segmentation model on the first point cloud, wherein the pre-training includes guiding the model to learn visual features aligned with a textual semantic space, and the textual semantic space is generated by a multi-modal encoder text branch encoding the textual label; a fine-tuning module configured to aggregate the target vocabulary list to generate a global scene vocabulary list and a textual embedding for a second point cloud of a complete scene of the target region, extract point feature embeddings of the second point cloud using the three-dimensional segmentation model, match the point feature embeddings with the most similar textual embedding in a shared visual-linguistic feature space, generate trusted point-textual label pairs, and fine-tune the three-dimensional segmentation model based on the trusted point-textual label pairs.
Citation Information
Cited By
Data labeling method and system
CN121545158A
Data annotation methods and systems
CN121545158B
A three-dimensional point cloud representation segmentation method and system
CN122391579A
A three-dimensional point cloud representation segmentation method and system
CN122391579B