Indoor point cloud scene open type semantic segmentation method based on text prompt

By combining local information with global context modeling and text prompts, the problem of unknown category recognition in existing point cloud scene segmentation models is solved, efficient open semantic segmentation is achieved in complex indoor scenes, and the applicability and segmentation accuracy of the model are improved.

CN120612489AActive Publication Date: 2025-09-09HANGZHOU NORMAL UNIVERSITY

Patent Information

Application Number
CN202511106461.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-09
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing neural network models can recognize semantic categories in training set scenarios but have difficulty segmenting unknown categories, which limits their applicability and generalization in complex scenarios. In addition, the irregularity, sparsity and complexity of point cloud data affect the segmentation effect.

Method used

By combining local information aggregation with global context modeling, using text prompts to generate high-level semantic information of the scene, and combining 3D U-Net and Mamba structures to extract point cloud features, the image caption generation model and the pre-trained CLIP text encoder are used for feature contrast learning to achieve open semantic segmentation of point cloud scenes.

Benefits of technology

It achieves effective segmentation of categories not seen during training in large-scale indoor point cloud scenes, improves the generalization ability and segmentation accuracy of the model, and maintains the integrity of the scene structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612489A_ABST
    Figure CN120612489A_ABST
Patent Text Reader

Abstract

The invention relates to an indoor point cloud scene open type semantic segmentation method based on text prompt. On the basis of large-scale indoor scene point cloud data, a 3D U-Net network structure based on Mama blocks is adopted to extract local features of a point cloud scene, global features of the scene point cloud data are extracted through multi-stage Mama blocks and down-sampling, and scene detail information is recovered through up-sampling and jump connection to fuse the local-global features of the point cloud data; generating a text title prompt of a scene multi-view view, associating a projection matrix between a 2D view and a 3D scene with a point cloud scene, enabling text prompt features to be aligned with corresponding point cloud data features, and loading a category text embedded weight into a scene segmentation head to perform a semantic segmentation task; in network training, binary classification loss is added on the basis of semantic segmentation loss so as to balance the recognition and understanding capability of the scene semantic segmentation network on a basic category and a new category; and finally, open semantic segmentation of the indoor point cloud scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual artificial intelligence and relates to a semantic segmentation method for indoor point cloud scenes based on text prompts. Background Art

[0002] Semantic segmentation of point cloud scenes is a fundamental task in 3D scene understanding, which aims to distinguish the semantic categories of scene objects, identify their locations, and infer their geometric properties from 3D scan data. However, neural network models trained on manually annotated datasets are typically only able to understand the semantic categories of scene objects in that dataset. Specifically, they can identify the semantic categories of objects contained in the training set scene but struggle to segment and understand the semantic categories of objects unseen in the training set scene. This significantly limits their applicability and generalization across diverse and complex scenes.

[0003] Contrastive learning based on multimodal features will provide a new solution for open semantic segmentation of point cloud scenes. By effectively utilizing scene text prompts as auxiliary information for scene understanding, the semantic segmentation network will be able to identify unknown semantic categories in the training data set. As a modal information that can provide high-level semantics of the scene, scene text prompts can provide prior knowledge such as scene object categories or scene semantics, thereby making up for the deficiency of scene point cloud data that only contains geometric semantic information. The present invention uses a large-scale pre-trained language model to generate a text description of high-level semantic information of the scene, which adds additional semantic feature information to the scene point cloud and assists the semantic segmentation network in learning fine-grained scene semantic feature representations. In addition, the introduction of scene text features will help reduce the segmentation network's dependence on a large amount of labeled data and improve the generalization ability of the model.

[0004] In order to overcome the defects of large-scale point cloud scene semantic segmentation such as the irregularity of scenic spot cloud data, the sparsity of point cloud data, and the complexity of the scene, a semantic segmentation method for indoor point cloud scenes based on text prompts is proposed. By combining the local information aggregation of scene point cloud data with global context modeling, the geometric features of scene point cloud data are effectively enhanced and the open semantic segmentation of indoor point cloud scenes is achieved. Summary of the Invention

[0005] The purpose of the present invention is to provide an open semantic segmentation method for indoor point cloud scenes based on text prompts, which can realize the global context information capture of point cloud scenes with linear complexity, by combining the local information aggregation of scene point cloud data with global context modeling, and using category text embedding to achieve open semantic segmentation of large-scale indoor point cloud scenes.

[0006] Based on large-scale scene point cloud data input by users, the present invention uses the SoftGroup method to voxelize the scene point cloud data and expand the voxel point cloud into a one-dimensional sequence. Then, the Hilbert space curve is used to map the high-dimensional space into a one-dimensional space through a recursive segmentation strategy. Secondly, the 3D U-Net and Mamba structures are combined in the network encoder to extract the global features of the scene point cloud data through multi-level Mamba blocks and downsampling, and the scene detail information is restored through upsampling and jump connections to fuse the local-global features of the point cloud. Then, with the help of the projection correspondence between the scene view and the scene point cloud, the image caption generation model ViT-GPT2 is used to generate text title prompts for multiple views of the scene. Then, the pre-trained CLIP text encoder is used to embed text features of the multi-view text title prompts, and the point cloud text adapter is used to align the point cloud-text features. The weights of the category text embedding are loaded into the point cloud semantic segmentation head, finally realizing open semantic segmentation of the point cloud scene.

[0007] The specific steps include:

[0008] Step 1: Obtain 2D view data of the scene from the 3D scene, generate text title prompts for the multi-view view scene, and then associate the scene point cloud with the scene text title;

[0009] Furthermore, step one is specifically as follows: obtaining multi-view projection views from indoor scene point cloud data by utilizing the projection matrix between the scene 2D view and the 3D scene, and then generating text title prompts for the scene multi-view images by using the image caption generation model ViT-GPT2, and then associating the scene with the point cloud scene with the help of the projection matrix between the 2D view and the 3D scene, and then associating the scene point cloud with the scene text title.

[0010] Step 2: Embed the Mamba block into the U-Net framework to build a point cloud scene semantic segmentation network. Use a binary classification loss based on the semantic segmentation loss to balance the scene semantic segmentation network's ability to recognize and understand basic categories and new categories.

[0011] Step 3: Use the Hilbert space curve to serialize the point cloud 3D structure and preserve its spatial relationship; the point cloud scene semantic segmentation network trained in step 2 extracts and fuses the local-global features of the point cloud;

[0012] Specifically, the unordered voxel point cloud is converted into a 1D serialized structure suitable for Mamba modeling while preserving its 3D spatial relationship. The point cloud scene semantic segmentation network trained in step 2 uses a 3D U-Net network structure based on Mamba blocks to extract local features of the point cloud scene. The global features of the scene point cloud data are extracted through multi-level Mamba blocks and downsampling. The scene details are then restored through upsampling and skip connections to fuse the local and global features of the point cloud.

[0013] Step 4: Extract text title hints to generate text embeddings and align point cloud-text features. Load the weights of the category text embeddings into the scene segmentation head for semantic segmentation tasks; ultimately, open semantic segmentation of large-scale indoor point cloud scenes is achieved.

[0014] Furthermore, we use the pre-trained CLIP text encoder Transformer to generate text embeddings for the text title prompts and use the point cloud-text adapter to align the point cloud-text features.

[0015] Traditional point cloud scene semantic segmentation methods only use dataset semantic labels for network training and semantic segmentation, while the method in this invention uses feature comparison learning between scene point clouds and text prompts to train the semantic segmentation network, ultimately achieving open semantic segmentation of large-scale indoor point cloud scenes.

[0016] Furthermore, the proposed point cloud scene semantic segmentation network, based on the U-Net framework, embeds Mamba blocks into the two convolutional layers at each level of the U-Net framework's downsampling process. This allows for global features of the scene point cloud to be extracted at different levels and local details to be recovered through skip connections. The Mamba blocks consist of a Hilbert curve layer, a linear projection layer, a 1D convolutional layer, a selective state space (SSM) block, skip connections, and a reshape block.

[0017] Furthermore, in the process of using the Hilbert space curve to convert the disordered voxel point cloud into a one-dimensional serialized structure suitable for Mamba modeling, the SoftGroup method is first used to perform a voxelization operation on the scene point cloud data and to unfold the voxel point cloud into a one-dimensional sequence. The Hilbert space curve is then used to map the high-dimensional space to the one-dimensional space through a recursive segmentation strategy, so that adjacent sampling points in the space remain adjacent in the unfolded one-dimensional sequence. Then, the Hilbert index is calculated for each sampling point, and the sampling points are sorted according to the index, finally obtaining a one-dimensional sequence that maintains the original point cloud space structure.

[0018] In the open semantic segmentation network for indoor point cloud scenes proposed in this paper, a 3D U-Net network structure based on Mamba blocks is used to extract local features of point cloud scenes, and global features of scene point cloud data are extracted through multi-level Mamba blocks and downsampling.

[0019] It can capture global contextual information of point cloud scenes with linear complexity. This is achieved by combining local information aggregation of scene point cloud data with global context modeling and using categorical text embedding to achieve open semantic segmentation of indoor point cloud scenes. The method provided by this invention uses multimodal features such as scene point cloud data and textual prompts to enhance the semantic segmentation network's understanding of open scenes through feature contrast learning, thus achieving open semantic segmentation of point cloud scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0021] Figure 2 This is an example process of converting scene point cloud data into text prompts in an embodiment;

[0022] Figure 3 Schematic diagram of the structure of the Mamba block in the point cloud scene semantic segmentation network;

[0023] Figure 4 Schematic diagram of the point cloud scene semantic segmentation network;

[0024] Figure 5 This is an example diagram of the semantic segmentation effect of point cloud data of a conference room scene in an embodiment;

[0025] Figure 6 This is an example diagram of the semantic segmentation effect of point cloud data for an office scene in the embodiment.

[0026] Figure 7 13 categories of color legends in the embodiment. DETAILED DESCRIPTION

[0027] The technical method and semantic segmentation effect of the present invention are further described and illustrated below with reference to the accompanying drawings.

[0028] like Figure 1 As shown in the figure, a method for semantic segmentation of indoor point cloud scenes based on text prompts consists of two branches: one branch extracts text features based on text title prompts obtained from multi-view views of the scene; the other branch performs voxelization on the input scene point cloud data and uses a visual backbone network (point cloud scene semantic segmentation network) to extract point cloud features; then, through point cloud-text feature alignment, the category text embedding weights obtained by encoding the category name through the encoder are used as the semantic segmentation head, finally achieving open semantic segmentation of large-scale scene point cloud data. Specifically, it includes the following steps:

[0029] Step 1: Use the projection matrix between the scene 2D view and the 3D scene to obtain multi-view projection views from the scene point cloud data. Then, use the image caption generation model ViT-GPT2 to generate text title prompts for the multi-view images of the scene, thereby associating the scene point cloud with the scene text title.

[0030] In this embodiment, the Scenes No. The image is input, and the pre-trained image caption generation model ViT-GPT2 can generate its corresponding language description ,in Generate a model for image captions. The generated results are as follows Figure 2 As shown in the figure, "a large room with tables and a chair in it" means "a large room with tables and chairs in it", "a room filled with board, desks, chairs, and a projector" means "a room with a blackboard, desks, chairs and a projector", and "a room filled with board, chairs,and tables" means "a room with a blackboard, desks and chairs".

[0031] Then, the projection matrix between the 2D view and the 3D scene is used to associate the point cloud scene, and then the scene point cloud is associated with the scene text title.

[0032] The projection equation from world coordinates to image coordinates in this embodiment is: , where the projection matrix ,in , K is the camera intrinsic parameter matrix, is the rotation matrix, is the translation vector, is the 3D point coordinate, are plane coordinates.

[0033] Step 2: Embed the Mamba block into the U-Net framework to build a point cloud scene semantic segmentation network. Use a binary classification loss based on the semantic segmentation loss to balance the scene semantic segmentation network's ability to recognize and understand basic categories and new categories.

[0034] The loss function used for scene point cloud semantic segmentation in this embodiment is ;

[0035] Semantic segmentation loss The semantic loss of each sampling point is calculated using the semantic label of the base category, where is the semantic score, is the semantic label, σ is the Softmax function, is the semantic segmentation head, is the point-by-point cloud feature;

[0036] Point cloud text contrast loss , then the global features of the scene point cloud are obtained through pooling operation and the contrast loss between the point cloud features and the matching text prompt features is calculated, where Indicates traversing point cloud-text paired samples and calculating point cloud features The text features that match it The similarity of Indicates traversing all text features for each fixed Calculating point cloud features With all possible text features The similarity of , is the set of points associated with the text, is the feature of the text-related point set, is a learnable temperature constant, is the number of title-hint pairs for a given point cloud;

[0037] Binary cross entropy loss Used to distinguish between the basic category and the new category, where the prediction score represents the probability that a sample point belongs to a new category, is the predicted label (1 for the base class and 0 for the new class).

[0038] like Figure 3 As shown in the figure, the Mamba block consists of a Hilbert curve layer, a linear projection layer, a 1D convolutional layer, a selective state-space SSM block, a skip connection, and a reshape block. The state-space SSM block in Mamba imposes a structured form on the state matrix and introduces a specific algorithm. Specifically, it uses the High-Order Polynomial Projection Operator (HIPPO) to construct and initialize the state matrix, thereby building a deep sequence model with efficient long-range reasoning capabilities.

[0039] like Figure 4As shown, the point cloud scene semantic segmentation network takes voxelized point cloud data as input. Based on the U-Net architecture, the scene semantic segmentation network embeds Mamba blocks into the two convolutional layers at each level of the U-Net architecture's downsampling process. This extracts global features of the point cloud at different levels while preserving local details through skip connections. While traditional convolutional methods focus only on local regions, the Mamba state-space model excels at processing long sequences. It helps capture global dependencies in large-scale point cloud data and has linear complexity compared to the Transformer.

[0040] Step 3: Use the Hilbert space curve to serialize the point cloud 3D structure and preserve its spatial relationship; the point cloud scene semantic segmentation network trained in step 2 extracts and fuses the local-global features of the point cloud;

[0041] Specifically, the Hilbert space curve is used to convert the disordered voxel point cloud into a 1D serialized structure suitable for Mamba modeling while preserving its 3D spatial relationship. The point cloud scene semantic segmentation network trained in step 2 uses a 3D U-Net network structure based on Mamba blocks to extract local features of the point cloud scene. The global features of the scene point cloud data are extracted through multi-level Mamba blocks and downsampling. The scene details are then restored through upsampling and skip connections to fuse the local and global features of the point cloud.

[0042] The specific process of converting an unordered voxel point cloud into a one-dimensional serialized structure suitable for Mamba modeling is as follows: first, the scene point cloud data is voxelized using the SoftGroup method and the voxel point cloud is unfolded into a one-dimensional sequence. Then, the Hilbert space curve is used to map the high-dimensional space into a one-dimensional space through a recursive segmentation strategy, which ensures that adjacent sampling points in the space remain adjacent in the unfolded one-dimensional sequence. Then, the Hilbert index is calculated for each sampling point and the sampling points are sorted according to the index, finally obtaining a one-dimensional sequence that maintains the original point cloud spatial structure.

[0043] Step 4: Extract text title hints to generate text embeddings and align point cloud-text features. Load the weights of the category text embeddings into the scene segmentation head for semantic segmentation tasks; ultimately, open semantic segmentation of large-scale indoor point cloud scenes is achieved.

[0044] In this embodiment, the pre-trained CLIP text encoder Transformer is used to generate text feature embeddings based on text title prompts, and the point cloud-text adapter is used to align scene point cloud-text features. Specifically: First, the pre-trained CLIP text encoder is used to generate a 512-dimensional vector for the text title prompt of each scene projection view. : , where t is the input text title prompt, A text encoder for CLIP is constructed; then, an MLP adapter is constructed to map 3D point cloud features to the text feature space to achieve feature alignment between point cloud and text modalities; finally, the weights of the category text feature embedding are loaded into the scene segmentation head for scene semantic segmentation.

[0045] like Figure 5 and Figure 7 As shown in the figure, the semantic segmentation effect of the conference room scene point cloud data in the S3DIS dataset is given. Figure 5 In the figure a is the input conference room scene point cloud data, Figure 5 b in the middle is the real background category of the conference room scene point cloud data. Figure 5 Figure c is an example of the semantic segmentation effect of the conference room point cloud scene achieved using the above method. Figure 5 It can be seen that the indoor point cloud scene semantic segmentation method proposed in the present invention can effectively segment the conference room scene point cloud data, and its segmentation result can maintain the complete structure of the scene. At the same time, the boundaries between different segmented objects are relatively clear; from the real background category and semantic segmentation effect example diagram, it can be seen that the chairs, tables, ceilings, floors, blackboards, walls, pillars and other categories in the conference room point cloud scene can be effectively segmented.

[0046] like Figure 6 and Figure 7 As shown in the figure, the semantic segmentation effect of the office scene point cloud data in the S3DIS dataset is given. Figure 6 In the figure a, it is the input point cloud data of the office scene. Figure 6 b in the middle is the real background category of the office scene point cloud data. Figure 6 Figure c is an example of semantic segmentation of an office point cloud scene using the above method. Figure 6 As can be seen, the semantic segmentation method for indoor point cloud scenes proposed in this paper can effectively segment office scene point cloud data. As can be seen from the example image comparing real background categories and semantic segmentation results, most architectural elements in the office scene are accurately segmented while maintaining their structural integrity. The proposed method not only effectively segments visible categories during training, but also accurately segments new, unseen categories during training, ultimately achieving open semantic segmentation for large-scale indoor point cloud scenes.

Claims

1. An open semantic segmentation method for indoor point cloud scenes based on textual cues, characterized by: The specific steps include: Step 1: Obtain 2D view data of the scene from the 3D scene, generate text title prompts for the multi-view view scene, and then associate the point cloud text title; Step 2: Embed the Mamba block into the U-Net framework to build a point cloud scene semantic segmentation network. Use a binary classification loss based on the semantic segmentation loss to balance the scene semantic segmentation network's ability to recognize and understand basic categories and new categories. Step 3: Use the Hilbert space curve to serialize the point cloud 3D structure and preserve its spatial relationship; the point cloud scene semantic segmentation network trained in step 2 extracts and fuses the local-global features of the point cloud; Step 4: Extract text title hints to generate text embeddings and align point cloud-text features. Load the weights of the category text embeddings into the scene segmentation head for semantic segmentation tasks; ultimately, open semantic segmentation of large-scale indoor point cloud scenes is achieved.

2. The open semantic segmentation method for indoor point cloud scenes based on textual prompts according to claim 1, characterized in that: The step 1 obtains a multi-view projection view from the indoor scene point cloud data by using the projection matrix between the scene 2D view and the 3D scene, and then generates a text title prompt of the scene multi-view image by using the image caption generation model ViT-GPT2, and then associates it with the point cloud scene with the projection matrix between the 2D view and the 3D scene, and then associates the point cloud text title.

3. The open semantic segmentation method for indoor point cloud scenes based on textual prompts according to claim 1, characterized in that: Based on the U-Net network framework, the Mamba block is embedded into the two convolutional layers at each level in the U-Net framework downsampling process, so as to extract the global features of the scene point cloud at different levels and restore its local detail information through skip connections.

4. The open semantic segmentation method for indoor point cloud scenes based on textual prompts according to claim 1 or 3, characterized in that: The Mamba block consists of a Hilbert curve layer, a linear projection layer, a 1D convolution layer, a selective state space SSM block, a skip connection, and a reshape block.

5. The open semantic segmentation method for indoor point cloud scenes based on textual cues according to claim 1, characterized in that: The step three specifically includes: converting the disordered voxel point cloud into a one-dimensional serialized structure suitable for Mamba modeling and retaining its point cloud 3D spatial relationship; the point cloud scene semantic segmentation network trained in step two uses a 3D U-Net network structure based on Mamba blocks to extract local features of the point cloud scene, and extracts global features of the scene point cloud data through multi-level Mamba blocks and downsampling, and then restores scene detail information through upsampling and skip connections to fuse the local and global features of the point cloud.

6. The open semantic segmentation method for indoor point cloud scenes based on textual prompts according to claim 5, characterized in that: In the process of using the Hilbert space curve to convert the disordered voxel point cloud into a one-dimensional serialized structure suitable for Mamba modeling, the SoftGroup method is first used to perform a voxelization operation on the scene point cloud data and to unfold the voxel point cloud into a one-dimensional sequence. The Hilbert space curve is then used to map the high-dimensional space into a one-dimensional space through a recursive segmentation strategy, which ensures that adjacent sampling points in the space remain adjacent in the unfolded one-dimensional sequence. Then, the Hilbert index is calculated for each sampling point, and the sampling points are sorted according to the index, ultimately obtaining a one-dimensional sequence that maintains the original point cloud spatial structure.

7. The open semantic segmentation method for indoor point cloud scenes based on textual prompts according to claim 1, characterized in that: In the process of extracting text embeddings in step 4, the pre-trained CLIP text encoder Transformer is used to generate text embeddings for the text title prompts and the point cloud-text adapter is used to align the point cloud-text features. The weights of the category text embeddings are loaded into the scene segmentation head to perform the semantic segmentation task; finally, open semantic segmentation of large-scale indoor point cloud scenes is achieved.

Citation Information

Patent Citations

  • Open vocabulary three-dimensional scene understanding method based on bimodal interaction

    CN118606900A

  • Three-dimensional point cloud semantic segmentation method and device, electronic equipment and storage medium

    CN118898717A

  • Multimodal feature embedded indoor three-dimensional scene understanding method and terminal

    CN118968271A

  • Three-dimensional open vocabulary semantic segmentation method based on cross-modal mask interaction

    CN119006822A

  • Three-dimensional scene digital fusion method for internal and external field environmental characteristics of intelligent test

    CN119338976A

Cited By

  • 3D scene semantic segmentation system and method for open vocabularies

    CN121236762A