A semantic segmentation method for indoor point cloud scenes based on patch context features

By extracting the patches of indoor scenes and learning the patch features using multi-scale structure and Transformer module, the problem of semantic segmentation of large-scale indoor point cloud scenes is solved, and efficient and accurate semantic segmentation effect is achieved.

CN115620287BActive Publication Date: 2025-05-16HANGZHOU NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211398672.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-05-16
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

The semantic segmentation of large-scale indoor point cloud scenarios faces problems such as huge data scale, uneven data distribution and complex and diverse object shapes, which makes it difficult for traditional methods to effectively process and segment.

Method used

A semantic segmentation method for indoor point cloud scenes based on the context features of the patch are proposed. The scene patch is extracted through the dynamic region growth algorithm, and the local and global features of the patch are learned using the multi-scale structure and the Transformer module, and efficient semantic segmentation is achieved through the encoder-decoder structure.

Benefits of technology

It realizes efficient and accurate semantic segmentation of large-scale indoor point cloud scenarios, especially when dealing with scenes with a large number of repetitive structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620287B_ABST
    Figure CN115620287B_ABST
Patent Text Reader

Abstract

The present invention relates to a semantic segmentation method for indoor point cloud scenes based on facet context features. Based on large-scale scene point cloud data input by users, the present invention uses a dynamic region growing algorithm to extract point clouds with consistent geometric features in indoor scenes as scene faces; in a neural network encoder, a multi-scale structure is used and a facet local feature aggregation module is embedded to effectively aggregate context information of adjacent facets of scenes at different scales; a Transformer module based on a multi-head self-attention mechanism is used to learn the global features of scene faces, and at the same time, the global features of the scene faces are connected and fused with the local features, and a multi-layer perceptron is used to reduce the feature dimensions of the scene faces; then, in a neural network decoder, interpolation upsampling is used to restore the resolution of the downsampled scene faces to the original facet resolution and a semantic label is assigned to each facet, so that large-scale indoor point cloud scenes with a large number of repeated structures can be efficiently and accurately segmented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field and relates to a semantic segmentation method for indoor point cloud scenes based on facet context features. Background Art

[0002] A large amount of point cloud data of indoor scenes (such as large theaters, indoor gymnasiums, and large shopping malls) obtained by 3D scanners and depth cameras usually leads to defects such as large point cloud data scale and uneven distribution due to the characteristics of large scene range, complex scene layout, and many scene objects. These will bring great challenges to the understanding of large-scale indoor point cloud scenes.

[0003] In the semantic segmentation of large-scale point cloud scenes, the traditional method of using discrete point clouds as scene data representation has the following defects: First, due to the huge amount of large-scale scene point cloud data, it is difficult to directly process it when training neural networks; second, due to the complex and diverse types of objects in indoor scenes, it is difficult to establish a unified model to effectively learn and segment different types of shapes, and there are phenomena such as fuzzy segmentation boundaries between objects. Specifically, it is reflected in: 1) The scale of scene point cloud data is too large, which makes it difficult to directly process it when training neural networks; 2) The shapes of three-dimensional objects in indoor point cloud scenes are complex and diverse, ranging from walls, beams, columns, windows to furniture and items. Many of them are difficult to learn and recognize their semantic information using a unified network model; 3) Due to the clutter of scene objects and mutual occlusion between objects in real indoor environments, the segmentation of scene point cloud data becomes more difficult. Therefore, how to characterize large-scale scene point cloud data? How to extract scene data feature information? How to perform effective scene semantic segmentation based on scene feature information? All of these urgently need to propose new ideas and methods.

[0004] In order to effectively overcome the defects of inefficient and time-consuming neural network training and hardware resource consumption in the semantic segmentation of large-scale indoor scene point cloud data, a method is proposed to use scene patches as a representation form of large-scale point cloud data. By extracting patches of indoor scenes and representing their contextual features, it helps to achieve efficient semantic segmentation of large-scale indoor point cloud scenes. Summary of the invention

[0005] The purpose of the present invention is to provide a semantic segmentation method for indoor point cloud scenes based on patch context features. In order to overcome the defects of large point cloud data size, uneven distribution of point cloud data and difficulty in effectively analyzing the contextual semantic relationship of indoor scenes in indoor point cloud scene understanding, it is possible to efficiently and accurately segment large-scale indoor scene point cloud data with a large number of repeated structures.

[0006] The present invention is based on large-scale scene point cloud data input by users, and uses a dynamic region growing algorithm to extract point clouds with consistent geometric features in indoor scenes as scene patches. Secondly, a multi-scale structure is used in a neural network encoder and a patch local feature aggregation module is embedded to effectively aggregate contextual information of adjacent scene patches at different scales. Then, a Transformer module based on a multi-head self-attention mechanism is used to learn the global features of scene patches. The module constructs a high-level semantic space of scene patches by stacking four attention modules in series and learns the feature similarity between scene patches. At the same time, the global features of scene patches are connected and fused with local features, and a multi-layer perceptron is used to reduce the feature dimension of scene patches. Then, interpolation upsampling is used in a neural network decoder to restore the resolution of the downsampled scene patches to the original resolution and a semantic label is assigned to each patch, thereby finally realizing large-scale scene semantic segmentation.

[0007] The specific steps include:

[0008] A large number of scene patches are extracted from indoor scene point cloud data through the region growing strategy; the point cloud segmentation MPTNet network based on the encoder-decoder structure is used to perform semantic segmentation of point cloud scenes. The encoder in the point cloud segmentation MPTNet network consists of a multi-scale feature aggregation layer and an intermediate layer patch Transformer module, and the decoder consists of a feature diffusion layer. The multi-scale feature aggregation layer in the encoder structure includes a patch local feature aggregation module, which is used to obtain local feature information of scene patches at different scales through downsampling and aggregate context information of adjacent scene patches at different scales. The intermediate layer patch Transformer module is used to learn the global features of the downsampled scene patches. The patch Transformer module can learn the scene patch features in the high-level semantic space with the help of the offset attention mechanism. At the same time, the global features of the scene patches are obtained through the average pooling layer and the maximum pooling layer, and the global features of the scene patches are connected and fused with the local features. In the decoder in the feature diffusion layer, the resolution of the downsampled scene patches is restored to the original patch resolution by interpolation upsampling and a semantic label is assigned to each scene patch, finally realizing the semantic segmentation of large-scale indoor point cloud scenes.

[0009] The point cloud segmentation network based on facet context features proposed in the present invention adopts a self-attention module to be suitable for scene facet feature extraction, which can effectively learn the feature similarity between scene faces and improve the effectiveness of scene segmentation from the perspective of geometric semantics; a local feature aggregation module that can effectively aggregate the context information of scene faces is adopted, which can effectively aggregate the context information feature information of adjacent facets of the scene at different scales and improve the accuracy of scene segmentation from the perspective of feature extraction. The scene semantic segmentation method provided by the present invention can efficiently and accurately segment large-scale indoor point cloud scenes with a large number of repeated structures. Brief Description of the Drawings

[0010] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0011] Figure 2 It is a schematic diagram of the point cloud segmentation MPTNet network proposed by the present invention;

[0012] Figure 3 It is a schematic diagram of the patch multi-scale feature aggregation MPLA module in the point cloud segmentation MPTNet network;

[0013] Figure 4 It is a schematic diagram of the patch local feature aggregation PLA module in the point cloud segmentation MPTNet network;

[0014] Figure 5 It is an example diagram of the semantic segmentation effect of the point cloud data for the classroom scene in the embodiment;

[0015] Figure 6 It is an example diagram of the semantic segmentation effect of the point cloud data for the corridor scene in the embodiment. Detailed Embodiment

[0016] The technical method and object detection effect of the present invention will be further described and explained below with reference to the drawings.

[0017] As Figure 1 shown, a method for semantic segmentation of indoor point cloud scenes based on patch context features is as follows:

[0018] Step 1: Extract a large number of scene patches from the indoor scene point cloud data through a region growing strategy;

[0019] During the region growing process, analyze the indoor scene point cloud data and sort it according to the curvature of its sampling points. Select the sampling point s with the largest curvature as the seed point, and set the initial patch Π as an empty set; then select the nearest neighbor point p outside the patch Π according to the seed point. Assume that the nearest neighbor point p satisfies the following conditions: N p ·N s > t1, (p - s)·N s < t2, (p - q)·N q < t3, #(П)<t4, then add the nearest neighbor point p to the patch Π until the number of point clouds in the patch reaches the threshold upper limit t4, then select another seed point and repeat the above operation until all the point cloud data in the scene is traversed.

[0020] Among them, q represents the last sampling point added to the patch Π in sequence, N represents the normal vector of the corresponding sampling point, # represents the number of sampling points in the point set, and t1, t2, t3, and t4 are threshold parameters.

[0021] Step 2: Based on the scene patch representation, the point cloud segmentation MPTNet network based on the encoder-decoder structure is used to finally achieve semantic segmentation of large-scale indoor point cloud scenes.

[0022] like Figure 2 As shown in the figure, in the point cloud segmentation MPTNet network, the encoder is composed of a multi-scale feature aggregation layer MPLA and an intermediate layer Transformer module, and the decoder is composed of a feature diffusion layer.

[0023] In the multi-scale feature aggregation MPLA layer, the number of input point cloud scene faces is N and its feature dimension is 64 dimensions; then, through two downsamplings, the number of scene faces is reduced from N to N / 4 and N / 16, and its feature dimension is increased from 64 dimensions to 128 dimensions and 256 dimensions; among them, in order to aggregate the local features of the scene faces, the PLA module is embedded in the downsampling process to obtain 128-dimensional and 256-dimensional scene face features; finally, a scene point cloud of N / 16 faces containing 256-dimensional feature information is obtained through a multi-layer perceptron and maximum pooling operations.

[0024] In the middle layer patch Transformer module, it is necessary to learn the scene feature information of N / 16 scene patches (whose feature dimension is 256 dimensions). In order to learn the global features of the downsampled scene patches, the patch Transformer module is first used to construct the high-level semantic space of the scene patches by stacking 4 attention modules in series, and the feature similarity between the scene patches is learned in the high-level semantic space to obtain the global features of the scene patches with N / 16 number of scene patches and 1024 dimensions. Then, a global feature of the scene patches with 1024 dimensions is obtained through the average pooling layer and the maximum pooling layer respectively, and the global feature is stacked into N / 16 scene patch features with 1024 dimensions by repeating (Repeat) operation. Finally, the local features of the scene patches are connected and fused with the global features into a scene patch feature with a number of N / 16 and a feature dimension of 2304 dimensions, and the scene point cloud data including N / 16 patches with 256-dimensional feature information is obtained through a multi-layer perceptron.

[0025] In the feature diffusion layer, the scene patches with a number of N / 16 and a feature dimension of 256 are first interpolated and upsampled twice, and the resolution of the downsampled scene patches is restored to the original patch resolution, that is, the number of scene patches increases from N / 16 to N / 4 and N, and the patch feature dimension decreases from 256 to 128 and 64 respectively; then, N scene patches including 13-dimensional feature information are obtained through a multi-layer perceptron and a semantic label is assigned to each scene patch, ultimately achieving semantic segmentation of large-scale indoor point cloud scenes.

[0026] like Figure 3 As shown in Figure 1, the multi-scale feature aggregation layer extracts patch features at different scales. Scene point cloud composed of patches As input (I stands for input), the input includes N I scene patches, each patch includes F I dimensional feature information and 3D patch centroid coordinate information; first, the downsampled patch point cloud is obtained through the farthest point sampling algorithm (S stands for downsampling). The downsampling process is performed by taking the input scene patch point cloud Π I Select a seed patch s from the sample and add it to the downsampled patch point cloud. Calculate the centroid coordinates of the seed patch s and Π I The Euclidean distance of the centroid coordinates of the remaining patches in the π S , and repeat this cycle until П S The number of faces in the S So far. The number of scene faces is N I Down to N S At the same time, the feature dimension of the downsampled patch is increased by 2 times to F through a multi-layer perceptron S dimensional feature information; then, in order to obtain the local information of the surrounding neighborhood, for each downsampled patch point cloud ∏ S Through the KNN algorithm, the scene patch point cloud is input I Get the k nearest neighboring faces from S A scene neighbor patch point cloud consisting of k neighbor patches Each patch includes F S Then, using Figure 4 The local feature aggregation (PLA) module shown in Figure 2 obtains N pla Locally aggregated features Each patch includes F pla dimensional feature information, and then reduce the feature dimension of the patch through the multi-layer perceptron and extract the global features of the scene patch through the maximum pooling function; finally, the downsampled patch point cloud ∏ S With aggregated patch point cloud pla Merge to get scene point cloud patch representation (O represents input), where each patch includes F O dimensional feature information and 3D patch centroid coordinate information.

[0027] like Figure 5 As shown in the figure, the semantic segmentation effect of classroom scene point cloud data is given. Figure 5 a is the input classroom scene point cloud data, Figure 5b is an example of the semantic segmentation effect of the point cloud scene achieved by the above method (the circle represents the area with better segmentation effect). Figure 5 c is an example of the local magnification effect; Figure 6 As shown in the figure, the semantic segmentation effect of the corridor scene point cloud data is given. Figure 6 a is the input corridor scene point cloud data, Figure 6 b is an example of the semantic segmentation effect of the point cloud scene achieved by the above method (the circle represents the area with better segmentation effect). Figure 6 c is an example of the local magnification effect.

[0028] Depend on Figure 5 It can be seen that the method proposed in this patent can effectively perform semantic segmentation on classroom scene point cloud data, and its segmentation results can maintain the complete structure of the classroom scene, while the boundaries of different object shapes are relatively clear; from the example picture of the local magnification effect, it can be seen that the point clouds of categories such as chairs, ceilings, floors, doors, walls, columns, etc. in the classroom scene can be effectively segmented, especially for chairs with a large number of repeated structures. The segmentation effect is better.

[0029] Depend on Figure 6 It can be seen that the method proposed in this patent can effectively perform semantic segmentation on the corridor scene point cloud data; from the example picture of the local magnification effect, it can be seen that most of the architectural elements in the corridor scene can be accurately segmented and the integrity of their structural information can be guaranteed, especially the wall elements (although obstructed by columns, beams, etc.) can also be effectively segmented.

Claims

1. A semantic segmentation method for indoor point cloud scenes based on patch context features, characterized by: The specific steps include: Scene patches are extracted from indoor scene point cloud data through the region growing strategy; the point cloud segmentation MPTNet network based on the encoder-decoder structure is used to perform semantic segmentation of point cloud scenes; the encoder in the point cloud segmentation MPTNet network consists of a multi-scale feature aggregation layer and an intermediate layer patch Transformer module, and the decoder consists of a feature diffusion layer; the multi-scale feature aggregation layer in the encoder structure includes a patch local feature aggregation module, which is used to obtain local feature information of scene patches at different scales through downsampling and aggregate context information of adjacent scene patches at different scales; the intermediate layer patch Transformer module is used to learn the global features of the downsampled scene patches. The patch Transformer module can learn the scene patch features in the high-level semantic space with the help of the offset attention mechanism, and at the same time obtain the global features of the scene patches through the average pooling layer and the maximum pooling layer respectively, and connect and fuse the global features of the scene patches with the local features; in the decoder feature diffusion layer, the resolution of the downsampled scene patches is restored to the original patch resolution by interpolation upsampling and a semantic label is assigned to each scene patch, finally realizing the semantic segmentation of large-scale indoor point cloud scenes; During the process of the regional growth strategy, first analyze the scene point cloud data and sort it according to the curvature of its sampling points. Select the sampling point s with the largest curvature as the seed point, and set the initial patch Π as an empty set. Then select the nearest neighbor point p outside the patch Π according to the seed point. Assume that the neighbor point p satisfies the following conditions: N p ·N s > t1, (p - s)·N s < t2, (p - q)·N q < t3, #(Π) < t4, then add the neighbor point p to the patch Π until the number of point clouds in the patch reaches the threshold upper limit t4, then select another seed point and repeat the above operation until all point cloud data in the scene are traversed; where, in the formula, q represents the last sampling point added to the patch ∏ in sequence, N represents the normal vector of the corresponding sampling point, # represents the number of sampling points in the point set, and t1, t2, t3, t4 are threshold parameters.

2. The method for semantic segmentation of indoor point cloud scenes based on patch context features according to claim 1, characterized in that: In the multi-scale feature aggregation layer, the number of input point cloud scene faces is N and its feature dimension is 64 dimensions; then, the number of scene faces is reduced from N to N / 4 and N / 16 respectively through two downsamplings, and its feature dimension is increased from 64 dimensions to 128 dimensions and 256 dimensions respectively; wherein, in order to aggregate the local features of the scene faces, the PLA module is embedded in the downsampling process to obtain 128-dimensional and 256-dimensional scene face features; finally, a scene point cloud of N / 16 faces containing 256-dimensional feature information is obtained through a multi-layer perceptron and a maximum pooling operation.

3. The method for semantic segmentation of indoor point cloud scenes based on patch context features according to claim 1, characterized in that: In the intermediate layer patch Transformer module, it is necessary to learn scene feature information of N / 16 scene patches; Firstly, the high-level semantic space of scene patches is constructed by using the patch Transformer module and stacking 4 attention modules in series. The feature similarity between scene patches is learned in the high-level semantic space to obtain the global features of scene patches with N / 16 number of patches and 1024 dimensions. Then, a global feature of scene patches with 1024 dimensions is obtained through the average pooling layer and the maximum pooling layer respectively, and the global features are stacked into N / 16 scene patch features with 1024 dimensions by repeated operations. Finally, the local features of scene patches are connected and fused with the global features into scene patch features with N / 16 number and 2304 dimensions, and the scene point cloud data with N / 16 patches and 256-dimensional feature information is obtained through a multi-layer perceptron.

4. The method for semantic segmentation of indoor point cloud scenes based on patch context features according to claim 1, characterized in that: In the feature diffusion layer, the scene patches with a number of N / 16 and a feature dimension of 256 are first interpolated and upsampled twice, and the resolution of the downsampled scene patches is restored to the original patch resolution, that is, the number of scene patches is increased from N / 16 to N / 4 and N, and the patch feature dimension is decreased from 256 to 128 and 64 respectively; then, N scene patches including 13-dimensional feature information are obtained through a multi-layer perceptron and a semantic label is assigned to each scene patch, finally realizing the semantic segmentation of large-scale indoor point cloud scenes.

5. The method for semantic segmentation of indoor point cloud scenes based on patch context features according to claim 2, characterized in that: The multi-scale feature aggregation layer extracts patch features at different scales, and the scene point cloud composed of patches As input, I represents input, which includes N I scene patches, each patch includes F I dimensional feature information and 3D face centroid coordinate information; First, the downsampled patch point cloud is obtained through the farthest point sampling algorithm S stands for downsampling; The downsampling process is performed by taking the input scene patch point cloud ∏ I Select a seed patch s from the sample and add it to the downsampled patch point cloud. Calculate the centroid coordinates of the seed patch s and ∏ I The Euclidean distance of the centroid coordinates of the remaining patches in the π S In this cycle, until Π S The number of faces in the S So far; the number of scene faces is N I Down to N S At the same time, the feature dimension of the downsampled patch is increased by 2 times to F through a multi-layer perceptron S dimensional feature information; then, in order to obtain the local information of the surrounding neighborhood, for each downsampled patch point cloud Π S Through the KNN algorithm, the scene patch point cloud Π I Get the k nearest neighboring faces from S A scene neighbor patch point cloud consisting of k neighbor patches Each patch includes F S dimensional feature information; then, the local feature aggregation module is used to obtain N pla Locally aggregated features Each patch includes F pla dimensional feature information, and then reduce the feature dimension of the patch through a multi-layer perceptron and extract the global features of the scene patch through the maximum pooling function; finally, the downsampled patch point cloud Π S With aggregated patch point cloud pla Merge to get scene point cloud patch representation O represents the input, where each patch includes F O dimensional feature information and 3D patch centroid coordinate information.

Citation Information

Patent Citations

  • A method for semantic segmentation of scene point cloud

    CN109410307A

  • RandLA-Net outdoor scene semantic segmentation method based on local feature enhancement

    CN114758129A