A 3D Semantic Segmentation Method Based on Channel Attention and Multi-Scale Fusion

By introducing channel attention and multi-scale fusion technology in three-dimensional point cloud semantic segmentation, using position adaptive convolution and multi-scale convolution context modules, the problems of low segmentation accuracy and serious information loss in the existing methods are solved, and more efficient calculations and more accurate segmentation results are achieved.

CN114743007BActive Publication Date: 2025-05-30XIANGTAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210418602.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-05-30
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud semantic segmentation method has problems such as low segmentation accuracy, serious information loss and low computational efficiency.

Method used

The three-dimensional point cloud semantic segmentation method based on channel attention and multi-scale fusion is adopted to extract features through position adaptive convolution, and the channel attention layer is introduced for feature recalibration, and rich features are extracted using the multi-scale convolution context module.

Benefits of technology

The segmentation accuracy of three-dimensional point cloud semantic segmentation is improved, information loss is reduced, and computing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743007B_ABST
    Figure CN114743007B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of three-dimensional point cloud data processing, and discloses a three-dimensional point cloud semantic segmentation method based on channel attention and multi-scale fusion. First, the point cloud data to be segmented is read, preprocessed and then input into the segmentation network. Then, it sequentially passes through four modules composed of an encoder and a channel attention layer, where the encoder includes a downsampling layer, a grouping layer, and a position adaptive convolution. Next, a multi-scale convolutional context module is used to extract the point cloud context information, and finally, it sequentially passes through four decoders composed of an upsampling layer and a unit PointNet network. The final segmentation result is obtained through a fully connected layer of size k (number of categories). The present invention not only makes full use of the position information of the point cloud, but also introduces a channel attention layer to recalibrate the point cloud features in the channel dimension, pays more attention to the channel information useful for the segmentation task, and further proposes a multi-scale convolutional context module, which parallelly captures features of different scales by using dilated convolutions with the same dilation rate but different kernel sizes, thereby improving the segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional point cloud data processing, and particularly relates to a three-dimensional point cloud semantic segmentation method based on channel attention and multi-scale fusion. Background Art

[0002] With the development and rise of artificial intelligence technology, 3D point cloud data analysis has attracted wide attention. Compared with two-dimensional images, 3D point clouds contain richer three-dimensional spatial information, are not affected by external factors such as light and perspective, and can depict models more accurately and comprehensively. And 3D point cloud segmentation, as a key content of scene understanding, is one of the forefront research directions of artificial intelligence and has wide applications in the fields of robotics, virtual reality, autonomous driving, and laser remote sensing measurement.

[0003] Point cloud segmentation methods can be divided into traditional point cloud segmentation and point cloud semantic segmentation. Traditional point cloud segmentation uses information such as the position and shape of the point cloud to segment different regional boundaries. There are mainly edge-based, region-based, and model-fitting-based segmentation methods. The segmentation results obtained do not contain any semantic information and require manual semantic annotation of the results, which is extremely inefficient in the case of large data scales. Point cloud semantic segmentation, on the basis of traditional point cloud segmentation, automatically labels different types of objects in three-dimensional space with different semantic labels, so that each object has specific category information. At present, it mainly uses deep learning as the implementation means, and the processing methods are mainly the following three:

[0004] (1) Voxel-based method: The three-dimensional scene is divided into voxel grids, the original three-dimensional point cloud is converted into voxels, and then a three-dimensional convolutional network is used for processing. However, the three-dimensional points are mainly concentrated on the surface of the object, and become very sparse after being converted into voxels, resulting in very low time and space utilization rates of the dense convolutional network, and some information is lost during the voxel conversion process, thus affecting the performance of the network.

[0005] (2) Multi-view projection-based method: First, a three-dimensional object is projected into multiple views, and then conventional two-dimensional convolutional neural networks are used to extract image features and perform target recognition and analysis. Due to the object occlusion problem in the real scene, some information is lost after the object is projected onto the two-dimensional plane, and the spatial structure information contained in the three-dimensional data cannot be fully utilized. The choice of the projection plane will also have a certain impact on the results of the algorithm.

[0006] (3) Original point cloud-based method: Without relying on intermediate data types, the three-dimensional point cloud data in the scene is directly processed. It requires less memory than the first two methods and does not cause information loss. It mainly uses multi-layer perceptrons or convolutional methods suitable for point cloud data to extract features. Summary of the Invention

[0007] The object of the present invention is to overcome the deficiencies of the above-mentioned existing technologies, and propose a three-dimensional point cloud semantic segmentation method based on channel attention and multi-scale fusion, so as to improve the segmentation accuracy.

[0008] The three-dimensional point cloud semantic segmentation method based on channel attention and multi-scale fusion of the present invention includes the following steps:

[0009] Step 1, read and preprocess the point cloud data;

[0010] Step 2, pass the point cloud data through an encoder composed of a downsampling layer, a grouping layer, and a position-adaptive convolution, which is mainly responsible for upsampling and feature extraction;

[0011] Step 2.1, use the downsampling layer to downsample the point cloud.

[0012] Step 2.2, use the grouping layer to divide the point set obtained in the previous step into several regions.

[0013] Step 2.3, use the position-adaptive convolution method to extract initial features for each region.

[0014] Step 3, use the channel attention layer to recalibrate the point cloud features, model the correlation between channel feature information, and change the corresponding proportion of different features in the overall feature expression by learning their weight values;

[0015] Step 4, repeat Steps 2 to 3 four times to extract point cloud features by layer-by-layer downsampling.

[0016] Step 5, input the feature vector output by the last channel attention layer into the multi-scale convolutional context module, which uses dilated convolutions with the same dilation rate but different kernel sizes to sample the features in parallel, gradually increasing the receptive field range to make up for the lost detailed information.

[0017] Step 6, pass the feature vector output by the multi-scale convolutional context module through a decoder composed of an upsampling layer and a unit PointNet network, which is mainly responsible for downsampling and feature decoding, and use skip connections to take the input of the encoder as another input of the decoder.

[0018] Step 6.1, use the upsampling layer to upsample the point cloud features.

[0019] Step 6.2, use the unit PointNet network to decode the features.

[0020] Step 7, repeat Step 6 four times to decode the point cloud features by layer-by-layer upsampling.

[0021] Step 8: Obtain classification scores for k classes through a fully connected layer of size k (number of classes), and then obtain the segmentation result.

[0022] The present invention has the following advantages compared with the prior art:

[0023] (1) The present invention uses position - adaptive convolution instead of the commonly used multi - layer perceptron to extract point cloud features, constructs the convolution kernel in a dynamically data - driven manner, makes full use of the position information of points, and flexibly utilizes the irregular geometric structure of 3D point clouds.

[0024] (2) The present invention introduces a channel attention layer, which can fully apply the channel information of features, increase the weight value of information that contributes greatly to the network model, and vice versa (reduce the weight value corresponding to features with less information), realizing the recalibration of features by the model.

[0025] (3) The present invention proposes a multi - scale convolutional context module for extracting point cloud context information. It uses dilated convolutions with the same dilation rate but different kernel sizes to capture features at different scales in parallel and improve the segmentation result. Brief Description of the Drawings

[0026] Figure 1 It is a schematic diagram of the 3D point cloud semantic segmentation network structure of the present invention.

[0027] Figure 2 For the present invention Figure 1 Schematic diagrams of the structures of the encoder and decoder in it.

[0028] Figure 3 For the present invention Figure 2 Schematic diagram of the structure of the position - adaptive convolution in it.

[0029] Figure 4 For the present invention Figure 1 Flow schematic diagram of the channel attention layer in it.

[0030] Figure 5 It is a comparison schematic diagram between the dilated convolution and the standard convolution in the multi - scale convolutional context module of the present invention.

[0031] Figure 6 For the present invention Figure 1 Schematic diagram of the structure of the multi - scale convolutional context module in it. Detailed Embodiments

[0032] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0033] Aiming at the problems existing in the original point cloud segmentation method based on deep learning, the present invention proposes a 3D point cloud semantic segmentation method based on channel attention and multi - scale fusion. The network structure is asFigure 1 As shown. Similar to the image segmentation method, the attention layer focuses on the channel information that is more beneficial to the task, while the multi-scale fusion module further samples the features using dilated convolutions with different receptive field sizes, emphasizes the ignored local information, and at the same time uses position-adaptive convolutions more suitable for point cloud data to extract preliminary features. The specific implementation process is as follows:

[0034] Step 1: Read and preprocess the point cloud data;

[0035] Existing point cloud datasets are mainly divided into indoor scenes and outdoor scenes. Among them, indoor datasets include S3DIS, ScanNet, and Semantics, etc. The storage formats mainly include TXT, PLY, OBJ, and BIN. Therefore, first, the point cloud format needs to be unified and the data needs to be read, and then preprocessing operations such as rotation and denoising are performed to simplify the point cloud data while maintaining geometric features, providing a robust data basis for subsequent processing.

[0036] Step 2: Pass the point cloud data through an encoder composed of a downsampling layer, a grouping layer, and a position-adaptive convolution as shown on the left. Figure 2 As shown on the left.

[0037] Step 2.1: Use the downsampling layer to downsample the point cloud.

[0038] Given the input points {x 1 , x 2 ,..., x n}, the farthest point sampling (FPS) method is used to select a subset consisting of m center points such that is the farthest point relative to other points in the set . Compared with random sampling, when the number of given centroids is the same, it can better cover the entire point set and generate receptive fields in a data-dependent manner.

[0039] Step 2.2: Use the grouping layer to divide the point set obtained in the previous step into several regions.

[0040] The input of this layer is a point set of size N×(d + C) and a set of centroid coordinates of size N 0 ×d. The output is a set of point sets of size N 0 ×K×(d + C), where each group corresponds to a local region and K is the number of points near the centroid point. The grouping method uses the boolean query method to select K points within a given radius, where the query distance is the metric distance, and at the same time, the K values of different local regions are different. Compared with the k-nearest neighbor (kNN) search, this query method ensures a fixed regional scale, making the local region features more general in space.

[0041] Step 2.3: Use the Position Adaptive Convolution method (PAConv) to extract initial features for each region.

[0042] As Figure 3 shown, PAConv first defines a weight bank (Weight Bank) consisting of weight matrices, and then the scoring network (ScoreNet) learns coefficient vectors according to the point positions to combine the weight matrices. Finally, the dynamic kernel is generated by combining the weight matrices and their related position adaptive coefficients. The obtained convolution kernel is applied to the input features and then the output features are obtained through max pooling. The detailed process is as follows:

[0043] The weight bank B = {B m | m = 1,..., M} is generated by random initialization, where each represents a weight matrix and M represents the number of matrices. ScoreNet is responsible for associating the relative positions of the points with the weight matrices. Given the position relationship between the central point p i and its adjacent point p j (p i , p j ) ∈ R Din , ScoreNet predicts the position adaptive coefficient m of B

[0044] S ij = α(θ(p i , p j )) (1)

[0045] In formula (1), θ represents a multi-layer perceptron (MLP) and α is a normalization operation implemented using the softmax function. The output vector where represents the coefficient of B i , p j ) when constructing the kernel K(p m ), and M is the number of weight matrices. The softmax function ensures that the coefficients are in the range of 0 to 1, ensuring that each weight matrix will be selected with a certain probability value. The larger the value, the stronger the relationship between the position input and the weight matrix. The kernel of PAConv is obtained according to formula (2) by combining the weight matrices in the weight bank with the position adaptive coefficients predicted by ScoreNet.

[0046]

[0047] Finally, the generated kernel is applied to the input features according to formula (3), and a new feature vector is obtained through max pooling.

[0048]

[0049] In formula (3), K represents the convolution kernel, represents the max pooling operation, P in and P out respectively represent the input and output features.

[0050] Step 3: Use the channel attention layer (L_SE layer) to recalibrate the point cloud features.

[0051] The L_SE layer consists of three parts: Squeeze, Excitation, and Reweight. Squeeze performs feature compression in the spatial dimension, turning each feature channel into a real number. This real number has a global receptive field to some extent, and the output dimension is the same as the number of input feature channels. Excitation generates a weight on each feature channel based on the correlation between feature channels, representing the importance of the feature channel. Reweight treats the weights output by Excitation as the importance of each feature channel, and then weights the previous features channel by channel through multiplication to complete the recalibration of the original features in the channel dimension. The detailed process is as follows:

[0052] For point cloud data, Squeeze is implemented by one-dimensional global average pooling, as shown in formula (4), to complete the correlation statistics of information between feature mapping channels:

[0053] P avg = AvgPool1D(P in ) (4)

[0054] Based on the information obtained through the Squeeze operation, in order to further capture the correlation information between channels, an operation is performed with the sigmoid activation function, as shown in formula (4).

[0055] P s = σ(L(δ(L(P avg )))) (5)

[0056] In formula (5), σ represents the sigmoid function, L represents the Linear linear function, and δ represents the Leaky_ReLU activation function. In the backpropagation process, different from the ReLU function of the original network, the Leaky_ReLU activation function selected in the present invention can also calculate the gradient in the part where the input is less than zero, rather than being 0 like ReLU, as shown in formulas (6) and (7), which can solve the problem of neuron "death".

[0057] ReLU(x) = max(0, x) (6)

[0058] Leaky_ReLU = max(0, αx) (7)

[0059] To reduce the complexity of the network model and improve the adaptability of the network to different data, the first Linear function reduces the input channel dimension to Then, through the Leaky_ReLU activation function, and then another Linear function is used to expand the dimension of the data to make it the same as the original input dimension. Finally, it is input into the sigmoid function to normalize the weight value to a value between 0 and 1, and then the weight value is weighted to the original channel information through formula (8) to complete the recalibration.

[0060]

[0061] P in formula (8) out is the new feature output by the L_SE layer, and the calculation process is as Figure 4 shown.

[0062] Step 4: Repeat steps 2 to 3 four times to extract point cloud features by downsampling layer by layer.

[0063] Step 5: Use the multi-scale convolutional context (MSCC) module to extract detailed information.

[0064] MSCC is designed to extract rich point cloud features. Different from standard convolution, one-dimensional dilated convolution is selected in the present invention. Dilated convolution is actually a process of sampling point cloud features, and the sampling frequency is set according to the parameter dilation rate (rate). When rate = 1, no information is lost in feature sampling, which is the standard convolution operation; when rate > 1, sampling is performed on every (rate - 1) point cloud in the original data, thereby increasing the range of the receptive field. The actual kernel size K is calculated according to formula (9).

[0065] kernel_size+(kernel_size - 1)(rate - 1) (9)

[0066] In formula (9), kernel_size is the initial kernel size. So when standard convolution is selected, K is equal to kernel_size, while K of dilated convolution is larger. The comparison diagram is as Figure 5 shown.

[0067] While increasing the receptive field, dilated convolution does not reduce the spatial dimension, nor does it increase the number of parameters, achieving a balance between accuracy and speed. The output point cloud size after convolution is calculated according to formula (10):

[0068] ·input: (B, c in , N in )

[0069] Output: (B, C out , N out )

[0070]

[0071] In formula (10), N is the number of point clouds, and dilation represents the rate. For different convolution kernel sizes, in order to keep N unchanged after output, dilation is set to 2 and padding is equal to (kernel_size-1).

[0072] The structure of MSCC is Figure 6 As shown in the figure, firstly, the global information is obtained by using the standard convolution with a kernel size of 1, and then the dilation rate is 2 and the kernel sizes are 3, 5, and 7 respectively. The dilation rate is 2 and the dilation rate is 3, 5, and 7 respectively. The dilation rate is 2 and the kernel sizes are ...

[0073] Step 6: The feature vector output by the multi-scale convolutional context module is processed as follows: Figure 2 The decoder shown on the right consists of an upsampling layer and a unit PointNet network, with the output of the encoder as another input to the decoder through a skip connection.

[0074] Step 6.1: Use the upsampling layer to upsample the point cloud features.

[0075] The interpolation method is used for upsampling to restore the original point cloud scale. According to the coordinates of the center point, the K nearest neighbor algorithm with K=3 is used for interpolation, as shown in formula (11):

[0076]

[0077] Step 6.2: Decode the features using the unit PointNet network.

[0078] The unit PointNet network is mainly composed of a transformation network (T-Net) and a multi-layer perceptron (MLP). T-Net is used to generate a transformation matrix and directly apply this transformation to the coordinates of the input point. Specifically, two-dimensional regularization is used, and in order to maintain the rotation invariance of the point cloud, an orthogonal matrix is ​​used as much as possible, as shown in formula (12). The T-Net network is used to align features to make them easier to extract.

[0079] P reg =||I-AA T || 2 (12)

[0080] In formula (10), P regThe converted feature matrix is \(F\), \(I\) is the identity matrix corresponding to the dimension of the input matrix, and \(A\) is the feature matrix to be converted.

[0081] The MLP is a neural network model composed of an input layer, a hidden layer, and an output layer, and the output is \(h w,b (x)\), where \(W\) represents the inter-layer weight matrix and \(b\) represents the bias. The unit PointNet network has 3 or 4 MLPs, which sequentially reduce the dimension of the feature vector.

[0082] Step 7: Repeat Step 6 seven times to upsample and decode the point cloud features layer by layer.

[0083] Step 8: Obtain the classification scores of \(k\) classes through a fully connected layer of size \(k\) (the number of classes), and then obtain the segmentation result.

[0084] Embodiment

[0085] The dataset used in this embodiment is the S3DIS dataset, which is collected from the indoor environments of three different buildings and contains 271 rooms in 6 regions. It has a total of 695,878,620 point clouds, each of which has corresponding coordinate and color information, as well as semantic labels such as chairs, tables, floors, and walls, for a total of 13 categories. In this embodiment, regions 1, 2, 3, 4, and 6 are selected for training, and region 5 is used for testing. During training, this embodiment samples the input points into a unified number of 4096 points, while all points are used during testing.

[0086] This embodiment is trained for 150 epochs on two GeForce RTX 2080Ti GPUs, with a batch size of 16, using the SGD optimizer with an initial learning rate of 0.05, a momentum of 0.9, and a weight decay rate of \(10 -4 ^{-4}\), and is implemented on the Pytorch platform using Linux. After training the network with the training set to obtain a model, the model performance is evaluated through the test set, and the mIoU (mean intersection over union) is selected as the evaluation metric. The IoU (intersection over union) of each category on the S3DIS dataset is shown in Table 1, and the mIoU is 64.8. It can be seen that the present invention can achieve good segmentation performance in the 3D point cloud semantic segmentation task.

[0087] Table 1: IoU results of each category on the S3DIS dataset

[0088]

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A 3D point cloud semantic segmentation method based on channel attention and multi-scale fusion, characterized in that, it includes the following steps: Step 1, read and preprocess the point cloud data; Step 2, pass the point cloud data through an encoder composed of a downsampling layer, a grouping layer, and a position-adaptive convolution, which is mainly responsible for upsampling and feature extraction; Step 3, use the channel attention layer to recalibrate the point cloud features, model the correlation between channel feature information, and change the corresponding proportion of different features in the overall feature expression by learning their weight values; Step 4, repeat steps 2 to 3 four times to extract point cloud features by layer-by-layer downsampling; Step 5, input the feature vector output by the last channel attention layer into the multi-scale convolutional context module, which uses dilated convolutions with the same dilation rate but different kernel sizes to sample the features in parallel, gradually increasing the receptive field range to make up for the lost detailed information; Step 6, pass the feature vector output by the multi-scale convolutional context module through a decoder composed of an upsampling layer and a unit PointNet network, which is mainly responsible for downsampling and feature decoding, and use skip connections to take the input of the encoder as another input of the decoder; Step 7, repeat step 6 four times to decode the point cloud features by layer-by-layer upsampling; Step 8, obtain the classification scores of k classes through a fully connected layer with the number of classes being k, and then obtain the segmentation result.

2. The method according to claim 1, characterized in that, in step 2, the position-adaptive convolution first defines a weight library composed of weight matrices, and then the scoring network ScoreNet learns a coefficient vector according to the point positions to combine the weight matrices, generates a dynamic kernel by combining the weight matrices and their related position-adaptive coefficients, and finally applies the obtained convolution kernel to the input features and then obtains the output features through max pooling. The detailed process is as follows: Weight library B = {B m |m=1,...,M} is generated by random initialization, where each represents a weight matrix, M represents the number of matrices; ScoreNet is responsible for associating the relative position of the point with the weight matrix. Given a center point p i Its adjacent point p j The position relationship (p i , p j )∈R Din , ScoreNet predicts B m Position adaptation coefficient for: S ij = α(θ(p i p j )) Where θ represents a multi-layer perceptron (MLP), α is the normalization operation that enables the softmax function; the output vector in Denotes the construction kernel K(p i , p j ) m The coefficients of , M is the number of weight matrices; the softmax function ensures that the coefficients range between 0 and 1, ensuring that each weight matrix will be selected with a certain probability value. The larger the value, the stronger the relationship between the position input and the weight matrix. The kernel of PAConv is obtained by combining the weight matrix in the weight library with the position adaptation coefficient predicted by ScoreNet: Apply the generated kernel to the input features and obtain a new feature vector through max pooling: where L represents the convolutional kernel, represents the max pooling operation, P in and P out represent the input and output features respectively.

3. The method according to claim 1, characterized in that, in step 3, the channel attention layer is composed of three parts: Squeeze, Excitation, and Reweight. Squeeze performs feature compression in the spatial dimension, turning each feature channel into a real number, which has a global receptive field to some extent, and the output dimension is the same as the number of input feature channels; Excitation generates a weight on each feature channel based on the correlation between feature channels to represent the importance of the feature channel; Reweight takes the weight output by Excitation as the importance of each feature channel, and then weights the previous features channel by channel through multiplication to complete the recalibration of the original features in the channel dimension. The detailed process is as follows: For point cloud data, Squeeze is implemented by one-dimensional global average pooling to complete the correlation statistics of information between feature mapping channels: P avg = AvgPool1D(P in ) Based on the information obtained through the Squeeze operation, in order to further capture the correlation information between channels, an operation is carried out with the help of the sigmoid activation function: P s = σ(L(δ(P avg )))) where σ represents the sigmoid function, L represents the Linear linear function, and δ represents the Leaky_ReLU activation function; in the backpropagation process, different from the ReLU function of the original network, the Leaky_ReLU function selected in this method can calculate the gradient when the input is less than zero, rather than the ReLU function having a zero gradient in this area, thus effectively solving the problem of neuron "vanishing": ReLU(x) = max(0, x) Leaky_ReLU = max((0, ax) To reduce the complexity of the network model and improve the adaptability of the network to different data, the first Linear function reduces the input channel dimension to Then, through the Leaky_ReLU activation function, and then another Linear function is used to expand the dimension of the data to make it the same as the original input dimension. Finally, it is input into the sigmoid function to normalize the weight value to a value between 0 and 1, and the weight value is weighted to the original channel information to complete the recalibration: Among which P out is the new feature output by the L_SE layer.

4. The method according to claim 1, characterized in that in step 5, the multi-scale convolutional context module is used to extract rich point cloud features. Different from standard convolution, this method selects one-dimensional dilated convolution. Dilated convolution is actually a process of sampling point cloud features, and the sampling frequency is set according to the parameter dilation rate. When rate = 1, no information is lost in feature sampling, which is the standard convolution operation; when rate > 1, sampling is performed on every rate - 1 point cloud in the original data, thereby increasing the range of the receptive field. The actual kernel size K is calculated according to the following formula: kernel_size+(kernel_size - 1)(rate - 1) where kernel_size is the initial kernel size. Therefore, when standard convolution is selected, L is equal to kernel_size, and K of dilated convolution is larger; while increasing the receptive field, dilated convolution does not reduce the spatial dimension, nor does it increase the number of parameters, achieving a balance between accuracy and speed. The output point cloud size after convolution is: input: (B,C in , N in ) ·output: (B, C out , N out ) where N is the number of point clouds, and dilation represents rate; for different kernel sizes, in order to keep N unchanged after output, dilation is set to 2 and padding is equal to kernel_size - 1. Based on the above settings, the multi-scale convolutional context module first obtains global information using a standard convolution with a kernel size of 1, and then parallelly samples using dilated convolutions with a dilation rate of 2 and kernel sizes of 3, 5, and 7 respectively, so as to extract context features with different receptive fields and strengthen the connection between adjacent point clouds.

Citation Information

Patent Citations

  • Three-dimensional point cloud semantic segmentation method based on multi-scale feature fusion

    CN114359902A

  • Method and system for scene image modification

    US20210142497A1