A Weakly Supervised Point Cloud Semantic Segmentation Method, System, Electronic Device and Medium

Through the combination of point center attention mechanism and high-dimensional position coding, the problem of insufficient utilization of labeled information in weakly supervised point cloud semantic segmentation is solved, and the segmentation accuracy and robustness are improved. It is suitable for computer vision and three-dimensional point cloud processing fields.

CN120236084BActive Publication Date: 2025-07-29NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510714365.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-29
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the existing weakly supervised point cloud semantic segmentation method, problems such as insufficient utilization of labeled information, blurred boundaries, misjudgment of categories, insensitive to rotation, and irregularity of local area of point clouds, resulting in insufficient segmentation accuracy and robustness.

Method used

Using a method based on point center and high-dimensional position coding, two feature embeddings are performed through the point center attention mechanism to generate high-dimensional semantic position coding, and adaptive correction is applied to the feature affine transformation module, and feature extraction and aggregation are performed in combination with the Transformer block.

Benefits of technology

The utilization rate of finite labeled information is improved, the model's expression ability and rotational robustness of the point cloud spatial structure is enhanced, the impact of local area inhomogeneity of point cloud on feature learning is alleviated, and the segmentation accuracy and stability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236084B_ABST
    Figure CN120236084B_ABST
Patent Text Reader

Abstract

This application belongs to the technical fields of computer vision, 3D point cloud processing, and weakly supervised learning. It discloses a weakly supervised point cloud semantic segmentation method, system, electronic device, and medium. This method adopts an encoder-decoder architecture and introduces an innovative point center attention mechanism, high-dimensional semantic position encoding, and feature affine transformation module in the encoder. The point center attention mechanism extracts global features from adjacent points through two feature embedding processes, considers multiple neighborhoods where they are located, and shares the center point features with adjacent points to make full use of limited annotation information. The high-dimensional semantic position encoding effectively captures local geometric information and enhances the robustness to rotational deformations through a 7D representation including relative XYZ coordinates, Euclidean distance, and direction angle, and integrates it into the attention calculation. The feature affine transformation module adaptively corrects local point cloud features based on the high-dimensional semantic centroid and feature variance to alleviate the influence of irregularities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, 3D point cloud processing, and weakly supervised learning, and particularly relates to a weakly supervised point cloud semantic segmentation method, system, electronic device, and medium based on point centers and high-dimensional position encoding. Background Art

[0002] With the continuous expansion of the scale of 3D point cloud datasets, the cost of fine-point-by-point annotation of point clouds is extremely high, which has greatly promoted the research on weakly supervised point cloud semantic segmentation methods. Although some existing weakly supervised methods have achieved certain results, due to the scarcity of annotation information, the point cloud semantic segmentation task often faces problems such as blurred boundaries and misclassified categories. Especially in the weakly supervised learning framework, the semantic information contained in spatially limited annotation points is not fully utilized, and it is difficult for the network to comprehensively understand the local and global semantics of the annotation points, which directly affects the accuracy and practicality of the final segmentation results.

[0003] In existing methods, some methods based on point center networks use the relationship between points and their neighborhoods for feature learning and perform well in the fully supervised scenario. However, in the weakly supervised environment, the labeled points are sparse and may be unevenly distributed, resulting in the inability to effectively use the semantic information of the labeled points to update and optimize the unlabeled neighborhoods. In addition, the exploration of the relationship between the center point and neighboring points and the insufficient utilization of global features limit the value of the limited annotation information. Existing local spatial representation methods (such as relative coordinates) are sensitive to the rotational deformation of point clouds and fail to fully capture the characteristics of the same object under different rotations. At the same time, the problems of uneven density and irregular geometric structure in the local regions of point clouds also pose challenges to feature learning. These problems together lead to room for improvement in the segmentation accuracy and robustness of existing weakly supervised point cloud semantic segmentation methods. Summary of the Invention

[0004] In view of the problems in existing weakly supervised point cloud semantic segmentation methods, such as insufficient utilization of annotation information, blurred boundaries, misclassified categories, insensitivity of local spatial representation to rotation, and irregularity of local regions of point clouds, the present invention proposes a weakly supervised point cloud semantic segmentation method and system based on point centers and high-dimensional position encoding.

[0005] In a first aspect, the present invention provides a weakly supervised point cloud semantic segmentation method, including the following steps:

[0006] Receiving input point cloud data;

[0007] Performing encoder processing on the input point cloud data, where the encoder includes multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, and in feature extraction and aggregation, a Transformer block including an attention mechanism is used;

[0008] In the encoder processing, a point - centered attention mechanism is applied, and the point - centered attention mechanism includes:

[0009] Perform a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished by associating with the central weights of the neighborhood where the adjacent points are located; and, perform a second feature embedding process to share the center point features to the adjacent points based on the global features extracted by the first feature embedding process;

[0010] In the encoder processing, generate a high - dimensional semantic position encoding for the local point cloud region; and incorporate the high - dimensional semantic position encoding into the calculation of the point - centered attention mechanism;

[0011] In the encoder processing, apply a feature affine transformation module to adaptively correct the features of the local point cloud region;

[0012] Perform decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels. As an optional implementation manner of the first aspect of this application, the first feature embedding process specifically includes: using a one - dimensional linear layer to calculate the central weights of the center point, the dimension of the center point features is C, and the dimension of the central weights is 1; using K - nearest neighbors to explore adjacent points for the center point; for the adjacent points in a neighborhood of the center point, aggregate the features of the adjacent points according to the sequential positions belonging to and spanning multiple neighborhoods to calculate the global features; and normalize the global features through a normalized exponential function to obtain neighborhood - specific adjacent point weights for the second feature embedding process.

[0013] As an optional implementation manner of the first aspect of this application, the second feature embedding process specifically includes: using the global features extracted by the first feature embedding process as the weights of the adjacent points; and performing matrix multiplication on the center point features and the global features to generate a representation in which the center point features are shared to the adjacent points.

[0014] As an optional implementation manner of the first aspect of this application, the generation of the high - dimensional semantic position encoding includes: for an adjacent point in the local point cloud region relative to a center point, determine the relative XYZ coordinates of the adjacent point relative to the center point; determine the Euclidean distance from the adjacent point to the center point; determine the direction angle ρ formed by the line connecting the center point and the adjacent point and the XOY plane xy ; determine the direction angle ρ formed by the line connecting the center point and the adjacent point and the XOZ plane xz; Determine the direction angle ρ formed by the line connecting the center point and the adjacent point and the YOZ plane yz ; And combine the relative XYZ coordinates, the Euclidean distance, and the above three direction angles to form the high-dimensional semantic position encoding.

[0015] As an alternative implementation of the first aspect of the present application, the high-dimensional semantic position encoding includes 7 dimensions, including 3 relative XYZ coordinate dimensions, 1 Euclidean distance dimension, and 3 direction angle dimensions.

[0016] As an alternative implementation of the first aspect of the present application, the integration of the high-dimensional semantic position encoding into the calculation of the attention mechanism includes: using a linear layer to transform the global features extracted by the first feature embedding process; combining the transformed global features, the representations generated by the second feature embedding process, and the high-dimensional semantic position encoding to calculate the original attention score; and normalizing the original attention score through a normalized exponential function to calculate the final attention weight of the point center attention mechanism.

[0017] As an alternative implementation of the first aspect of the present application, the application of the feature affine transformation module specifically includes: calculating the high-dimensional semantic centroid of the local point cloud region, the high-dimensional semantic centroid being based on the point features within the region and the high-dimensional semantic position encoding; calculating the feature variance within the local point cloud region; based on the high-dimensional semantic centroid and the feature variance, using 2D learnable parameters to perform an affine transformation on the point features within the local point cloud region; and concatenating the features after the affine transformation with the center point features to generate local features.

[0018] In a second aspect, an embodiment of the present application provides a weakly supervised point cloud semantic segmentation system, including:

[0019] An input module for receiving input point cloud data;

[0020] A processing module for performing encoder processing and decoder processing on the input point cloud data;

[0021] The processing module includes an encoder unit and a decoder unit;

[0022] The encoder unit is configured to perform multiple downsampling stages, each downsampling stage including point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks including an attention mechanism;

[0023] The encoder unit includes:

[0024] The point - centered attention sub - unit is configured to apply a point - centered attention mechanism, and the point - centered attention mechanism includes: performing a first feature embedding process to extract global features from adjacent points, where the extraction of the global features is distinguished from the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and performing a second feature embedding process to share the center - point features to the adjacent points based on the global features extracted by the first feature embedding process.

[0025] The position encoding sub - unit is configured to generate a high - dimensional semantic position encoding for the local point cloud region; and incorporate the high - dimensional semantic position encoding into the calculation of the attention mechanism.

[0026] The feature affine transformation sub - unit is configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region.

[0027] The decoder unit is configured to perform decoder processing on the features processed by the encoder to obtain densely output semantic segmentation labels.

[0028] The output module is used to output the semantic segmentation labels.

[0029] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0030] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0031] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in:

[0032] 1. Through the two embedding processes of the point - centered attention network, the semantic information of the labeled points and their neighborhoods is enriched, effectively solving the problem that it is difficult to update the unlabeled neighborhoods under weak supervision, expanding the receptive field of the labeled points, and improving the utilization rate of limited labeled information.

[0033] 2. The proposed high - dimensional semantic position encoding module, represented by 7 dimensions including relative position, distance, and direction angle, can not only accurately describe the spatial structure of the local point cloud, but also learn the different rotational deformation characteristics of the same object, enhancing the model's ability to express the spatial structure of the point cloud and its robustness to rotation.

[0034] 3. The designed feature affine module adaptively corrects local point cloud features based on high-dimensional semantic centroids and feature variances, effectively alleviating the impact of uneven local point cloud density and geometric structure irregularity on feature learning, and improving the stability and accuracy of the model. Description of the Drawings

[0035] Figure 1 is a schematic diagram of the overall network structure of a weakly supervised point cloud semantic segmentation based on point centers and high-dimensional position encoding according to an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of the downsampling module structure according to an embodiment of the present invention;

[0037] Figure 3 is a schematic diagram of the Transformer block structure according to an embodiment of the present invention;

[0038] Figure 4 is a schematic diagram of the upsampling module structure according to an embodiment of the present invention;

[0039] Figure 5 is a detailed schematic diagram of the enhanced feature module (point center attention network part) according to an embodiment of the present invention;

[0040] Figure 6 is a schematic diagram of the high-dimensional position encoding module structure according to an embodiment of the present invention;

[0041] Figure 7 is a schematic diagram of the feature affine module structure according to an embodiment of the present invention

[0042] Figure 8 is a schematic diagram of the structure of a weakly supervised point cloud semantic segmentation system based on point centers and high-dimensional position encoding according to an embodiment of the present invention. Detailed Embodiments

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0044] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally indicates an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0045] Embodiment 1

[0046] The present invention proposes a weakly supervised point cloud semantic segmentation method based on point centers and high-dimensional position encoding. The overall technical route of this method is as Figure 1 shown. An encoder-decoder architecture is adopted, and the input is the XYZ coordinates and RGB color information of the point cloud. In the encoder part stage, multiple Transformer blocks are integrated into each downsampling stage to extract the basic features of the point cloud. Then, these encoded features are sent to the decoder part, which contains upsampling blocks for restoring the point cloud density and outputting dense semantic segmentation labels. Finally, a classifier is used to obtain the categories after semantic segmentation of the point cloud.

[0047] Among them, the downsampling process is as Figure 2 shown, which is mainly used to reduce the number of points in the point cloud, reduce the dimension and computational amount of the data, and at the same time retain important feature information. The downsampling scale is defined as 4, and the number of points is reduced by 4 in each downsampling layer. Then the selected points are grouped, and max pooling is used to aggregate the corresponding features. Among them, Po 、 Fo respectively represent the number of input point clouds and the original features; Pd 、 FdThey are respectively the number of point clouds and the point cloud features after the downsampling process. After the downsampling process, a point set with reduced quantity but still representative and its features are obtained. The specific module functions are: farthest point sampling (FPS), k-nearest neighbors (KNN), multi-layer perceptron (MLP), and max pooling. Farthest point sampling (FPS) is the first step of downsampling, which constructs a subset by selecting the points with the farthest distance in the point cloud. Farthest point sampling can ensure that the sampled points are evenly distributed in space, thus better retaining the overall structural information of the point cloud. K-nearest neighbors (KNN), for the points selected by FPS, find their k nearest neighbor points. The KNN operation can help determine the local area around each sampled point, providing local context information for subsequent feature extraction. Multi-layer perceptron (MLP) is used to learn and transform the features of the points in the local area determined by KNN, extracting more meaningful local features. These local features will be used for subsequent processing. Max pooling, within each local area, through the max pooling operation, selects the point with the largest feature value as the representative of the area. Max pooling can further reduce the number of points while retaining the most significant feature information. After this series of operations, the downsampled point set and its features are obtained.

[0048] Among them, the Transformer block is as follows Figure 3 shown, which is a module commonly used to process sequence data and can also be used for feature extraction and transformation in point cloud processing to enhance the model's ability to capture global information. Multi-layer perceptron (MLP), in the Transformer block, MLP is usually used to perform non-linear transformation on the input features, increasing the model's expressive ability. The first MLP performs preliminary feature extraction and transformation on the input features, extracting more meaningful feature representations. Convolutional neural network with attention (CNA), this step includes convolutional operations to capture local features, then normalization operations to stabilize the training process, and finally non-linearity is introduced through activation functions to further enhance the model's learning ability. The second MLP further transforms and extracts the features after being processed by CNA, further optimizing the feature representations to make them more suitable for subsequent tasks. Dropout is a regularization technique used in neural networks, mainly used to prevent the model from overfitting. Its core idea is to randomly "drop out" (temporarily turn off) a part of the neurons in the neural network during the training process, forcing the network not to overly rely on certain specific nodes, thereby improving the generalization ability. The residual connection structure helps to solve the problem of gradient disappearance in deep networks, enabling the model to be better trained and optimized. Among them, Pd 、 Fd are respectively the number of point clouds and the point cloud features after the downsampling process; Pd 、 F’dThey are the number of point clouds and the enhanced features after feature enhancement, respectively.

[0049] Among them, the upsampling process is as follows Figure 4 As shown, it is mainly used to increase the density of the point cloud, make the point cloud data more refined, so as to better perform subsequent processing such as semantic segmentation. Interpolation, which is one of the key steps of upsampling. Since the original point set is relatively sparse, through the interpolation algorithm, according to the information of the denser points in the middle, the features of the midpoints are estimated and supplemented, so as to obtain a richer feature representation. Concatenate, concatenate the features obtained after interpolation with the original features. In this way, the information of the original features can be retained, and at the same time, the new features obtained through interpolation are incorporated, providing a more comprehensive feature input for subsequent processing. Multilayer perceptron (MLP), the multilayer perceptron is a deep learning model used to further learn and transform the concatenated features. It can extract higher-level and more abstract features, so as to better meet the requirements of tasks such as semantic segmentation. After being processed by the MLP, the final updated features are obtained, which together with the original point set are used as the output of upsampling. Among them Po 、 Fo represent the number of input point clouds and the original features, respectively; Pd 、 F’d They are the number of point clouds and the enhanced features after feature enhancement, respectively; Po 、 F’o represent the number of input point clouds and the final features after upsampling, respectively.

[0050] In this embodiment, the received point cloud data is processed by an encoder, which includes multiple downsampling stages. Each downsampling stage includes point sampling, neighborhood construction, feature extraction and aggregation, and the feature extraction and aggregation use a Transformer block including an attention mechanism. In the encoder processing, a point-centered attention mechanism is applied, which includes: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished from the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process.

[0051] It is understandable that in the weakly supervised point cloud segmentation task, the core challenge lies in how to efficiently optimize the features of unlabeled points when training relying only on a limited number of labeled points, and then achieve accurate segmentation of the entire point cloud. Existing point center-based networks use local attention mechanisms to calculate weights separately within each neighborhood. This attention mechanism has achieved remarkable results in fully supervised large-scale point cloud segmentation because labeled points are often closely connected and have small distances from each other. However, in a weakly supervised environment, its limitations also emerge. Therefore, this method proposes a point center attention mechanism that innovatively integrates two embedding processes, aiming to overcome the problems caused by sparse point annotations in weakly supervised point clouds.

[0052] As Figure 5 shown, Pi 、 Fi are respectively expressed as the number of point clouds and features input into the network. Pij 、 Fij respectively represent the number of point clouds and features after KNN. g1(Fi) represents the result obtained by passing the features of the input points through the linear layer g1 . g2(e1) represents e1 after passing through the linear layer g2 the resulting representation. Specifically, in the local attention mechanism, the weights of adjacent points within a given neighborhood are calculated only based on the features of that neighborhood itself, without considering the internal connections between different neighborhoods. The features of the central point are integrated by weighted summation of the weights of adjacent points. However, in the case of sparse annotations, the backpropagation of the training loss only depends on a limited number of labeled points. This leads to the situation that only when the central point is labeled, its adjacent unlabeled points can obtain corresponding updates and optimizations. Conversely, if the central point is not labeled while its adjacent points are labeled, the method cannot fully exploit the value of these labeled points. When both the central point and adjacent points are not labeled, the problem becomes even more serious, and this situation is called an "unlabeled neighborhood". To address the above problems, the method based on centralized attention uses two embedding processes to properly handle the sparse annotation problem. In the initial embedding stage, the problem of unlabeled neighbors is successfully solved, and the relationships between different neighborhoods are comprehensively considered. Previous studies only obtained global features from the central point, while the present invention starts from another perspective, extracts global features based on neighboring points, and then shares them within the neighborhood. Substantially, by using the existing labeled neighboring points, the features of unlabeled points in multiple neighborhoods are improved, regardless of whether the central point is labeled. In this way, even in an unlabeled neighborhood, adjacent points can be strengthened by the global features from other regions that may contain labeled points. Moreover, these advantages also extend to unlabeled central points, which are further optimized through deep fusion with their corresponding neighborhoods, thus significantly improving the performance of the model in the weakly supervised point cloud segmentation task.

[0053] However, global feature extraction encounters an obstacle in neighborhood construction. In three-dimensional space, a point can act as an adjacent point in neighborhoods with multiple different sequential positions. Therefore, if global features are simply aggregated from all these neighborhoods, the result will be a mixture of all points without any meaningful semantic context. To overcome this problem, the present invention differentiates between different neighborhoods by embedding the features of the different neighborhoods into corresponding central weights. Specifically, the central weights are derived from the features of the central points, initially N×C-dimensional, and then transformed into N×1-dimensional, where N represents the number of point clouds and C represents the dimension of the point cloud features. The global features represented by the neighborhood points are denoted as e1 , and are obtained by aggregating and calculating according to the sequential positions of these features within their respective neighborhoods. In particular, Figure 5 illustrates the proposed attention mechanism with two embedding stages. For the central point Pi , the central weights are calculated through a linear layer g1 with only one dimension. Then, KNN is used to explore the adjacent points, , to obtain K points. In the first embedding, global features are extracted through the central weights and the features of the neighboring points. These features are then used as the neighborhood weights for the second embedding through the normalized exponential function σ. The process formulas are shown in (1) and (2) as follows:

[0054] (1)

[0055] (2)

[0056] where, represents the result obtained after the features i of the th input point pass through the linear layer g1 .

[0057] Although e1 improves the representation of unlabeled neighboring points in weakly supervised point clouds through valuable global features, it often lacks the key local features specific to separated neighborhoods and weakens the influence of the central points. Therefore, compared with previous studies, the effectiveness of calculating the adjacent point weights may not be optimal. To address this limitation, the present invention introduces a second embedding process, aiming to share the central point features to their respective neighborhoods through the above global features e1 . This method uses these global features as the weights of the adjacent points, and then performs matrix multiplication in combination with the central points. Therefore, the second embedding (denoted as e2 ) ensures that the central point features are effectively propagated to each point in the corresponding neighborhood under the guidance of appropriate global features.

[0058] To construct the attention weights for adjacent points that cover global features and valuable central weights, the present invention merges these embedding processes with positional encoding as spatial supplementation. Specifically, through g2 linear layer transformation of global features e1 , and using learnable parameters , b and integrating with the center sharing mechanism e2 . This integration enhances local features through center point sharing while also retaining global features. In addition, positional encoding is incorporated into the attention weights to supplement appropriate geometric features according to the position of the points. Then, the attention weights are multiplied by the value feature map through the normalized exponential function σ . Finally, a weighted sum is taken over all adjacent points, avoiding the problem of irregular point ordering, and the output feature is obtained. The process formula is shown in (3) and (4) as follows:

[0059] (3)

[0060] (4)

[0061] In this embodiment, during the encoder processing, a high-dimensional semantic positional encoding of the local point cloud region is generated; and, the high-dimensional semantic positional encoding is incorporated into the calculation of the attention mechanism.

[0062] It can be understood that a simple representation of the local point cloud is the relative coordinates of adjacent points with respect to the center. By converting the original three-dimensional coordinates into their relative coordinates, a simple local representation of the local point cloud is obtained, that is, the position relative to the centroid. The present invention expects the local representation to be insensitive to rotational deformations, that is, it can learn different rotational characteristics of the same object. Specifically, as Figure 6 shown, the point cloud representation is supplemented with the distance and direction angles around the XYZ axes, where is the Euclidean distance from the adjacent point to the center point, is the direction angle between the line connecting the center point and the adjacent point and the plane formed by the X and Y axes, is the direction angle between the line connecting the center point and the adjacent point and the plane formed by the X and Z axes, is the direction angle between the line connecting the center point and the adjacent point and the plane formed by the Y and Z axes.

[0063] To overcome the problem that simple relative coordinates are sensitive to rotation, the present invention introduces high-dimensional semantic positional encoding (position + angle + distance).

[0064] (1) 7D representation of NeRF

[0065] A simple representation of the local point cloud is the relative coordinates of neighboring points with respect to the center. Converting its relative coordinates into a simple representation of the local point cloud, are respectively represented as the changed coordinates, corresponding to formulas (5), (6), (7):

[0066] (5)

[0067] (6)

[0068] (7)

[0069] The present invention proposes a method for representing local point clouds based on centroid reference. By establishing a relative coordinate system with the local centroid as the origin, the original point cloud is converted into a position-normalized local representation. That is to say, different rotation attributes of the same object can be learned. Specifically, as shown in formulas (8), (9), (10), (11), the point cloud representation is supplemented with distances and direction angles around the XYZ axes:

[0070] (8)

[0071] (9)

[0072] (10)

[0073] (11)

[0074] where is the Euclidean distance from the adjacent point to the center point. , , are respectively the direction angles of the straight line connecting the center point and the adjacent point with the XOY plane, XOZ plane and YOZ plane.

[0075] (2) NeRF Focus Aggregation

[0076] The present invention represents the local point cloud scene as a 7D vector and aggregates local feature information through NeRF (Neural Radiance Field). Specifically, the local point cloud learns the NeRF 7D representation, including spatial density and adaptive weights. The NeRF learning process of the present invention is shown in the following formula (12):

[0077] (12)

[0078] For convenience, the present invention represents the coordinate and direction information as the following formulas (13), (14):

[0079] (13)

[0080] (14)

[0081] Among them, represents the relative coordinates, represents the direction angle.

[0082] In order to limit the positional relationship in the local space of network learning, the initial input is only the relational function of relative positions, and the process formula is as shown in (15):

[0083] (15)

[0084] Among them, represents the relative position encoding.

[0085] Considering that the local point cloud density distribution is uneven, and the density has a great correlation with the distance, the spatial position encoding is related to the distance. By constraining the spatial position encoding and the distance, the density distribution relationship formula of the point cloud is learned as follows in (16):

[0086] (16)

[0087] Among them, represents the learned density.

[0088] Considering the direction of learning the feature aggregation weight, the spatial position encoding and the direction are linked. To learn the point cloud feature weight with an undistorted arrangement, the process formula is as shown in (17):

[0089] (17)

[0090] Among them, represents the learned weight information.

[0091] Due to the complexity and variability of large-scale scenes, learning aggregation is difficult. The present invention uses NeRF 7-D and absolute coordinates. On the basis of the original features, the following formulas (18) and (19) are used to enrich the semantic feature information of the local point cloud scene:

[0092] (18)

[0093] (19)

[0094] Among them, is the extended feature, represents the original feature, represents the final feature.

[0095] Due to the sparsity of the point cloud in space, it is discretized into a point cloud with a certain number of points, and more accurate position information of the point cloud in space is given through position encoding. The process formula is shown in (20):

[0096] (20)

[0097] In this embodiment, in the encoder processing, a feature affine transformation module is applied to adaptively correct the features of the local point cloud region.

[0098] It can be understood that in view of the feature learning challenges brought by the irregular geometric structure and uneven density of the local point cloud, the present invention proposes an adaptive feature correction method based on a high-dimensional semantic centroid, as Figure 7 shown. The core lies in mapping the features with non-normal distribution to the optimized feature space through affine transformation. The following formulas (21) and (22) are used to adaptively change the local point features for affine transformation:

[0099] (21)

[0100] Where is the high-dimensional semantic centroid of the region, represents the original feature.

[0101] (22)

[0102] Where, is the variance of the sample. N represents the number of points, K represents the number of selected neighbors, and D represents the dimension of the feature. The feature after affine transformation can be obtained through the following formula (23):

[0103] (23)

[0104] Where, is a numerically stable decimal, and at the same time and are learnable parameters with a dimension of 2D. By calculating the local semantic elements that conform to the normal distribution, the local point cloud semantic features can be scaled to an appropriate range. This can alleviate the differences in searching the local domain of the center point using the KNN algorithm. Considering the particularity of the center point feature, it can be cascaded with the feature after affine transformation, as shown in formula (24):

[0105] (24)

[0106] Finally, the feature affine module generates the feature after affine transformation with rich information .

[0107] In summary, in the semantic segmentation of point clouds, how to make the method of this paper better understand the semantic categories of point clouds has always been the focus of attention in the field. Based on a large-scale point cloud dataset, the present invention adopts a weakly supervised semantic segmentation method, which needs to solve the problems of unclear semantic information of point clouds in space, incorrect segmentation of point clouds at the category boundary, and fuzzy semantic information. Based on this, the present invention proposes a method of point center and high-dimensional position encoding. First, through two feature embedding processes, the present invention enriches the neighborhood feature information of the labeled points and expands the receptive field of the labeled points. Second, by designing a high-dimensional position encoding module, while ensuring that the dimension of the input point cloud information remains unchanged, the model structure of the position encoding is changed. Combining the relevant content of the neural radiance field, it is raised from three dimensions to seven dimensions, which can fully represent the local features of the point cloud in space and can achieve the effect of having discriminative ability for different rotational deformations of the same target object through self-learning. Then, considering the sparsity and irregular geometric structure of the local area of the point cloud, directly processing the features learned by the previous module may affect the final accuracy and reduce the overall stability, so a feature affine module is designed.

[0108] Embodiment 2

[0109] Please refer to Figure 8 , which shows a schematic structural diagram of a weakly supervised point cloud semantic segmentation system proposed in the second embodiment of the present application. The system includes the following key modules:

[0110] An input module 100 for receiving input point cloud data;

[0111] A processing module 200 for performing encoder processing and decoder processing on the input point cloud data;

[0112] The processing module 200 includes an encoder unit 210 and a decoder unit 220;

[0113] The encoder unit 210 is configured to perform multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize a Transformer block including an attention mechanism;

[0114] The encoder unit 210 includes:

[0115] A point center attention sub-unit 211 configured to apply a point center attention mechanism, and the point center attention mechanism includes: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished from the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted by the first feature embedding process;

[0116] A position encoding subunit 212 is configured to generate a high-dimensional semantic position encoding for a local point cloud region; and integrate the high-dimensional semantic position encoding into the calculation of the attention mechanism;

[0117] A feature affine transformation subunit 213 is configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region;

[0118] The decoder unit 220 is configured to perform decoder processing on the features processed by the encoder to obtain a densely output semantic segmentation label;

[0119] An output module 300 is used to output the semantic segmentation label.

[0120] A weak supervision point cloud semantic segmentation system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiment of the present application.

[0121] A weak supervision point cloud semantic segmentation system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0122] A weak supervision point cloud semantic segmentation system provided by an embodiment of the present application can implement Figure 1 each process implemented by a weak supervision point cloud semantic segmentation method in the method embodiment. To avoid repetition, it will not be elaborated here.

[0123] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned weak supervision point cloud semantic segmentation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0124] An embodiment of the present application also provides a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the various processes of the above embodiment of a weakly supervised point cloud semantic segmentation method are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0125] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0126] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may also be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the above embodiment can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0128] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A weakly supervised point cloud semantic segmentation method, characterized in that, It includes the following steps: Receiving input point cloud data; Performing encoder processing on the input point cloud data, where the encoder includes multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, and feature extraction and aggregation utilize Transformer blocks containing an attention mechanism; In the encoder processing, applying a point-centered attention mechanism, which includes: Performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process; In the encoder processing, generating a high-dimensional semantic position encoding for the local point cloud region; and integrating the high-dimensional semantic position encoding into the calculation of the point-centered attention mechanism; In the encoder processing, applying a feature affine transformation module to adaptively correct the features of the local point cloud region; Performing decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels.

2. The method according to claim 1, wherein The first feature embedding process specifically includes: Using a one-dimensional linear layer to calculate the central weights of the center points, where the dimension of the center point features is C and the dimension of the central weights is 1; Exploring adjacent points for the center points using K-nearest neighbors; For the adjacent points in a neighborhood of the center point, aggregating the features of the adjacent points according to the sequential positions belonging to and spanning multiple neighborhoods to calculate the global features; and, Normalizing the global features through a normalized exponential function to obtain neighborhood-specific adjacent point weights for the second feature embedding process.

3. The method according to claim 1 or 2, characterized in that, The second feature embedding process specifically includes: Using the global features extracted in the first feature embedding process as the weights of the adjacent points; and, Performing matrix multiplication on the center point features and the global features to generate a representation of sharing the center point features to the adjacent points.

4. The method according to claim 1, characterized in that, The generation of the high-dimensional semantic position encoding includes: For an adjacent point relative to a center point in the local point cloud region, determining the relative XYZ coordinates of the adjacent point relative to the center point; Determining the Euclidean distance from the adjacent point to the center point; Determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOY plane xy ; Determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOZ plane xz ; Determine the direction angle ρ formed by the line connecting the center point and the adjacent point and the YOZ plane yz ; and Combining the relative XYZ coordinates, the Euclidean distance, and the above three direction angles to form the high-dimensional semantic position encoding.

5. The method according to claim 4, wherein The high-dimensional semantic position encoding contains 7 dimensions, including 3 relative XYZ coordinate dimensions, 1 Euclidean distance dimension, and 3 direction angle dimensions.

6. The method according to claim 1 or 4, characterized in that, The integration of the high-dimensional semantic position encoding into the calculation of the attention mechanism includes: Using a linear layer to transform the global features extracted in the first feature embedding process; Combining the transformed global features, the representation generated in the second feature embedding process, and the high-dimensional semantic position encoding to calculate the original attention scores; and, Normalize the original attention scores through a normalization exponential function to calculate the final attention weights of the point center attention mechanism.

7. The method according to claim 1, characterized in that The application feature affine transformation module specifically includes: Calculate the high-dimensional semantic centroid of the local point cloud region, where the high-dimensional semantic centroid is based on the point features within the region and the high-dimensional semantic position encoding; Calculate the feature variance within the local point cloud region; Based on the high-dimensional semantic centroid and the feature variance, perform an affine transformation on the point features within the local point cloud region using learnable parameters of dimension 2D; and, Concatenate the features after the affine transformation with the center point features to generate local features.

8. A weakly supervised point cloud semantic segmentation system, characterized in that, Include: An input module for receiving input point cloud data; A processing module for performing encoder processing and decoder processing on the input point cloud data; The processing module includes an encoder unit and a decoder unit; The encoder unit is configured to perform multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks containing an attention mechanism; The encoder unit includes: A point center attention sub-unit configured to apply a point center attention mechanism, where the point center attention mechanism includes: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process; A position encoding sub-unit configured to generate the high-dimensional semantic position encoding of the local point cloud region; and, incorporate the high-dimensional semantic position encoding into the calculation of the attention mechanism; A feature affine transformation sub-unit configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region; The decoder unit is configured to perform decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels; An output module for outputting the semantic segmentation labels.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the steps of a weakly supervised point cloud semantic segmentation method according to any one of claims 1-7.

10. A readable storage medium, characterized in that, The program or instruction is stored on the readable storage medium, and when the program or instruction is executed by the processor, it implements the steps of a weakly supervised point cloud semantic segmentation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Road point cloud segmentation method based on multi-task learning

    CN117576400A

  • Point cloud semantic segmentation method for improving RandLA-Net network

    CN119579879A