Weak supervision point cloud semantic segmentation method and system, electronic equipment and medium

By introducing point-center attention mechanism, high-dimensional semantic position coding module and feature affine transformation module in the weakly supervised point cloud semantic segmentation method, the problems such as insufficient utilization of labeled information, blurred boundaries, and misjudgment in the existing methods are solved, and higher segmentation accuracy and robustness are achieved.

CN120236084AActive Publication Date: 2025-07-01NANCHANG UNIV

Patent Information

Application Number
CN202510714365.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the existing weakly supervised point cloud semantic segmentation methods, there are problems such as insufficient utilization of label information, blurred boundaries, misjudgment of categories, insensitive to rotation, and irregularity of local area of ​​point clouds, resulting in insufficient segmentation accuracy and robustness.

Method used

A weakly supervised point cloud semantic segmentation method based on point center and high-dimensional position coding is proposed. Through the point center attention mechanism and the high-dimensional semantic position coding module, the semantic information of the annotated points and their neighborhoods is enriched, the model's expression ability and robustness of the point cloud spatial structure and rotation are enhanced, and the feature affine transformation module is used to alleviate the influence of local area density unevenness and geometric structure irregularity on feature learning.

Benefits of technology

It effectively solves the problem of insufficient utilization of labeled information under weak supervision, improves segmentation accuracy and robustness, enhances the model's ability to express point cloud spatial structure and rotation deformation, and improves the stability and accuracy of feature learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236084A_ABST
    Figure CN120236084A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, three-dimensional point cloud processing and weak supervision learning, and discloses a weak supervision point cloud semantic segmentation method and system, electronic equipment and a medium. An innovative point center attention mechanism, a high-dimensional semantic position coding module and a feature affine transformation module are introduced into the encoder; the point center attention mechanism extracts global features from adjacent points and considers a plurality of neighborhoods where the global features are located through two feature embedding processes, and shares center point features to the adjacent points so as to make full use of limited annotation information; high-dimensional semantic position coding is expressed by 7D including relative XYZ coordinates, Euclidean distance and direction angle, local geometric information is effectively captured, robustness to rotational deformation is enhanced, and attention calculation is fused; the feature affine transformation module performs adaptive correction on local point cloud features based on the high-dimensional semantic centroid and the feature variance, and relieves the irregularity influence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, 3D point cloud processing, and weakly supervised learning, and particularly relates to a weakly supervised point cloud semantic segmentation method, system, electronic device, and medium based on point centers and high-dimensional position encoding. Background Art

[0002] With the continuous expansion of the scale of 3D point cloud datasets, the cost of fine per-point annotation of point clouds is extremely high, which has greatly promoted the research on weakly supervised point cloud semantic segmentation methods. Although some existing weakly supervised methods have achieved certain results, due to the scarcity of annotation information, the point cloud semantic segmentation task often faces problems of blurred boundaries and misclassification of categories. Especially in the weakly supervised learning framework, the semantic information contained in spatially limited labeled points is not fully utilized, and it is difficult for the network to comprehensively understand the local and global semantics of the labeled points, which directly affects the accuracy and practicality of the final segmentation result.

[0003] In existing methods, some methods based on point center networks use the relationship between points and their neighborhoods for feature learning and perform well in the fully supervised scenario. However, in the weakly supervised environment, the labeled points are sparse and may be unevenly distributed, resulting in the inability to effectively use the semantic information of the labeled points to update and optimize the unlabeled neighborhoods. In addition, the exploration of the relationship between the center point and neighboring points and the insufficient utilization of global features limit the value of the limited annotation information. Existing local spatial representation methods (such as relative coordinates) are sensitive to the rotational deformation of point clouds and fail to fully capture the characteristics of the same object under different rotations. At the same time, the problems of uneven density and irregular geometric structure in the local regions of point clouds also pose challenges to feature learning. These problems together lead to room for improvement in the segmentation accuracy and robustness of existing weakly supervised point cloud semantic segmentation methods. Summary of the Invention

[0004] Aiming at the problems existing in the existing weakly supervised point cloud semantic segmentation methods, such as insufficient utilization of annotation information, blurred boundaries, misclassification of categories, insensitivity of local spatial representation to rotation, and irregularity of local regions of point clouds, the present invention proposes a weakly supervised point cloud semantic segmentation method and system based on point centers and high-dimensional position encoding.

[0005] In a first aspect, the present invention provides a weakly supervised point cloud semantic segmentation method, including the following steps: Receiving input point cloud data; Performing encoder processing on the input point cloud data, where the encoder includes multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, and feature extraction and aggregation utilize Transformer blocks including an attention mechanism; In the encoder processing, a point-centered attention mechanism is applied, and the point-centered attention mechanism includes: Perform a first feature embedding process to extract global features from adjacent points. The extraction of the global features is distinguished by associating with the central weights of the neighborhoods where the adjacent points are located; and, perform a second feature embedding process to share the central point features to the adjacent points based on the global features extracted in the first feature embedding process; In the encoder processing, generate a high-dimensional semantic position encoding for the local point cloud region; and, incorporate the high-dimensional semantic position encoding into the calculation of the point-centered attention mechanism; In the encoder processing, apply a feature affine transformation module to adaptively correct the features of the local point cloud region; Perform decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels. As an optional implementation manner of the first aspect of this application, the first feature embedding process specifically includes: using a one-dimensional linear layer to calculate the central weights of the central points, the dimension of the central point features is C, and the dimension of the central weights is 1; using K-nearest neighbors to explore adjacent points for the central points; for the adjacent points in a neighborhood of the central points, aggregate the features of the adjacent points according to the sequential positions belonging to and spanning multiple neighborhoods to calculate the global features; and, normalize the global features through a normalized exponential function to obtain neighborhood-specific adjacent point weights for the second feature embedding process.

[0006] As an optional implementation manner of the first aspect of this application, the second feature embedding process specifically includes: using the global features extracted in the first feature embedding process as the weights of the adjacent points; and, performing matrix multiplication on the central point features and the global features to generate a representation in which the central point features are shared to the adjacent points.

[0007] As an optional implementation manner of the first aspect of this application, the generation of the high-dimensional semantic position encoding includes: for an adjacent point in the local point cloud region relative to a central point, determine the relative XYZ coordinates of the adjacent point relative to the central point; determine the Euclidean distance from the adjacent point to the central point; determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOY plane xy ; determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOZ plane xz ; determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the YOZ plane yz ; and, combine the relative XYZ coordinates, the Euclidean distance, and the above three direction angles to form the high-dimensional semantic position encoding.

[0008] As an alternative implementation of the first aspect of the present application, the high-dimensional semantic position encoding includes 7 dimensions, including 3 relative XYZ coordinate dimensions, 1 Euclidean distance dimension, and 3 direction angle dimensions.

[0009] As an alternative implementation of the first aspect of the present application, the integration of the high-dimensional semantic position encoding into the calculation of the attention mechanism includes: using a linear layer to transform the global features extracted in the first feature embedding process; combining the transformed global features, the representations generated in the second feature embedding process, and the high-dimensional semantic position encoding to calculate the original attention scores; and normalizing the original attention scores through a softmax function to calculate the final attention weights of the point-centered attention mechanism.

[0010] As an alternative implementation of the first aspect of the present application, the application of the feature affine transformation module specifically includes: calculating the high-dimensional semantic centroid of the local point cloud region, where the high-dimensional semantic centroid is based on the point features within the region and the high-dimensional semantic position encoding; calculating the feature variance within the local point cloud region; performing an affine transformation on the point features within the local point cloud region using learnable parameters of dimension 2D based on the high-dimensional semantic centroid and the feature variance; and concatenating the features after the affine transformation with the center point features to generate local features.

[0011] In a second aspect, an embodiment of the present application provides a weakly supervised point cloud semantic segmentation system, including: An input module for receiving input point cloud data; A processing module for performing encoder processing and decoder processing on the input point cloud data; The processing module includes an encoder unit and a decoder unit; The encoder unit is configured to perform a plurality of downsampling stages, each downsampling stage including point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks including an attention mechanism; The encoder unit includes: A point-centered attention sub-unit configured to apply a point-centered attention mechanism, where the point-centered attention mechanism includes: performing a first feature embedding process to extract global features from adjacent points, where the extraction of the global features is distinguished by associating with the central weights of the neighborhoods where the adjacent points are located; and performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process; A position encoding subunit, configured to generate a high-dimensional semantic position encoding for a local point cloud region; and, incorporate the high-dimensional semantic position encoding into the calculation of the attention mechanism; A feature affine transformation subunit, configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region; The decoder unit is configured to perform decoder processing on the features processed by the encoder to obtain a densely output semantic segmentation label; An output module for outputting the semantic segmentation label.

[0012] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in: 1. Through the two embedding processes of the point center attention network, the semantic information of the labeled points and their neighborhoods is enriched, effectively solving the problem that it is difficult to update the unlabeled neighborhoods under weak supervision, expanding the receptive field of the labeled points, and improving the utilization rate of limited labeled information.

[0015] 2. The proposed high-dimensional semantic position encoding module, represented by 7 dimensions including relative position, distance, and direction angle, can not only accurately describe the spatial structure of the local point cloud, but also learn the different rotational deformation characteristics of the same object, enhancing the model's expression ability for the point cloud spatial structure and robustness to rotation.

[0016] 3. The designed feature affine module adaptively corrects the local point cloud features based on the high-dimensional semantic centroid and feature variance, effectively alleviating the influence of uneven local point cloud density and irregular geometric structure on feature learning, and improving the stability and accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of an overall network structure of a weakly supervised point cloud semantic segmentation based on point center and high-dimensional position encoding according to an embodiment of the present invention; Figure 2 is a schematic diagram of the structure of a downsampling module according to an embodiment of the present invention; Figure 3 is a schematic diagram of the structure of a Transformer block according to an embodiment of the present invention; Figure 4 is a schematic structural diagram of an upsampling module according to an embodiment of the present invention; Figure 5 is a detailed structural diagram of an enhanced feature module (point center attention network part) according to an embodiment of the present invention; Figure 6 is a structural diagram of a high-dimensional position encoding module according to an embodiment of the present invention; Figure 7 is a schematic structural diagram of a feature affine module according to an embodiment of the present invention Figure 8 is a structural diagram of a weakly supervised point cloud semantic segmentation system based on point center and high-dimensional position encoding according to an embodiment of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0019] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0020] Embodiment 1 The present invention proposes a weakly supervised point cloud semantic segmentation method based on point center and high-dimensional position encoding. The overall technical route of this method is as follows Figure 1 shown. An encoder-decoder architecture is adopted, and the input is the XYZ coordinates and RGB color information of the point cloud. In the encoder part stage, multiple Transformer blocks are integrated into each downsampling stage to extract the basic features of the point cloud. Then, these encoded features are sent to the decoder part, which contains upsampling blocks for restoring the point cloud density and outputting dense semantic segmentation labels. Finally, a classifier is used to obtain the categories after semantic segmentation of the point cloud.

[0021] Among them, the downsampling process is as follows Figure 2As shown, it is mainly used to reduce the number of points in the point cloud, lower the dimension and computational volume of the data, while retaining important feature information. Define the downsampling scale as 4, and reduce the number of points by 4 in each downsampling layer. Then group the selected points and aggregate the corresponding features using max pooling. Among them, Po 、 Fo respectively represent the number of input point clouds and the original features; Pd 、 Fd respectively represent the number of point clouds and the point cloud features after the downsampling process. After the downsampling process, a point set with reduced number but still representative and its features are obtained. The specific module functions are: farthest point sampling (FPS), k-nearest neighbor (KNN), multi-layer perceptron (MLP), and max pooling (Max Pool). Farthest point sampling (FPS) is the first step of downsampling, which constructs a subset by selecting the points farthest from each other in the point cloud. Farthest point sampling can ensure that the sampled points are evenly distributed in space, thus better retaining the overall structural information of the point cloud. K-nearest neighbor (KNN), for the points selected by FPS, find their k nearest neighbor points. The KNN operation can help determine the local area around each sampled point, providing local context information for subsequent feature extraction. Multi-layer perceptron (MLP), used to learn and transform the features of the points in the local area determined by KNN, extracting more meaningful local features. These local features will be used for subsequent processing. Max pooling (Max Pool), within each local area, through the max pooling operation, select the point with the largest feature value as the representative of the area. Max pooling can further reduce the number of points while retaining the most significant feature information. After this series of operations, the downsampled point set and its features are obtained.

[0022] Among them, the Transformer block is as follows Figure 3As shown, it is a module commonly used for processing sequence data and can also be used for feature extraction and transformation in point cloud processing to enhance the model's ability to capture global information. Multilayer perceptron (MLP), in the Transformer block, MLP is usually used to perform non-linear transformation on input features to increase the model's expressive power. The first MLP performs preliminary feature extraction and transformation on the input features to extract more meaningful feature representations. Enhanced Feature Module (CNA), this step includes convolution operations to capture local features, then normalization operations to stabilize the training process, and finally non-linearity is introduced through activation functions to further enhance the model's learning ability. The second MLP further transforms and extracts the features after being processed by CNA to further optimize the feature representation and make it more suitable for subsequent tasks. Dropout is a regularization technique used in neural networks, mainly used to prevent model overfitting. Its core idea is to randomly "discard" (temporarily turn off) a part of the neurons in the neural network during the training process, forcing the network not to overly rely on certain specific nodes, thereby improving the generalization ability. Residual connection structures help solve the problem of gradient disappearance in deep networks, enabling the model to be better trained and optimized. Among them, Pd 、 Fd are the number of point clouds and the point cloud features respectively after the downsampling process; Pd 、 F’d are the number of point clouds and the enhanced features respectively after feature enhancement.

[0023] Among them, the upsampling process is as follows Figure 4 As shown, it is mainly used to increase the density of the point cloud, make the point cloud data more refined, so as to better perform subsequent processing such as semantic segmentation. Interpolation, this is one of the key steps in upsampling. Since the original point set is relatively sparse, through the interpolation algorithm, according to the information of the denser points in the middle, the features of the middle points are estimated and supplemented to obtain a richer feature representation. Concatenate, concatenate the features obtained after interpolation with the original features. In this way, the information of the original features can be retained, and at the same time, the new features obtained through interpolation can be incorporated to provide a more comprehensive feature input for subsequent processing. Multilayer perceptron (MLP), the multilayer perceptron is a deep learning model used to further learn and transform the concatenated features. It can extract higher-level and more abstract features, so as to better meet the requirements of tasks such as semantic segmentation. After being processed by MLP, the final updated features are obtained, which together with the original point set are used as the output of upsampling. Among them Po 、 Fo represent the number of input point clouds and the original features respectively; Pd 、 F’d are the number of point clouds and the enhanced features respectively after feature enhancement; Po 、F’o They are respectively expressed as the number of input point clouds and the final features after upsampling.

[0024] In this embodiment, the received point cloud data is processed by an encoder, which includes multiple downsampling stages. Each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks containing an attention mechanism. In the encoder processing, a point-centered attention mechanism is applied, which includes: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished from the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process.

[0025] It can be understood that in the weakly supervised point cloud segmentation task, the core challenge lies in: when training only relying on a limited number of labeled points, how to efficiently optimize the features of unlabeled points, and then achieve accurate segmentation of the entire point cloud. Existing point-centered networks use a local attention mechanism to calculate weights separately within each neighborhood. This attention mechanism has achieved remarkable results in fully supervised large-scale point cloud segmentation because the labeled points are often closely connected and have a small distance from each other. However, in a weakly supervised environment, its limitations are also manifested. Therefore, this method proposes a point-centered attention mechanism, which innovatively integrates two embedding processes to address the problem caused by sparse point annotations in weakly supervised point clouds.

[0026] Such as Figure 5 shown Pi 、 Fi They are respectively expressed as the number of point clouds and features input into the network. Pij 、 Fij They are respectively expressed as the number of point clouds and features after KNN. g1(Fi) It is expressed as the result obtained by passing the features of the input points through a linear layer g1 The obtained result is expressed as g2(e1) It is expressed as e1 after passing through a linear layer g2The obtained results show that. Specifically, in the local attention mechanism, the weights of adjacent points within a given neighborhood are calculated only based on the characteristics of that neighborhood itself, without considering the internal connections between different neighborhoods. The features of the central point are integrated by weighted summation of the weights of adjacent points. However, in the case of sparse annotations, the backpropagation of the training loss only depends on a limited number of labeled points. This leads to the situation that only when the central point is labeled, its adjacent unlabeled points can obtain corresponding updates and optimizations. Conversely, if the central point is not labeled while its adjacent points are labeled, this method cannot fully exploit the value of these labeled points. When both the central point and its adjacent points are not labeled, the problem becomes even more serious, and this situation is called an "unlabeled neighborhood". To address the above problems, the method based on centralized attention, through two embedding processes, properly handles the sparse annotation problem. In the initial embedding stage, the problem of unlabeled neighbors is successfully solved, and the relationships between different neighborhoods are comprehensively considered. Previous studies only obtained global features from the central point, while the present invention starts from another perspective, extracts global features based on neighboring points, and then shares them within the neighborhood. Essentially, by using the existing labeled neighboring points, the features of unlabeled points in multiple neighborhoods are improved, regardless of whether the central point is labeled. In this way, even in an unlabeled neighborhood, adjacent points can be enhanced by the global features from other regions that may contain labeled points. Moreover, these advantages also extend to unlabeled central points, which are further optimized through deep fusion with their corresponding neighborhoods, thus significantly improving the performance of the model in the weakly supervised point cloud segmentation task.

[0027] However, the extraction of global features encounters an obstacle in neighborhood construction. In three-dimensional space, a point can act as an adjacent point in multiple neighborhoods with different sequential positions. Therefore, if global features are simply aggregated from all these neighborhoods, the result will be a mixture of all points without any meaningful semantic context. To overcome this problem, the present invention differentiates them by embedding the features of different neighborhoods into corresponding central weights. Specifically, the central weights are derived from the central point features, initially N×C-dimensional, and then transformed into N×1-dimensional, where N represents the number of point clouds and C represents the point cloud feature dimension. The global features represented by neighborhood points are denoted as e1 , which are obtained by aggregating and calculating according to the sequential positions of these features within their respective neighborhoods. In particular, Figure 5 illustrates the proposed attention mechanism with two embedding stages. For the central point Pi , the central weights are calculated through a linear layer g1 with only one dimension. Then, using KNN to explore the adjacent points, , , K neighboring points are obtained. In the first embedding, global features are extracted through the central weights and neighboring point features. These features are then used as the neighboring point weights for the second embedding through the normalized exponential function σ. The process formulas are shown in (1) and (2) as follows: (1) (2) wherein, represents the feature of the i th input point after passing through the linear layer g1 and is represented by the obtained result.

[0028] Although e1 the representation of unlabeled neighboring points in weakly supervised point clouds is improved by valuable global features, it often lacks key local features specific to separated neighborhoods and weakens the influence of the central point. Therefore, compared with previous studies, the effectiveness of calculating adjacent point weights may not be optimal. To address this limitation, the present invention introduces a second embedding process, aiming to share the central point features with their respective neighborhoods through the above global features e1 This method uses these global features as the weights of adjacent points, and then performs matrix multiplication in combination with the central point. Therefore, the second embedding (denoted as e2 ) ensures that the central point features are effectively propagated to each point in the corresponding neighborhood under the guidance of appropriate global features.

[0029] To construct the attention weights of adjacent points covering global features and valuable central weights, the present invention combines these embedding processes with position encoding as spatial supplementation. Specifically, the global features are transformed through g2 a linear layer e1 , and the learnable parameters , b are integrated with the central sharing mechanism e2 . This integration enhances local features through central point sharing while also retaining global features. In addition, the position encoding is incorporated into the attention weights to supplement appropriate geometric features according to the position of the points. Then, the attention weights are multiplied by the value feature map through the normalized exponential function σ . Finally, a weighted sum of all adjacent points is performed to avoid the problem of irregular point sorting and obtain the output feature . The formulas for this process are shown in (3) and (4) as follows: (3) (4) In this embodiment, during the encoder processing, a high-dimensional semantic position encoding of the local point cloud region is generated; and, the high-dimensional semantic position encoding is incorporated into the calculation of the attention mechanism.

[0030] It is understandable that a simple representation of the local point cloud is the relative coordinates of neighboring points with respect to the center. By converting the original three-dimensional coordinates into their relative coordinates, a simple local representation of the local point cloud is obtained, that is, the position relative to the centroid. The present invention expects the local representation to be insensitive to rotational deformation, that is, it can learn different rotational characteristics of the same object. Specifically, as Figure 6 shown, the point cloud representation is supplemented with distances and direction angles around the XYZ axes, where is the Euclidean distance from the neighboring point to the center point, is the direction angle between the line connecting the center point and the neighboring point and the plane formed by the X and Y axes, is the direction angle between the line connecting the center point and the neighboring point and the plane formed by the X and Z axes, is the direction angle between the line connecting the center point and the neighboring point and the plane formed by the Y and Z axes.

[0031] To overcome the problem that the simple relative coordinates are sensitive to rotation, the present invention introduces high-dimensional semantic position encoding (position + angle + distance).

[0032] (1) 7D representation of NeRF A simple representation of the local point cloud is the relative coordinates of neighboring points with respect to the center. Converting its relative coordinates into a simple representation of the local point cloud, are respectively represented as the changed coordinates, corresponding to formulas (5), (6), (7): (5) (6) (7) The present invention proposes a method for representing local point clouds based on centroid reference. By establishing a relative coordinate system with the local centroid as the origin, the original point cloud is converted into a position-normalized local representation. That is to say, different rotational properties of the same object can be learned. Specifically, as shown in formulas (8), (9), (10), (11), the point cloud representation is supplemented with distances and direction angles around the XYZ axes: (8) (9) (10) (11) where is the Euclidean distance from the adjacent point to the center point. , , are respectively the direction angles between the line connecting the center point and the adjacent point and the XOY plane, XOZ plane, and YOZ plane.

[0033] (2)NeRF Focus Aggregation The present invention represents the local point cloud scene as a 7D vector and aggregates local feature information through NeRF (Neural Radiance Field). Specifically, the local point cloud learns the NeRF 7D representation, including spatial density and adaptive weights. The NeRF learning process of the present invention is shown in the following formula (12): (12) For convenience, the present invention represents the coordinate and direction information as the following formulas (13) and (14): (13) (14) Wherein, represents the relative coordinate, represents the direction angle.

[0034] To limit the network from learning the position relationship in the local space, the initial input is only the relational function of the relative position, and the process formula is as shown in (15): (15) Wherein, represents the relative position encoding.

[0035] Considering that the local point cloud density distribution is uneven and the density has a great correlation with the distance, the spatial position encoding is related to the distance. By constraining the spatial position encoding and the distance, the density distribution relationship formula (16) of the point cloud is learned as follows: (16) Wherein, represents the learned density.

[0036] Considering the direction of learning the feature aggregation weight, the spatial position encoding and the direction are linked. To learn the point cloud feature weight with undistorted arrangement, the process formula is as shown in (17): (17) Wherein, represents the learned weight information.

[0037] Due to the complexity and variability of large scenes, learning aggregation is difficult. The present invention uses NeRF 7-D and absolute coordinates, and on the basis of the original features, uses the following formulas (18) and (19) to enrich the semantic feature information of the local point cloud scene: (18) (19) Wherein, is the extended feature, Represents the original feature, represents the final feature.

[0038] Due to the sparsity of the point cloud in space, it is discretized into a point cloud with a certain number of points, and more accurate position information of the point cloud in space is given through position encoding. The process formula is as shown in (20): (20) In this embodiment, in the encoder processing, a feature affine transformation module is applied to adaptively correct the features of the local point cloud region.

[0039] It can be understood that in view of the feature learning challenges brought by the irregularity of the local point cloud geometry and the non-uniform density, the present invention proposes an adaptive feature correction method based on the high-dimensional semantic centroid, as Figure 7 shown. The core lies in mapping the features with non-normal distribution to the optimized feature space through affine transformation. The following formulas (21) and (22) are used to adaptively change the local point features for affine: (21) where is the high-dimensional semantic centroid of the region, represents the original feature.

[0040] (22) where, is the variance of the sample. N represents the number of points, K represents the number of selected neighbors, and D represents the dimension of the feature. The feature after affine can be obtained through the following formula (23): (23) where, is a numerically stable decimal, and at the same time and are learnable parameters with a dimension of 2D. By calculating the local semantic elements that conform to the normal distribution, the local point cloud semantic features can be scaled to an appropriate range. This can alleviate the differences in searching the local domain of the center point using the KNN algorithm. Considering the particularity of the center point features, they can be cascaded with the features after affine, as shown in formula (24): (24) Finally, the feature affine module generates the features after affine with rich information .

[0041] In summary, in the semantic segmentation of point clouds, how to make the method of this paper better understand the semantic categories of point clouds has always been the focus of attention in the field. Based on a large-scale point cloud dataset, the present invention adopts a weakly supervised semantic segmentation method, which needs to solve the problems of unclear semantic information of point clouds in space, incorrect segmentation of point clouds at the category boundary, and fuzzy semantic information. Based on this, the present invention proposes a method of point center and high-dimensional position encoding. First, through two feature embedding processes, the present invention enriches the neighborhood feature information of the labeled points and expands the receptive field of the labeled points. Secondly, by designing a high-dimensional position encoding module, while keeping the dimension of the input point cloud information unchanged, the model structure of the position encoding is changed. Combining the relevant content of the neural radiance field, it is raised from three dimensions to seven dimensions, which can fully represent the local features of the point cloud in space and can achieve the effect of being discriminative for different rotational deformations of the same target object through self-learning. Then, considering the sparsity and irregular geometric structure of the local area of the point cloud, directly processing the features learned by the previous module may affect the final accuracy and reduce the overall stability, so a feature affine module is designed.

[0042] Embodiment 2 Please refer to Figure 8 , which shows a schematic structural diagram of a weakly supervised point cloud semantic segmentation system proposed in the second embodiment of the present application. The system includes the following key modules: Input module 100, configured to receive input point cloud data; Processing module 200, configured to perform encoder processing and decoder processing on the input point cloud data; The processing module 200 includes an encoder unit 210 and a decoder unit 220; The encoder unit 210 is configured to perform a plurality of downsampling stages, each downsampling stage including point sampling, neighborhood construction, feature extraction and aggregation, wherein feature extraction and aggregation utilize a Transformer block including an attention mechanism; The encoder unit 210 includes: Point center attention sub-unit 211, configured to apply a point center attention mechanism, the point center attention mechanism including: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished from the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted by the first feature embedding process; Position encoding sub-unit 212, configured to generate high-dimensional semantic position encoding of the local point cloud region; and, integrating the high-dimensional semantic position encoding into the calculation of the attention mechanism; The feature affine transformation subunit 213 is configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region; The decoder unit 220 is configured to perform decoder processing on the features processed by the encoder to obtain a densely output semantic segmentation label; The output module 300 is used to output the semantic segmentation label.

[0043] A weakly supervised point cloud semantic segmentation system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., and the embodiments of the present application do not make specific limitations.

[0044] A weakly supervised point cloud semantic segmentation system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application do not make specific limitations.

[0045] A weakly supervised point cloud semantic segmentation system provided in an embodiment of the present application can implement Figure 1 each process implemented by a weakly supervised point cloud semantic segmentation method in the method embodiment. To avoid repetition, it will not be elaborated here.

[0046] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned weakly supervised point cloud semantic segmentation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0047] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned weakly supervised point cloud semantic segmentation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0048] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0049] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0050] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0051] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A weakly supervised point cloud semantic segmentation method, characterized in that, It includes the following steps: Receive input point cloud data; Perform encoder processing on the input point cloud data. The encoder includes multiple downsampling stages, and each downsampling stage includes point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks containing an attention mechanism; In the encoder processing, apply a point-centered attention mechanism, which includes: Execute the first feature embedding process to extract global features from adjacent points. The extraction of the global features differentiates the features of different neighborhoods by associating with the central weights of the neighborhoods where the adjacent points are located; and, execute the second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process; In the encoder processing, generate high-dimensional semantic position encoding for the local point cloud region; and incorporate the high-dimensional semantic position encoding into the calculation of the point-centered attention mechanism; In the encoder processing, apply a feature affine transformation module to adaptively correct the features of the local point cloud region; Perform decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels.

2. The method according to claim 1, characterized in that The first feature embedding process specifically includes: Use a one-dimensional linear layer to calculate the central weights of the center points. The dimension of the center point features is C, and the dimension of the central weights is 1; Explore adjacent points for the center point using K-nearest neighbors; For the adjacent points in a neighborhood of the center point, aggregate the features of the adjacent points according to the sequential positions belonging to and spanning multiple neighborhoods to calculate the global features; and, Normalize the global features through a normalized exponential function to obtain neighborhood-specific adjacent point weights for the second feature embedding process.

3. The method according to claim 1 or 2, characterized in that The second feature embedding process specifically includes: Use the global features extracted in the first feature embedding process as the weights of the adjacent points; and, Perform matrix multiplication on the center point features and the global features to generate a representation in which the center point features are shared to the adjacent points.

4. The method according to claim 1, wherein The generation of the high-dimensional semantic position encoding includes: For an adjacent point relative to a center point in the local point cloud region, determine the relative XYZ coordinates of the adjacent point relative to the center point; Determine the Euclidean distance from the adjacent point to the center point; Determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOY plane xy ; Determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the XOZ plane xz ; Determine the direction angle ρ formed by the line connecting the central point and the adjacent point and the YOZ plane yz ; and Combine the relative XYZ coordinates, the Euclidean distance, and the above three direction angles to form the high-dimensional semantic position encoding.

5. The method according to claim 4, wherein The high-dimensional semantic position encoding contains 7 dimensions, including 3 relative XYZ coordinate dimensions, 1 Euclidean distance dimension, and 3 direction angle dimensions.

6. The method according to claim 1 or 4, characterized in that, The incorporation of the high-dimensional semantic position encoding into the calculation of the attention mechanism includes: Use a linear layer to transform the global features extracted in the first feature embedding process; Combine the transformed global features, the representation generated in the second feature embedding process, and the high-dimensional semantic position encoding to calculate the original attention scores; and, Normalize the original attention scores through the softmax function to calculate the final attention weights of the point center attention mechanism.

7. The method according to claim 1, wherein The application feature affine transformation module specifically includes: Calculate the high-dimensional semantic centroid of the local point cloud region, where the high-dimensional semantic centroid is based on the point features within the region and the high-dimensional semantic position encoding; Calculate the feature variance within the local point cloud region; Based on the high-dimensional semantic centroid and the feature variance, perform an affine transformation on the point features within the local point cloud region using learnable parameters of dimension 2D; and, Concatenate the features after the affine transformation with the center point features to generate local features.

8. A weakly supervised point cloud semantic segmentation system, characterized in that, Include: An input module for receiving input point cloud data; A processing module for performing encoder processing and decoder processing on the input point cloud data; The processing module includes an encoder unit and a decoder unit; The encoder unit is configured to perform multiple downsampling stages, each downsampling stage including point sampling, neighborhood construction, feature extraction, and aggregation, where feature extraction and aggregation utilize Transformer blocks containing an attention mechanism; The encoder unit includes: A point center attention sub-unit configured to apply a point center attention mechanism, where the point center attention mechanism includes: performing a first feature embedding process to extract global features from adjacent points, and the extraction of the global features is distinguished by associating with the central weights of the neighborhoods where the adjacent points are located; and, performing a second feature embedding process to share the center point features to the adjacent points based on the global features extracted in the first feature embedding process; A position encoding sub-unit configured to generate a high-dimensional semantic position encoding for the local point cloud region; and, incorporate the high-dimensional semantic position encoding into the calculation of the attention mechanism; A feature affine transformation sub-unit configured to apply a feature affine transformation module to adaptively correct the features of the local point cloud region; The decoder unit is configured to perform decoder processing on the features after the encoder processing to obtain densely output semantic segmentation labels; An output module for outputting the semantic segmentation labels.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored on the memory and executable on the processor. When the program or instructions are executed by the processor, the steps of a weakly supervised point cloud semantic segmentation method according to any one of claims 1-7 are implemented.

10. A readable storage medium, characterized in that, The program or instructions are stored on the readable storage medium. When the program or instructions are executed by the processor, the steps of a weakly supervised point cloud semantic segmentation method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Road point cloud segmentation method based on multi-task learning

    CN117576400A

  • Point cloud semantic segmentation method for improving RandLA-Net network

    CN119579879A

  • Point cloud semantic segmentation method and system based on local neighborhood attention

    CN119785032A

  • Local context feature extraction module for semantic segmentation in 3D point cloud scenario

    WO2024113078A1

Cited By

  • Point cloud weak supervision semantic segmentation method based on dual self-attention mechanism

    CN120580440A