Online Dense Point Cloud Semantic Segmentation System and Method Incorporating Temporal Features

Through the online dense point cloud semantic segmentation method that integrates timing features in the point cloud semantic segmentation system, the problem of inconsistent point cloud semantic segmentation results and high computational cost in the existing technology is solved, and more accurate and consistent segmentation results and the effect of reducing computational costs are achieved, meeting the real-time processing needs of autonomous driving.

CN115116013BActive Publication Date: 2025-06-13SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110305389.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-19
Publication Date
2025-06-13
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

The existing point cloud semantic segmentation technology is prone to produce inconsistent segmentation results when processing point clouds of continuous frames, and the calculation cost is high, which cannot meet the accuracy requirements and real-time processing requirements of autonomous driving decisions.

Method used

A online dense point cloud semantic segmentation system that integrates timing features is proposed. Through an adaptive frame scheduler, static segmentation module, timing feature aggregation network and partial feature update network, the temporal correlation aggregation features between continuous point cloud frames is used to enhance the accuracy and consistency of segmentation results and reduce computational costs.

Benefits of technology

It realizes more accurate and consistent point cloud semantic segmentation results, reduces computing costs, can process point cloud time series in real time, and meets the accuracy requirements of autonomous driving decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116013B_ABST
    Figure CN115116013B_ABST
Patent Text Reader

Abstract

An online dense point cloud semantic segmentation system and method integrating temporal features, including: an adaptive frame scheduler, a static segmentation module, a temporal feature aggregation network, and a partial feature update network. The adaptive frame scheduler determines whether the next frame is regarded as a key frame or a non-key frame based on the ratio of updated parts to non-updated parts of the information from the partial feature update network. The static segmentation module extracts features from the original point cloud of the key frame through static point cloud segmentation by its backbone network to obtain the semantic segmentation result of a single frame. The temporal feature aggregation network utilizes the temporal correlation between two adjacent frames to precisely optimize the semantic segmentation result by aggregating the features on the key frame. The partial feature update network quickly updates the semantic segmentation result by selectively updating the inherited features according to the locally important key points of the non-key frame between two adjacent frames. The present invention aggregates features by utilizing the temporal correlation between consecutive point cloud frames on the basis of simultaneously considering feature consistency and computational cost. The features of each frame after aggregation are enhanced, making the segmentation more accurate and the features between frames consistent, and eliminating the flicker when processing a series of point clouds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of information processing, specifically an online dense point cloud semantic segmentation system and method integrating temporal features. Background Art

[0002] To better perceive the driving environment, most autonomous vehicles are equipped with lidar sensors to continuously obtain point cloud data, and the performance of point cloud semantic segmentation algorithms is crucial for autonomous vehicles to make correct decisions in real time. Due to the discreteness and irregular distribution of point cloud data, compared with semantic segmentation of images, point cloud semantic segmentation is a more difficult task. Practical point cloud semantic segmentation methods should meet the following two requirements. First, the segmentation results should be accurate so that autonomous vehicles can make correct driving decisions based on the segmentation results. Second, the method should be able to process the time series of point clouds in real time. Summary of the Invention

[0003] Aiming at the problems that the existing point cloud semantic segmentation technology mainly focuses on a single, static point cloud, is prone to inconsistent segmentation results and high computational costs when processing consecutive frames of point clouds, the present invention proposes an online dense point cloud semantic segmentation system and method integrating temporal features. On the basis of considering both feature consistency and computational cost, it aggregates features using the temporal correlation between consecutive point cloud frames. After aggregation, the features of each frame are enhanced, making the segmentation more accurate and the features between frames consistent, and eliminating the flicker when processing the point cloud series. The present invention as a whole solves the deficiencies of the existing technology that the accuracy of semantic segmentation results cannot meet the requirements of autonomous driving decisions and cannot process the time series of point clouds in real time. Compared with the existing technology, the present invention realizes a lightweight and easily deployable online point cloud semantic segmentation framework (TempNet) and a temporal feature aggregation network, which simultaneously utilize the motion information between temporal frames to enhance point cloud features, improve the accuracy of semantic segmentation, and further reduce the calculation by the motion continuity information of multiple frames to improve the speed of semantic segmentation.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to an online dense point cloud semantic segmentation system integrating temporal features, including: an adaptive frame scheduler, a static segmentation module, a temporal feature aggregation network, and a partial feature update network, where: the adaptive frame scheduler determines whether the next frame is regarded as a key frame or a non-key frame based on the ratio of the updated part to the unupdated part of the information from the partial feature update network; the static segmentation module extracts features from the original point cloud of the key frame through static point cloud segmentation by its backbone network to obtain the semantic segmentation result of a single frame; the temporal feature aggregation network utilizes the temporal correlation between two adjacent frames to precisely optimize the semantic segmentation result by aggregating the features on the key frame; the partial feature update network quickly updates the semantic segmentation result by selectively updating the inherited features according to the local important key points of the non-key frame in two adjacent frames.

[0006] The present invention relates to an online dense point cloud semantic segmentation method integrating temporal features based on the above system. By performing complete feature extraction and aggregation on the key frame and propagating the aggregated enhanced features to the non-key frame, when the lightweight difference evaluation detects that the non-key frame contains non-negligible information, further partial update is performed on the non-key frame. Description of the Drawings

[0007] Figure 1 It is a schematic diagram of the system of the present invention;

[0008] Figure 2 It is a schematic diagram of static point cloud segmentation in the embodiment;

[0009] In the figure: (a) is for full feature extraction and aggregation of the key frame, and (b) is for partial feature update and aggregation of the non-key frame;

[0010] Figure 3 It is a schematic diagram of the temporal feature aggregation network in the embodiment;

[0011] In the figure: The key points of two consecutive frames i and frame j are used as inputs. For each key point of frame j, its position is used to find adjacent key points in the space of the previous frame i. The attention convolutional network is used to encode the collected motion information and calculate the residual of feature aggregation;

[0012] Figure 4 It is a schematic diagram of the partial feature update network in the embodiment;

[0013] In the figure: All points P j of the current non-key frame j and the key points of the previous frame i are used as inputs. The weight-sharing convolutional network is used to encode the collected spatial information and calculate the consistency estimator;

[0014] Figure 5 and Figure 6Schematic diagram of the effects of the embodiments. Detailed implementation manners

[0015] As Figure 1 shown, this embodiment relates to an online dense point cloud semantic segmentation system integrating temporal features, including: an adaptive frame scheduler AFS, a static segmentation module, a temporal feature aggregation network TFA, and a partial feature update network PFU, where: the adaptive frame scheduler determines the next frame as a key frame or a non-key frame based on the ratio of the updated part to the unupdated part of the information from the partial feature update network; the static segmentation module extracts features from the original point cloud of the key frame through static point cloud segmentation by its backbone network to obtain the semantic segmentation result of a single frame; the temporal feature aggregation network utilizes the temporal correlation between two adjacent frames to precisely optimize the semantic segmentation result by aggregating the features on the key frame; the partial feature update network quickly updates the semantic segmentation result by selectively updating the inherited features according to the local important key points of the non-key frame between two adjacent frames.

[0016] The described adaptive frame scheduler determines key frames and non-key frames at equal time intervals and dynamically adjusts the number of key frames according to the difference degree of the recently observed non-key frames.

[0017] The described adaptive frame scheduler calculates the ratio of the updated part to the unupdated part in the partial update network: when the ratio is large, it indicates that most points have been updated, and the next frame has a large difference from the current frame and should be regarded as a key frame, otherwise it is a non-key frame.

[0018] This embodiment adopts the point cloud semantic segmentation method RandLA-Net as the static segmentation module.

[0019] The described static segmentation module is implemented by using but not limited to the SqueezeSegV2 point cloud semantic segmentation model. This model takes the original point cloud data as input, extracts the proximity relationship between points, encodes the local spatial geometric structure, and finally obtains the semantic label of each point through neural network processing.

[0020] As Figure 2 shown, the described static point cloud segmentation includes step (a) feature extraction through a pre-trained backbone network and step (b) semantic segmentation by a detection network composed of multiple output branches. The arrows in the figure represent the feature aggregation flow, Figure 2 All in (a) are key frame features, Figure 2 In (b), they are key frame features, partially updated non-key frame features, and inherited features in sequence.

[0021] The aggregation mentioned above means: First, measure the position and feature differences between two consecutive frames caused by motion, and then use them to calculate the attention score. Such an attention score is used as the summary weight of the key points sampled from the two features, so that the key points with consistent motion contribute more to the summary. Specifically: For two consecutive frames i and j, the predicted key-frame feature of frame j where: ⊙ is element-wise multiplication; is the temporal feature aggregation network; W i→j and W j→j are regularization weights, and the predicted key-frame feature recursively aggregates the historical feature and the current feature. H is the point cloud feature vector space (matrix), and W is the weighting coefficient.

[0022] As Figure 3 shown, the temporal feature aggregation specifically includes:

[0023] ① Encode the position difference between two frames through the differential position matrix. The differential position matrix M pos (p j ) = mlp(concat(p i,l , p j ))), where: is the key-point neighbor in the point cloud P j collected by the KNN algorithm for the key point p i . i and j are two consecutive frames, and mlp refers to: multi-layer perceptron.

[0024] The KNN search radius is 1.6, and the maximum number of sampled points is 64.

[0025] ② Encode the feature difference between two frames through the differential feature matrix. The differential feature matrix M fea (p j ) = concat(concat(h i,l , h j ))), where: For each adjacent key point in, connect the encoded relative point position with its corresponding point information to obtain an enhanced feature vector.

[0026] ③ Connect the above two matrices to obtain the motion difference matrix M diff (p j ) = concat(M pos (p j ), M fea (p j ))), where: concat() is the matrix concatenation function. The motion difference matrix M diff(p)Implement feature enhancement of key points.

[0027] In this embodiment, the attention mechanism is adopted to further determine which adjacent points in the motion difference matrix have a greater impact on the current key point. The specific steps are as follows:

[0028] 1) Calculate the attention score: According to the motion difference matrix Learn the unique attention score of each neighboring point by calculating the shared function g(·). This shared function where: H, M, P are sets of vectors, and h, m, p, s are individual vectors in the corresponding sets.

[0029] 2) Perform weighted summation on the attention scores: The representation vector of the key point after update

[0030] The partial feature update network mentioned above determines the spatial consistency index of the calculated feature Q i→j to judge whether the feature H i→j transmitted from the previous frame i is a good approximation of the frame j, so as to selectively update the inherited feature.

[0031] The spatial consistency index mentioned above where: is the spatial correlation detection network implemented by the partial feature update network; P i is the key point of frame i and X i , X j are the point cloud data of frames i and j respectively. For each p i ∈P i , examine the similarity of its inter-frame local spatial features. When the spatial consistency index is less than or equal to the consistency threshold, that is, Q i→j (p i )≤τ, it is considered that there is an inconsistency between the aggregated feature h i and the current feature h j , that is, it means that h i is updated with the applied feature h j .

[0032] As Figure 4 shown, the above-mentioned partial feature update determines the part that needs to be updated through the following steps, which specifically include:

[0033] i) According to all the points Xi, Xj of two adjacent frames i, j and the key points P i of the previous frame i, for each key point p i ∈P i of frame i, search for neighbors in the point cloud through KNN to be spliced into an adjacency matrix wherein: X is a set of geometric coordinates of the input point cloud (x - y - z three - dimensional coordinates), P is a set of attribute features of the input point cloud (such as additional attributes like reflectivity, density, distance, etc.), and H is a set of representation vectors after feature extraction of the point cloud.

[0034] The adjacency matrix mentioned above i.e., the local spatial information encoding matrix, encodes the spatial relationship between the key point P in frame i i and its neighboring points.

[0035] ii) Correspondingly, for each key point in frame j, construct a local spatial information encoding matrix

[0036] iii) According to the local spatial information encoding matrix and construct a convolutional - fully connected layer (ConvFC) to measure the consistency, specifically: where: Q i→j Use the indicator function I(·) to generate the feature update mask U i→j = I(Q i→j ≤ τ), when Q i→j (p i ) meets the threshold requirement, then the current frame inherits the key point p i ∈ P i and its feature vector h i ∈ H i , otherwise discard the feature point.

[0037] The convolutional - fully connected layer includes two 3×3 two - dimensional convolutional layers, and each convolutional layer is followed by a pooling layer to reduce the size to half of the original. In addition, two FC layers are designed to predict the feature consistency index, where: the feature consistency index is restricted to [0,1] through regularization. Therefore, the threshold determined by the mask can be specified between [0,1], and each key point needs to retain its updated mask information. When the threshold is set to 1, the inheritance network is not inherited, and the entire feature vector needs to be calculated through the feature network.

[0038] iv) Re - apply the feature extraction network to the discarded feature points in the current frame, collect local spatial features to supplement the key points, that is, encode the spatial geometric structure in the local space of the point cloud so as to be processed by the neural network and obtain the semantic segmentation label.

[0039] The feature extraction network mentioned above is implemented by but not limited to RandLA - Net.

[0040] The described partial feature update propagation mechanism satisfies: where: U i→j is a binary variable of 0 - 1. When the consistency index is greater than the preset threshold, it is 1; otherwise, it is 0. When the value of U is taken as 0, the model inherits features from before. When the value of U is taken as 1, the model re - extracts features at the current time point. represents the predicted value.

[0041] The number of key frames in the described adaptive frame scheduler is dynamically determined by the consistency estimate value in the nearest non - key frame. Specifically: To determine whether the current frame i should be regarded as a key frame, the ratio of the updated number of key points to the total number of key points is used where: Frame k is the previous key frame, N i is the number of key points in the current frame i. When r k→i is greater than the threshold η, the frequency of setting key frames is reduced; otherwise, the frequency of setting key frames is increased, thereby saving computational costs.

[0042] As Figure 5 shown, it is the quantization results of the two models of the present invention, TempNet and SqueezeSegV2, on consecutive frames. The semantic segmentation effect of the present invention is better.

[0043] As Figure 6 shown, it is the comparison between the present invention's TempNet and two temporal processing algorithms (DFA, dense feature aggregation, aggregating information of all frames; DFP, direct feature propagation, directly copying the features of the previous frame for the processing of the latter frame). The present invention is biased towards the upper right corner, indicating that it can achieve feature aggregation enhancement well and quickly.

[0044] In summary, through the online point cloud series semantic segmentation framework TempNet, the present invention has the characteristics of being lightweight and easy to implement on the existing single - frame segmentation scheme; through the temporal feature aggregation network, it effectively aggregates two point cloud frames in motion by using the continuity of motion and attention pooling.

[0045] Those skilled in the art can make local adjustments to the above - mentioned specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above - mentioned specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.

Claims

1. An online dense point cloud semantic segmentation system integrating temporal features, characterized in that, it includes: an adaptive frame scheduler, a static segmentation module, a temporal feature aggregation network, and a partial feature update network, where: the adaptive frame scheduler determines whether the next frame is a key frame or a non-key frame based on the ratio of the updated part to the unupdated part of the information from the partial feature update network; the static segmentation module extracts features from the original point cloud of the key frame through static point cloud segmentation by its backbone network to obtain the semantic segmentation result of a single frame; the temporal feature aggregation network utilizes the temporal correlation between two adjacent frames to precisely optimize the semantic segmentation result by aggregating the features on the key frame; the partial feature update network quickly updates the semantic segmentation result by selectively updating the inherited features according to the local important key points of the non-key frame in two adjacent frames; the adaptive frame scheduler determines key frames and non-key frames at equal time intervals and dynamically adjusts the number of key frames according to the degree of difference of the most recently observed non-key frames; The aggregation mentioned above refers to: first measuring the position and feature differences between two consecutive frames caused by motion, and then using them to calculate the attention score; such an attention score is used as the summary weight of the key points sampled from the two features, so that the key points with consistent motion contribute more to the summary. Specifically, for two consecutive frames i and j, the predicted key-frame feature of frame j where: ⊙ is element-wise multiplication; is the temporal feature aggregation network; W i→j and W j→j are regularization weights, and the predicted key-frame feature recursively aggregates the historical feature and the current feature. The matrix H is the point cloud feature vector space, and W is the weighting coefficient; the temporal feature aggregation specifically includes: ①Encode the position difference between two frames through a differential position matrix, the differential position matrix M pos (p j ) = mlp(concat(p i,l , p j ))), where: are the key-point neighbors of the point cloud P j of the key point p i in, i and j are two consecutive frames, and mlp refers to: multi-layer perceptron; ②Encode the feature difference between two frames through the differential feature matrix, the differential feature matrix M fea (p j ) = concat(concat(h i,l , h j ))), where: For each adjacent key point in, connect the encoded relative point position with its corresponding point information to obtain an enhanced feature vector; ③ Connect the above two matrices to obtain the motion difference matrix M diff (p j ) = concat(M pos (p j ), M fea (p j )) where: concat() is the matrix concatenation function, and the motion difference matrix M diff (p) realizes the feature enhancement of key points; the motion difference matrix further determines the adjacent points that have a greater impact on the current key point through an attention mechanism. The specific steps are: 1) Calculate the attention score: Based on the motion difference matrix Learn the unique attention score for each neighboring point by calculating the shared function g(·), and the shared function where: H, M, P are sets of vectors, and h, m, p, s are individual vectors in the corresponding sets; 2) Weighted summation of attention scores: The representation vector after key point update The described partial feature update network calculates the spatial consistency index of feature Q i→j to determine whether the feature H i→j transmitted from the previous frame i is a good approximation of frame j, and thus selectively updates the inherited feature; The spatial consistency index Wherein: is a spatial correlation detection network implemented by a partial feature update network; P i is the key point of frame i and X i , X j are the point cloud data of frames i and j respectively. For each p i ∈P i , check the similarity of the inter-frame local spatial features. When the spatial consistency index is less than or equal to the consistency threshold, that is, Q i→j (p i )≤τ, it is considered that there is an inconsistency between the aggregated feature h i and the current feature h j , that is, it means that h i applies the feature h j for update; the partial feature update determines the part that needs to be updated through the following steps, specifically including: i) Based on all points Xi, Xj of two adjacent frames i, j and the key points P of the previous frame i i , for each key point p of frame i i ∈P i , search for neighbors in the point cloud through KNN to splice into an adjacency matrix where: x i,l ∈X i , X is the set of geometric coordinates of the input point cloud (x-y-z three-dimensional coordinates), P is the set of attribute features of the input point cloud, and H is the set of representation vectors after the point cloud undergoes feature extraction; The adjacency matrix described That is, the local spatial information encoding matrix encodes the spatial relationship between the key point p in frame i i and its neighboring points; ii) Correspondingly, for each key point of frame j, construct a local spatial information encoding matrix (iii) Encode the matrix according to local spatial information and Construct a convolutional - fully connected layer (ConvFC) to measure consistency, specifically: where: Q i→j Use the indicator function I(·) to generate a feature update mask U i→j = I(Q i→j ≤ τ), when Q i→j (p i ) meets the threshold requirement, then the current frame inherits the key point p i ∈ P i and its feature vector h i ∈ H i , otherwise discard the key point p i ; the convolutional - fully connected layer includes two 3×3 two-dimensional convolutional layers, and each convolutional layer is followed by a pooling layer. After reducing the size to half of the original, two layers of FC are used to predict the feature consistency index, where: the feature consistency index is restricted to [0,1] through regularization; therefore, the threshold determined by the mask can be specified between [0,1], and each key point needs to retain its updated mask information; iv) Reapply the feature extraction network to the discarded feature points in the current frame, collect local spatial features to supplement the key points, that is, encode the spatial geometric structure in the local space of the point cloud so that it can be processed by the neural network to obtain semantic segmentation labels.

2. The online dense point cloud semantic segmentation system integrating temporal features according to claim 1, characterized in that, the static point cloud segmentation includes step (a) feature extraction through a pre-trained backbone network and step (b) semantic segmentation by a detection network composed of multiple output branches.

3. The online dense point cloud semantic segmentation system integrating temporal features according to claim 1, characterized in that, The described partial feature update propagation mechanism satisfies: Where: U i→j is a binary variable of 0-1. When the consistency index is greater than the preset threshold, it is 1; otherwise, it is 0. When the value of U is 0, the model inherits features from before. When the value of U is 1, the model reextracts features at the current time point. represents the predicted value.

4. The online dense point cloud semantic segmentation system integrating temporal features according to claim 1, characterized in that, The number of key frames in the described adaptive frame scheduler is dynamically determined by the consistency estimate value in the most recent non-key frame. Specifically, to determine whether the current frame i should be regarded as a key frame, the ratio of the updated number of key points to the total number of key points is used. where: frame k is the previous key frame, N i is the number of key points in the current frame i. When r k→i is greater than the threshold η, the frequency of setting key frames is reduced; otherwise, the frequency of setting key frames is increased, thereby saving computational costs.