Dynamic scene-oriented point cloud semantic enhancement segmentation method

Through a multi-level spatiotemporal feature aggregation module, a dynamic interaction relationship modeling unit, and a lightweight post-processing optimization strategy, the problems of insufficient dynamic target capture capability and multi-category object segmentation accuracy in dynamic scenes are solved, and efficient and accurate point cloud semantic segmentation is achieved, which is suitable for autonomous driving and robot navigation.

CN120599259APending Publication Date: 2025-09-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510692925.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing point cloud semantic segmentation methods have insufficient ability to capture fast-moving targets in dynamic scenes, limited segmentation accuracy of multi-category objects, and insufficient fusion of spatiotemporal information, which affects the accuracy and robustness of the segmentation results.

Method used

It adopts a multi-level spatiotemporal feature aggregation module, a dynamic interaction relationship modeling unit and a lightweight post-processing optimization strategy, captures dynamic target features through an improved attention mechanism and graph convolutional network, optimizes multi-category object segmentation, and improves segmentation accuracy through sparse sampling and fast fusion modules.

Benefits of technology

It significantly improves the ability to capture fast-moving targets in dynamic scenes and the segmentation accuracy of multi-category objects, ensuring efficient and accurate semantic segmentation results, and is suitable for fields such as autonomous driving and robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599259A_ABST
    Figure CN120599259A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of point cloud semantic segmentation, in particular to a dynamic scene-oriented point cloud semantic enhancement segmentation method, which comprises a multi-level spatial-temporal feature aggregation module, a dynamic interaction relationship modeling unit and a lightweight post-processing optimization strategy. Dynamic target feature expression is enhanced through an improved attention mechanism, an interactive relationship between targets is captured by using a graph convolutional network, and a segmentation result is optimized through sparse sampling. The method can remarkably improve the capturing capability of a fast moving target in a dynamic scene and the segmentation precision of multiple types of objects, guarantees the real-time performance, is suitable for the fields of automatic driving, robot navigation and the like, and has high practicability and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and three-dimensional data processing, and specifically is a point cloud semantic enhancement segmentation method for dynamic scenes. Background Art

[0002] With the rapid development of three-dimensional point cloud data processing technology, point cloud semantic segmentation is increasingly being used in dynamic scenarios such as autonomous driving, robot navigation, and smart cities. However, existing point cloud semantic segmentation methods still have many shortcomings when dealing with dynamic scenes. For example, their ability to capture dynamic targets is limited, making it difficult to effectively handle fast-moving targets; the segmentation accuracy of multi-category objects in complex scenes is insufficient, especially when the targets are diverse and unevenly distributed; and the fusion of spatiotemporal information is insufficient, failing to fully utilize the temporal continuity and spatial correlation in dynamic scenes. These issues severely limit the performance of existing methods in real-world dynamic scenes.

[0003] After searching, it was found that the patent with publication number CN115170585B proposed a three-dimensional point cloud semantic segmentation method. The method establishes a neural network that integrates multiple point cloud expressions, combines multi-frame point clouds that have been voxelized as input, uses image information and time series information for semantic segmentation, and post-processes the segmentation results through a clustering algorithm. However, the technical solution mainly relies on point cloud feature extraction and time series information fusion in static scenes, and has weak ability to capture fast-moving targets in dynamic scenes. It also does not fully consider the interaction between targets in dynamic scenes, which may lead to insufficient accuracy and robustness of the segmentation results. In addition, the method uses a clustering algorithm in the post-processing stage, which may introduce additional computational overhead and affect real-time performance.

[0004] Another patent with the publication number CN118247512B proposes a method for establishing a point cloud semantic segmentation model and a point cloud semantic segmentation method. The method simultaneously trains CNN and Transformer models, combines cross-learning and adaptive class-balanced sampling strategies, and alleviates the problems of uneven spatial distribution of point cloud data and imbalanced data categories, thereby improving segmentation accuracy. However, the technical solution is mainly aimed at optimizing point cloud semantic segmentation in static scenes, and does not fully consider the changes in target motion state and the continuity of spatiotemporal information in dynamic scenes, which may lead to poor segmentation of dynamic targets. In addition, the method relies on partial annotated data to generate pseudo-labels for training, which may cause accumulation of annotation errors in dynamic scenes, further affecting segmentation accuracy.

[0005] The above issues indicate that existing point cloud semantic segmentation methods, when dealing with dynamic scenes, generally suffer from insufficient dynamic target capture, limited segmentation accuracy for multi-category objects in complex scenes, and insufficient integration of spatiotemporal information. Therefore, a point cloud semantic enhancement segmentation method and system for dynamic scenes is urgently needed. By introducing a dynamic scene perception mechanism, this method enhances the ability to capture dynamic targets, optimizes the segmentation accuracy of multi-category objects, and effectively integrates spatiotemporal information, thereby meeting the requirements for efficient and accurate point cloud semantic segmentation in dynamic scenes. Summary of the Invention

[0006] This invention aims to provide a semantically enhanced point cloud segmentation method for dynamic scenes. By introducing a novel dynamic target perception module and spatiotemporal information fusion mechanism, it significantly improves the ability to capture fast-moving targets in dynamic scenes and optimizes the segmentation accuracy of multiple object categories. This method not only effectively processes point cloud data in complex scenes but also achieves high-precision semantic segmentation while ensuring real-time performance, thus meeting the practical needs of fields such as autonomous driving and robotic navigation.

[0007] To address the problems in the prior art, the present invention proposes a method based on an adaptive dynamic feature extraction network. Specifically, the method first constructs a multi-level spatiotemporal feature aggregation module to jointly model point cloud data in both spatial and temporal dimensions. This module utilizes an improved attention mechanism to focus on enhancing the feature expression of dynamic target areas while suppressing the impact of background noise on segmentation results. The attention weight is calculated using the following formula: Among them, x s,i represents the spatial eigenvector of the i-th point, x t,i Represents its time series feature vector, w s and w t are the learnable weight matrices for spatial and temporal features, respectively, and N is the number of points in the point cloud. Through this formula, the model can automatically assign importance weights to different regions, thereby highlighting the key features of dynamic targets.

[0008] Preferably, the present invention further designs a dynamic interaction relationship modeling unit for capturing the interaction relationship between targets in a dynamic scene. The unit introduces a graph convolutional network (GCN) to model dynamic targets in the point cloud and construct a correlation graph between dynamic targets. On this basis, an improved edge weight update strategy is adopted to dynamically adjust the connection strength between targets to better reflect changes in the target motion state. The specific edge weight update formula is as follows: in, represents the edge weight between target i and target j in the kth iteration, f i and f j are the feature vectors of target i and target j respectively, θ ij is the relative motion angle between the two, η is the learning rate, and σ is the activation function. Through this formula, the model can dynamically capture the interaction between objects, thereby improving the robustness of the segmentation results.

[0009] Furthermore, the present invention proposes a lightweight post-processing optimization strategy that avoids the additional computational overhead associated with traditional clustering algorithms. This strategy directly optimizes the segmentation results by designing a fast fusion module based on sparse sampling, ensuring smoother and more accurate output segmentation masks. Experiments have shown that this strategy can further improve segmentation accuracy by approximately 3%-5% without significantly increasing computational cost.

[0010] In summary, the present invention achieves efficient and accurate semantic segmentation of point cloud data in dynamic scenes through the organic combination of a multi-level spatiotemporal feature aggregation module, a dynamic interaction relationship modeling unit, and a lightweight post-processing optimization strategy. Compared with existing technologies, this invention not only significantly improves the ability to capture dynamic targets but also effectively solves the problem of insufficient segmentation accuracy for multi-category objects in complex scenes, demonstrating its high practicality and potential for widespread adoption. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a schematic diagram of the overall process of the method of the present invention, showing the complete processing process from point cloud input to semantic segmentation output.

[0012] Figure 2 This is a structural diagram of the multi-level spatiotemporal feature aggregation module, which details the process of joint modeling of spatial features and temporal features and the application of the attention mechanism.

[0013] Figure 3 This is a working principle diagram of the dynamic interaction relationship modeling unit, showing the construction of the target association graph and the specific implementation of the edge weight update strategy.

[0014] Figure 4 This is the algorithm flow chart corresponding to the improved edge weight update formula, which describes the role of the target eigenvector and relative motion angle in edge weight adjustment.

[0015] Figure 5 This is the module structure diagram of the lightweight post-processing optimization strategy, which reflects the optimization process of the segmentation results by the sparse sampling and fast fusion modules.

[0016] Figure 6 This is the overall block diagram of the system architecture of the present invention, showing the connection relationship between various functional modules and the direction of data flow. DETAILED DESCRIPTION

[0017] The present invention provides a semantically enhanced point cloud segmentation method for dynamic scenes. Its core lies in the organic integration of a multi-level spatiotemporal feature aggregation module, a dynamic interaction relationship modeling unit, and a lightweight post-processing optimization strategy. The following detailed description of the present invention is combined with the accompanying drawings to illustrate its operating principles and process in practical applications.

[0018] First, if Figure 1 As shown, the overall process of the present invention includes complete steps from point cloud data input to final semantic segmentation result output. Point cloud data is usually collected by a lidar or a depth camera, and contains spatial three-dimensional coordinate information and possible attributes such as reflection intensity. These data first enter the multi-level spatiotemporal feature aggregation module for preliminary processing. The core function of this module is to jointly model the point cloud data in the spatial dimension and the temporal dimension, so as to extract a high-dimensional feature vector that can reflect the characteristics of the dynamic target. Specifically, the structure of the multi-level spatiotemporal feature aggregation module is as follows Figure 2 As shown in Figure 1, it mainly consists of a spatial feature extraction submodule and a temporal feature extraction submodule. The spatial feature extraction submodule generates a spatial feature vector for each point by analyzing the spatial distribution characteristics of the point cloud; while the temporal feature extraction submodule uses the correlation between consecutive frames to generate a temporal feature vector for each point. Subsequently, the importance weight of each point is calculated using an improved attention mechanism, as shown in the following formula: Among them, w s and w t are the learnable weight matrices for spatial and temporal features, respectively, and N is the number of points in the point cloud. Through this formula, the model can automatically assign importance weights to different regions, thereby highlighting the key features of dynamic targets and suppressing the influence of background noise.

[0019] Next, the feature vector processed by the multi-level spatiotemporal feature aggregation module is input into the dynamic interaction relationship modeling unit. The design purpose of this unit is to capture the interaction relationship between objects in a dynamic scene. Its working principle is as follows: Figure 3 As shown. The unit first constructs a correlation graph between dynamic targets through a graph convolutional network (GCN), where each point represents a dynamic target and the edge weight represents the connection strength between targets. In order to dynamically adjust the connection strength between targets, the present invention adopts an improved edge weight update strategy, which is formulated as follows: in, represents the edge weight between target i and target j in the kth iteration, fi and f j are the feature vectors of target i and target j respectively, θ ij is the relative motion angle between the two, η is the learning rate, and σ is the activation function. Through this formula, the model can dynamically adjust the edge weights according to the motion state of the target, thereby more accurately reflecting the interaction relationship between targets. Figure 4 The algorithm flow of the improved edge weight update formula is demonstrated, and the role of the target eigenvector and relative motion angle in edge weight adjustment is described.

[0020] After completing the dynamic interaction relationship modeling, the segmentation results enter the lightweight post-processing optimization strategy module for further optimization. The design purpose of this module is to avoid the additional computational overhead brought by traditional clustering algorithms while ensuring that the segmentation mask is smoother and more accurate. Figure 5 As shown in Figure 1, the lightweight post-processing optimization strategy module mainly includes two submodules: sparse sampling and fast fusion. The sparse sampling submodule randomly samples the segmentation results, retaining key points and removing redundant information, thereby reducing the computational complexity of subsequent processing. The fast fusion submodule directly optimizes the segmentation results by designing an efficient fusion algorithm based on the sampled point cloud data. Experiments show that this strategy can further improve segmentation accuracy by approximately 3%-5% without significantly increasing computational cost.

[0021] Finally, the architecture of the entire system is as follows Figure 6 As shown, the connection relationship and data flow direction between the various functional modules are demonstrated. From the point cloud data input to the multi-level spatiotemporal feature aggregation module, to the dynamic interaction relationship modeling unit, the semantic segmentation result is finally output through the lightweight post-processing optimization strategy module. The modules are seamlessly connected through an efficient data transmission mechanism to ensure the real-time and high precision of the entire system. For example, in an autonomous driving scenario, the present invention can be applied to environmental perception tasks around the vehicle. After the lidar collects the point cloud data, the system first extracts the key features of the dynamic target through the multi-level spatiotemporal feature aggregation module, and then uses the dynamic interaction relationship modeling unit to capture the interaction relationship between the targets, and finally generates a smooth and accurate segmentation mask through the lightweight post-processing optimization strategy module. This process can not only effectively identify fast-moving targets, but also achieve accurate segmentation of multiple categories of objects in complex scenes, thereby providing reliable support for autonomous driving decisions.

[0022] In summary, the present invention realizes efficient and accurate semantic segmentation of point cloud data in dynamic scenes through the organic combination of a multi-level spatiotemporal feature aggregation module, a dynamic interaction relationship modeling unit, and a lightweight post-processing optimization strategy. The multi-level spatiotemporal feature aggregation module strengthens the feature expression of dynamic target areas through an improved attention mechanism, while suppressing the influence of background noise; the dynamic interaction relationship modeling unit captures the interaction relationship between targets through a graph convolutional network and an improved edge weight update strategy; the lightweight post-processing optimization strategy further optimizes the segmentation results through sparse sampling and a fast fusion module. The comprehensive application of these technical means enables the present invention to significantly improve the ability to capture dynamic targets and the segmentation accuracy of multiple categories of objects while ensuring real-time performance, and has high practicality and promotion value.

Claims

1. A point cloud semantic enhancement segmentation method for dynamic scenes, characterized by The following steps are involved: The point cloud data is jointly modeled in the spatial and temporal dimensions through a multi-level spatiotemporal feature aggregation module to generate the spatial feature vector and temporal feature vector of each point. The improved attention mechanism is used to calculate the importance weight of each point. The formula is: Among them (F i ) represents the spatial eigenvector of the (i)th point, (T i ) represents its time series feature vector, (W s ) and (W t ) are the learnable weight matrices of spatial and temporal features, respectively, and (N) is the number of points in the point cloud.

2. The method according to claim 1, characterized in that It also includes capturing the interaction relationship between targets in dynamic scenes through a dynamic interaction relationship modeling unit. This unit constructs a correlation graph between dynamic targets through a graph convolutional network and uses an improved edge weight update strategy to adjust the connection strength between targets. The formula is in represents the edge weight between target i and target j in the tth iteration, and are the feature vectors of target i and target j respectively, (θ ij ) is the relative motion angle between the two, (α) is the learning rate, and (σ) is the activation function.

3. The method according to claim 1 or 2, characterized in that It also includes optimizing the segmentation results through a lightweight post-processing optimization strategy, which retains key points and removes redundant information through a sparse sampling submodule, and directly optimizes the segmentation results based on the sampled point cloud data through a fast fusion submodule.

4. A system for implementing the method according to any one of claims 1 to 3, characterized in that It includes a multi-level spatiotemporal feature aggregation module (2) for joint modeling of spatial and temporal dimensions of point cloud data; a dynamic interaction relationship modeling unit (3) for capturing the interaction relationship between objects in dynamic scenes; and a lightweight post-processing optimization strategy module (5) for optimizing the segmentation results.

5. The system according to claim 4, characterized in that The multi-level spatiotemporal feature aggregation module (2) includes a spatial feature extraction submodule and a temporal feature extraction submodule, which are used to generate a spatial feature vector and a temporal feature vector for each point respectively.

6. The system according to claim 4, characterized in that The dynamic interaction relationship modeling unit (3) constructs a correlation graph between dynamic targets through a graph convolutional network, and adjusts the connection strength between targets through an improved edge weight update strategy.

7. The system according to claim 4, characterized in that The lightweight post-processing optimization strategy module (5) includes a sparse sampling submodule and a fast fusion submodule, which are respectively used to retain key points and remove redundant information and directly optimize the segmentation results based on the sampled point cloud data.

8. The system according to claim 4, characterized in that The functional modules are seamlessly connected through an efficient data transmission mechanism, ensuring the real-time performance and high precision of the entire system.

9. The system according to claim 4, characterized in that The system architecture is shown in Figure 6, which shows the connection relationship and data flow direction between various functional modules.

10. The system according to claim 4, characterized in that It is suitable for environmental perception tasks in autonomous driving scenarios. After collecting point cloud data through lidar, it passes through a multi-level spatiotemporal feature aggregation module (2), a dynamic interaction relationship modeling unit (3), and a lightweight post-processing optimization strategy module (5) in sequence, and finally generates a smooth and accurate segmentation mask.

Citation Information

Patent Citations

  • 3D Point Cloud Semantic Segmentation Method

    CN115170585B

  • Point cloud semantic segmentation model establishment method and point cloud semantic segmentation method

    CN118247512B