Three-dimensional scene flow estimation bidirectional learning network based on point aggregation
By using a two-way learning network based on point aggregation in point cloud flow estimation, the technical means of furthest point sampling, point aggregation, bidirectional flow embedding and GRU loop units are used to solve the problem of inaccurate point cloud flow estimation in the existing technology, and more efficient dynamic change capture and time-dependent processing are achieved.
Patent Information
- Application Number
- CN202510249154.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively capture small-scale topographic features and details in point cloud flow estimation, and is susceptible to noise and environmental factors when dealing with scene point clouds, resulting in inaccurate estimation.
A three-dimensional scene flow estimation bidirectional learning network based on point aggregation is adopted to improve the representation of point cloud through the furthest point sampling module and point aggregation module. Combining the bidirectional flow embedding module and the GRU cycle unit, the context information and timing information are fully utilized to improve the accuracy of flow estimation.
It improves the representation of point cloud and the accuracy of flow estimation, can more effectively capture dynamic changes and time dependencies in the scene, reduce noise and discontinuity, and improve the estimation effect of the model.
Smart Images

Figure CN120070492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of autonomous driving and robot navigation, and particularly relates to a bidirectional learning network for three-dimensional scene flow estimation based on point aggregation. Background Art
[0002] Point cloud flow estimation is of great significance in scene understanding and dynamic object tracking. In autonomous driving and robot navigation, it can help the system distinguish and predict static backgrounds and dynamic objects in the environment, and optimize path planning and obstacle avoidance strategies. Point cloud flow estimation not only enhances the understanding of the dynamic world, but also promotes technological innovation and application development in multiple industries.
[0003] Three-dimensional point cloud flow can be estimated by methods such as voxel, multi-modal, point-to-point correspondence, and flow embedding. The voxel-based method converts three-dimensional point cloud data into a voxel grid, thereby simplifying complex spatial information and using techniques similar to traditional two-dimensional image processing for analysis. This method first voxelizes the point cloud, and then extracts information by calculating features of each voxel such as density or surface normal. By comparing the changes in voxel features in consecutive frames, the motion of the object can be estimated. However, this method cannot accurately capture small-scale terrain features and details when processing scene point clouds.
[0004] The multi-modal-based method jointly estimates two-dimensional optical flow and three-dimensional point cloud flow through the method of fusing RGB images and point clouds. Since two-dimensional and three-dimensional motions are highly correlated, that is, two-dimensional motion can be regarded as the projection of three-dimensional motion on the image plane. Therefore, various modal data can be integrated through data synchronization, preprocessing, and the adoption of different fusion strategies (such as early fusion and late fusion). Through this fusion, a more comprehensive understanding and prediction of dynamic changes can be achieved. However, this method is limited by sensor quality and environmental factors when processing scene point clouds, resulting in a decline in the quality of depth and motion data.
[0005] The point-to-point correspondence-based method can directly find the points in the second frame point cloud corresponding to the points in the previous frame point cloud. However, when processing scene point clouds, point-to-point correspondence is easily affected in an environment with a lot of noise, and incorrect matching points may lead to inaccurate flow estimation.
[0006] The method based on flow embedding is used to associate the point clouds of two consecutive frames. In actual data, due to reasons such as changes in perspective and occlusion, there is usually no one-to-one correspondence between the point clouds of two frames. However, the flow can still be estimated by finding points in the second frame that are spatially similar to the points in the first frame. The role of the flow embedding layer is to learn how to aggregate the feature similarities and spatial relationships of these related points, so as to generate an embedding vector encoding the point motion. The flow of the point cloud is obtained by learning the embedding vector through a deep learning network. However, this method is difficult to effectively capture the temporal dependencies and dynamic changes between data when processing scene point clouds.
[0007] By analyzing the existing methods for flow estimation of the above-mentioned scene point clouds and human point clouds, the following problems still exist in the current work: 1) The typical point cloud sampling mode has the problem of over-sampling in dense areas and under-sampling in sparse areas, resulting in the sampled point cloud not being able to well represent the original point cloud.
[0008] 2) The process of flow estimation is difficult to effectively capture the temporal dependencies and dynamic changes between data, resulting in insufficient memory capacity of the model when processing time-dependent data. Summary of the Invention
[0009] The technical problem to be solved by the present invention is how to provide a three-dimensional scene flow estimation bidirectional learning network with stronger time series processing capabilities, which can cope with complex dynamic changes, thereby improving the scene flow estimation effect.
[0010] To solve the above technical problem, the technical solution adopted by the present invention is: a three-dimensional scene flow estimation bidirectional learning network based on point aggregation, the network includes: two farthest point sampling modules, two point aggregation modules, and several groups of bidirectional flow embedding units, each group of bidirectional flow embedding units includes an upsampling module and a bidirectional flow embedding module connected in series; first, the two input frames of point clouds are respectively processed by the farthest point sampling module for farthest point sampling, secondly, the points after the farthest point sampling are sent to the point aggregation module, and the extracted point cloud is re-transformed to generate downsampled points that can better represent the original point cloud and adapt to the dynamic changes in the scene. Thirdly, the point cloud processed by the point aggregation module is input into the bidirectional flow embedding unit for processing. The bidirectional flow embedding unit first processes through the upsampling module and then through the bidirectional flow embedding module, and implements flow estimation by receiving the bidirectional flow embedding features, thereby improving the scene flow estimation effect.
[0011] The beneficial effects of adopting the above technical solutions are as follows: In the point cloud extraction stage of the present application, the point aggregation module is used to re-transform the point cloud, thereby improving the representativeness of the point cloud. In the flow embedding stage, a bidirectional flow embedding method is adopted, which makes full use of the context information in the point cloud sequence task and improves the accuracy of estimation. In the final flow estimation stage, the GRU recurrent unit is used, and its advantages in processing point cloud sequences are utilized to smoothly process the sequence data, reduce noise and discontinuities, thereby improving the estimation effect of the model. Brief Description of the Drawings
[0012] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0013] Figure 1 is the framework diagram of the network according to the embodiment of the present invention; Figure 2 is the framework diagram of the point aggregation module in the network according to the embodiment of the present invention; Figure 3 is the aggregation schematic diagram of the point aggregation module in the network according to the embodiment of the present invention; Figure 4 is the architecture diagram of the bidirectional flow embedding module in the network according to the embodiment of the present invention; Figure 5 is the feature enhancement schematic diagram in the network according to the embodiment of the present invention; Figure 6 is the framework diagram of the GRU recurrent unit in the network according to the embodiment of the present invention. Detailed Embodiments
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0015] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art may make similar extensions without departing from the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0016] As Figure 1 shown, the embodiment of the present invention discloses a bidirectional learning network for three-dimensional scene flow estimation based on point aggregation. The network includes: two farthest point sampling modules, two point aggregation modules, and several groups of bidirectional flow embedding units. Each group of bidirectional flow embedding units includes an upsampling module and a bidirectional flow embedding module connected in series; First, the farthest point sampling module is used to continuously downsample the input two frames of point clouds using PointNet++. Secondly, the points processed by the farthest point sampling are fed into the point aggregation module to re-transform the extracted point cloud to generate downsampled points that can better represent the original point cloud and adapt to the dynamic changes in the scene. Thirdly, the point cloud processed by the point aggregation module is input into the bidirectional flow embedding unit for processing. The bidirectional flow embedding unit is first processed by the upsampling module and then by the bidirectional flow embedding module. The flow estimation is implemented by receiving the bidirectional flow embedding features to improve the effect of scene flow estimation.
[0017] 1) Feature extraction layer First, the input of the feature extraction layer (including the farthest point sampling module and the point aggregation module) is two consecutive frames of point clouds in the autonomous driving scene and , where is the number of points in the point cloud. Then, in the farthest point sampling module, PointNet++ farthest point sampling is used to continuously downsample the point cloud features four times. Among them, the downsampling in the farthest point sampling module can be expressed by the following formula: (1) where represents an operation: downsample points from the point set and extract features using in the attention area of each point. represents the point cloud output by the th downsampling. The number of points output by each layer is N i = [ 2048 , 512 , 256 , 64 ] , and the number of layers used in the attention area of each downsampling layer is with the number of layers being MLP i = [ [ 64 ] , [ 128 ] , [ 256 ] , [ 256 ] ] .
[0018] The final output is and . Among them, the idea of the PointNet++ algorithm is: iteratively select points on and the point cloud . Each time, select the point with the largest distance from each point in the currently selected point set , and add it to the set . The set is used as the finally sampled point cloud. After sampling, the points in the extracted point cloud are aggregated and transformed by the final connection point aggregation module, and the role is to make the features more representative.
[0019] 2) Point aggregation module The framework diagram of the point aggregation module is as shown in Figure 2 The input of the point aggregation module is the point cloud after farthest point sampling and For each point in it, a set of points are clustered, and the point feature vectors of points are calculated using a shared MLP layer , and a pooled feature vector is calculated using a max pooling operator . Then, using the point feature vector and the pooled feature vector , a set of k weights are estimated, and the weighted average of the points in the cluster is used to synthesize a new representative point. The above operations are performed on each point in and , and finally a new set of continuous point cloud coordinates and will be obtained. After that, each new synthesized point in these two sets of point clouds redefines its own region of interest according to its new coordinates through PointNet++. Using MLP = [ 256 ] All points in the feature extraction region of interest are obtained, and the features and are used as the initial input of the point aggregation module for the bidirectional flow embedding module.
[0020] The principle effect diagram of the point aggregation module is as shown in Figure 3 Given a set of point cloud data with the shape of (A2), this module first generates a set of downsampled points through the farthest point sampling algorithm (shown as dark red points in Figure a). Subsequently, as shown in Figure b, the module performs k-nearest neighbor clustering on each downsampled point to obtain its adjacent k points. Then, the point aggregation module fuses these adjacent points to generate stable corresponding points (shown as blue points in Figure c), and these corresponding points will be used for scene flow estimation. Finally, based on the spatial position of the synthesized points, the module uses the PointNet network to reconstruct the point cloud features in the region of interest of the blue points, so as to obtain the feature representation of the synthesized points.
[0021] 3) Bidirectional flow embedding module Different from traditional unidirectional feature propagation, the bidirectional flow embedding layer provides rich context information. The entire model has 4 layers of bidirectional flow embedding modules. The input of the module is the continuous point cloud features after upsampling of layers and scene flow and point cloud transformation correction information , where Figure 4As shown
[0022] 3-1) Bidirectional Feature Enhancement Different from the traditional correlation extraction that only uses the unidirectional features between two consecutive frames, bidirectional feature enhancement enriches the context information by exchanging the information of two point clouds.
[0023] As Figure 5 shown, assuming this is the th layer, the input of bidirectional feature enhancement is the consecutive two-frame point coordinates and corresponding features of the previous layers and . For each point in the two frames, it is embedded into the other point cloud respectively to perform nearest neighbor clustering to aggregate the features of the other frame, and then the features of the point cloud group obtained by clustering are concatenated with the features of the embedded points (from the other frame point cloud), and then the output features and are obtained through shared perception . The bidirectional enhanced features in the current th layer are as follows: S l = SetConv ([ t i − s , g i l − 1 , f l − 1 ]) T l = SetConv ([ s j − t , f j l − 1 , g l − 1 ]) (2) where consists of a shared multi-layer perceptron and a MaxPooling layer, and represent the points of the first and second frame point clouds, and the subscripts , represent the indices of the points in the neighbor group.
[0024] 3-2) Point Cloud Feature Association Module The input of the point cloud feature association module is the enhanced point cloud features and obtained by the bidirectional flow embedding module. The original flow embedding method is used to associate the two-frame point clouds to match the correlation of the points in the two point clouds, so as to obtain the flow embedding features with enhanced features for subsequent scene flow estimation. The formal formula is expressed as: (3) where is the flow embedding enhanced feature, and represent the two-frame point clouds with enhanced features.
[0025] 3-3) GRU Recurrent Unit In scene flow estimation, GRU can effectively capture temporal information. Secondly, GRU can smoothly process the sequential data of point clouds, reducing noise and discontinuities. This is particularly important for scene flow estimation because accurate and continuous flow estimation in point cloud data is required to accurately reflect the motion of objects and scenes.
[0026] Specifically, as Figure 6 shown, the input of the GRU recurrent unit is the corrected information of the output of the previous layer in the previous time step , and the enhanced feature of the flow embedding in the current iteration as the input, and the output is the final corrected information of this layer . The corrected information calculates the scene flow through the mlp layer as the input of the next layer.
[0027] 4) Loss function The network described in this application trains the proposed model in a multi-scale supervised manner. In addition, the estimation results of all intermediate iterations at each scale are also supervised. Specifically, for each layer , the estimated scene flow and the corresponding ground truth of the scene flow The difference between them is quantified by the L2 metric and used as the optimization objective. The final loss function is calculated as follows: (4) where represents the number of layers of the network, is the number of points in the th layer, represents the weight parameter of the iteration results of different network layers, with the default setting , , , .
[0028] The method first performs farthest point sampling on the input two frames of point clouds. Secondly, the sampled points are fed into the point aggregation module. The extracted point clouds are re-transformed to generate downsampled points that can better represent the original point clouds to adapt to the dynamic changes in the scene, which is beneficial to the subsequent matching process of the two frames of point clouds. Thirdly, traditional flow prediction methods do not perform well in dealing with time-dependent features. To address this problem, an improved bidirectional flow embedding module is used in this paper, which introduces bidirectional gated recurrent unit flow estimation to estimate the scene flow. By receiving bidirectional flow embedding features to perform flow estimation, it has stronger time series processing ability to handle complex dynamic changes, thus improving the effect of scene flow estimation. Experimental results show that the performance of this network on the KITTI dataset is better than other existing methods.
Claims
1. A bidirectional learning network for 3D scene flow estimation based on point aggregation, characterized by The network includes: two farthest point sampling modules, two point aggregation modules and several groups of bidirectional stream embedding units, each group of bidirectional stream embedding units includes an upsampling module and a bidirectional stream embedding module connected in series; first, the farthest point sampling modules are used to perform farthest point sampling processing on the two input frames of point clouds respectively; secondly, the points processed by the farthest point sampling are sent to the point aggregation module, and the extracted point clouds are re-transformed to generate down-sampling points that can better represent the original point clouds to adapt to dynamic changes in the scene; thirdly, the point clouds processed by the point aggregation module are input to the bidirectional stream embedding unit for processing, and the bidirectional bidirectional stream embedding unit is first processed by the upsampling module and then processed by the bidirectional stream embedding module, and flow estimation is implemented by receiving the bidirectional stream embedding features to improve the effect of scene flow estimation.
2. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 1, characterized in that: The processing method of the farthest point sampling module comprises the following steps: First, the input of the farthest point sampling module is two consecutive frames of point cloud in the autonomous driving scene. and ,in is the number of point cloud points; Then, in the farthest point sampling module, PointNet++ farthest point sampling is used to downsample the point cloud features four times in a row. The downsampling in the farthest point sampling module is expressed by the following formula: (1); in, Represents an operation, from the point set Downsampling points and use the area of interest at each point Perform feature extraction; Indicates The point cloud output by downsampling is , each downsampling layer focuses on the area using The number of layers is The final output is and .
3. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 1, characterized in that: The processing method of the point aggregation module comprises the following steps: The input of the point aggregation module is the point cloud after sampling the farthest point and Cluster each point to get a set points, calculated using a shared MLP layer The point feature vector of , use the maximum pooling operator to calculate a pooling feature vector ; Then, use the point feature vector And the pooled feature vector Estimate a set of k weights and use these k weights to cluster The weighted average of the points is used to synthesize a new representative point; and Perform the above operation on each point in the , and finally get a new set of continuous point cloud coordinates and ; After that, each new synthetic point of the two sets of point clouds redefines its own focus area according to its new coordinates through PointNet++ Feature extraction focuses on all points in the area and obtains features and The output of the point aggregation module is used as the initial input of the bidirectional stream embedding module.
4. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 1, characterized in that: The bidirectional stream embedding module includes two bidirectional feature enhancement modules, a point cloud feature association module and a GRU cycle unit. The input of the bidirectional stream embedding module is Continuous point cloud features after layer upsampling and , Scene Flow And point cloud transformation correction information ,in Indicates the number of layers, scene flow and point cloud transformation correction information. The initial value is 0.
5. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 4, characterized in that: The processing method of the bidirectional feature enhancement module comprises the following steps: Assume this is the Layer, the input of the bidirectional feature enhancement is The coordinates of two consecutive frames of the layer and the corresponding continuous point cloud features and continuous point cloud features ; For each point in the two frames, embed it into another point cloud and perform Nearest neighbor clustering aggregates the features of another frame, then concatenates the features of the clustered point cloud group with the embedded point features, and finally uses shared perception Get the output enhanced point cloud features and ; In the current Layer bidirectional enhancement features are as follows: (2); in It consists of a shared multilayer perceptron and a MaxPooling layer. and Indicates the points of the first and second frame point clouds, subscript , Represents the index of the point in the neighbor group.
6. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 4, characterized in that: The processing method of the point cloud feature association module comprises the following steps: The input of the point cloud feature association module is the enhanced point cloud features obtained by the bidirectional stream embedding module. and , use the flow embedding method to associate the two frames of point clouds, match the correlation of the points in the two frames of point clouds, and obtain the feature-enhanced flow embedding features for subsequent scene flow estimation. The formal formula is expressed as: (3); in Enhanced features for stream embedding, and Two-frame point cloud showing feature enhancement.
7. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 4, characterized in that: The processing method of the GRU cycle unit comprises the following steps: In scene flow estimation, the GRU cycle unit is used to capture timing information. Secondly, the GRU cycle unit is used to smoothly process the sequence data of the point cloud. The input of the GRU cycle unit is the correction information of the previous layer output. , and the stream embedding enhancement feature of this iteration As input, the output is the final correction information of this layer , correction information Calculate the scene flow through the mlp layer as input to the next layer.
8. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 1, characterized in that: The network model is trained in a multi-scale supervised manner.
9. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network as claimed in claim 8, characterized in that: When the network model is trained in a multi-scale supervised manner, the estimation results of all intermediate iterations at each scale are supervised.
10. The point aggregation-based three-dimensional scene flow estimation bidirectional learning network according to claim 9, characterized in that: All estimated flows are supervised by the ground truth and L2 measure; For each layer , the estimated scene flow The corresponding scene flow true value The difference between them is quantified by the L2 metric and used as the optimization target. The training loss is defined as: (4); in, represents the number of layers in the network, It is The number of layer points, Represents the weight parameters of the iteration results of different network layers, the default setting , , , .