Temporal Fusion Point Cloud 3D Object Detection Method, System, Terminal and Medium

The processing of point cloud data through aligned and consistent timing data enhancement, and the time-sequence characteristics are fused with deformable attention mechanism, which solves the problem of few objects and unbalanced distribution in 3D scenes, significantly improving the performance of 3D object detection.

CN115984637BActive Publication Date: 2025-06-20SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211650983.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2025-06-20
Estimated Expiration
2042-12-21

AI Technical Summary

Technical Problem

The existing 3D object detection algorithms have few objects and unbalanced distribution in 3D scenes, and ignore position movement caused by dynamic object movement in multi-frame timing fusion.

Method used

By obtaining the time-sequential point cloud data, aligning it to the same coordinated coordinate system, and using a time-sequential data enhancement method during the training of the object detection model. Then, the data-enhanced point cloud data is encoded into a bird's-eye view feature map, and the deformable attention mechanism is used to dynamically fuse the features of past moments for the feature map of the current frame.

Benefits of technology

It significantly improves the performance of the single-frame detection algorithm, solves the problem of few objects and unbalanced distribution, and takes into account position movement caused by dynamic object movement in multi-frame timing fusion, providing more reliable 3D object detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984637B_ABST
    Figure CN115984637B_ABST
Patent Text Reader

Abstract

The present invention provides a temporal fusion point cloud 3D object detection method, system, terminal and medium, including: obtaining point cloud data of a time series; aligning the point cloud data to the same coordinate system; during the training process, using temporally consistent data augmentation for training to solve the problem of uneven object distribution; after encoding the point cloud into a bird's-eye view feature map, using a deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame; sending the fused feature map into the detection head to predict objects. The present invention significantly enhances the detection performance, and this method can be applied to any bird's-eye view detection method and can be extended to time series of any length.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and specifically, to a method, system, terminal, and medium for 3D object detection by fusing point clouds in time series. Background Art

[0002] 3D object detection is a key module in the context of autonomous driving and is crucial for subsequent decision-making and path planning. 3D object detection aims to identify objects in 3D space and predict the category and 3D bounding box of the objects. Currently, autonomous driving vehicles are usually equipped with lidar sensors to perceive the surrounding environment, and point cloud data is collected through laser reflection, which has accurate 3D spatial positions. However, point clouds are usually sparse and unevenly distributed, and only a very small number of points can be collected for distant objects and small objects. Nowadays, many algorithms use the point cloud data collected at a certain moment as the output to predict the objects in the surrounding environment. Although these algorithms have achieved good performance, single-frame algorithms ignore the importance of temporal information. In actual situations, due to occlusion and other conditions, it is often difficult for objects to be successfully identified relying on the point cloud data collected at the current moment. For example, at the current moment, a pedestrian in front is blocked by the vehicle in front and is not collected by the lidar. Only relying on the point cloud collected at the current moment, it is impossible to detect that there is a pedestrian in front, which is a major hidden danger to safe driving. However, at a previous moment, the pedestrian appeared completely within the laser collection range, and the algorithm could well identify the pedestrian. Therefore, using temporal information can achieve more reliable detection performance, especially for moving small targets or distant objects, which can provide more reliable guarantees for safe autonomous driving.

[0003] After retrieval, there are also methods for 3D object detection of point clouds by fusing time series in the prior art. For example, the Chinese invention patent with the publication number CN111429514A discloses a method for real-time 3D object detection of lidar by fusing multi-frame time-series point clouds, which can effectively overcome the problem of data sparsity of single-frame point clouds and obtain high accuracy in object detection under severe occlusion and long distances, achieving higher accuracy than single-frame point cloud detection. However, the point cloud completion method in this patent only complements the missing annotations caused by occlusion and other reasons through temporal information, and does not solve the problem of few objects and unbalanced distribution in 3D scenes. Moreover, in the multi-frame temporal fusion, it calculates the feature weights by calculating the pre-similarity at corresponding positions, ignoring the position movement caused by the movement of dynamic objects. Summary of the Invention

[0004] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method, system, terminal, and medium for 3D object detection of point clouds by fusing time series, which significantly enhances the performance of the detection algorithm.

[0005] According to one aspect of the present invention, there is provided a method for 3D object detection of point clouds by fusing time series, including:

[0006] Obtain point cloud data of a time series;

[0007] Align the point cloud data to the same coordinate system;

[0008] During the training process of the target detection model, use a temporally consistent data augmentation method to augment the point cloud data;

[0009] After encoding the augmented point cloud data into a bird's-eye view feature map, use a deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame;

[0010] Send the fused feature map into the detection head to predict objects.

[0011] Optionally, the aligning the point cloud data to the same coordinate system includes:

[0012] Use a parameter matrix to transform the point cloud data of the past moment to the lidar coordinate system of the current frame, so that the target detection model focuses on learning the correlation of objects in temporal movement.

[0013] Optionally, the aligning the point cloud data to the same coordinate system is specifically as follows:

[0014]

[0015] Wherein, is the transformation matrix for converting the point cloud from the lidar coordinate system to the ego-vehicle coordinate system at the (t - 1)th frame moment, is the transformation matrix for converting the point cloud from the ego-vehicle coordinate system to the global coordinate system at the (t - 1)th frame moment; conversely, is the transformation matrix for converting the point cloud from the global coordinate system to the current ego-vehicle coordinate system at the tth moment, is the transformation matrix for converting the point cloud from the ego-vehicle coordinate system to the lidar coordinate system at the tth moment, p t is the point cloud data of the current frame, p t-1 is the point cloud data of the past moment.

[0016] Optionally, the using a temporally consistent data augmentation method to augment the point cloud data means: when training the target detection model, paste additional objects into the current scene, and use the augmented point cloud data as the training dataset.

[0017] Optionally, the using a deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame includes:

[0018] For a certain position q = (x, y) of the feature map F t at the tth moment, its feature is z q, at position l q ,

[0019]

[0020] For position q, generate corresponding sampling offsets Δp for each historical feature map through a linear layer mΔtqk and corresponding weights A mΔtqk , and finally obtain the fused features through weighted sum; K refers to the number of sampling points for each position, M refers to the number of attention heads of the multi-head attention mechanism, F t-Δt (l q +Δp mΔtqk ) refers to taking the features at the corresponding position on the feature map F t-Δt , and W m and W' m are both linear layers.

[0021] According to the second aspect of the present invention, a temporal fusion point cloud 3D object detection system is provided, including:

[0022] Data acquisition module: acquire a temporal sequence of point cloud data;

[0023] Alignment module: align the point cloud data to the same coordinate system;

[0024] Data enhancement module: during the training of the object detection model, use a temporally consistent data enhancement method to enhance the point cloud data;

[0025] Feature fusion module: after encoding the enhanced point cloud data into a bird's-eye view feature map, use the deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame;

[0026] Detection module: send the fused feature map into the detection head to predict objects.

[0027] According to the third aspect of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it is used to execute the temporal fusion point cloud 3D object detection method, or, run the temporal fusion point cloud 3D object detection system.

[0028] According to the fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it is used to execute the temporal fusion point cloud 3D object detection method, or, run the temporal fusion point cloud 3D object detection system.

[0029] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0030] The above detection method of the present invention solves the problem of few objects and unbalanced distribution in the 3D scene, and considers the position movement caused by the movement of dynamic objects in the temporal fusion of multiple frames; by using the information collected at past moments to enhance the bird's-eye view feature map at the current moment, the performance of single-frame detection is significantly improved, and this method can be applied to any bird's-eye view detection method and can be extended to time series of any length.

[0031] The above detection method of the present invention, through the alignment of point cloud data, transforms the point cloud data at past moments to the current ego-vehicle coordinate system through a parameter matrix, aligning the data at different moments, and solves the problem of consistent alignment in the process of processing time-series input point clouds; further performs temporally consistent data augmentation. During the training process, single-frame detectors often use data augmentation methods such as pasting and copying. From the temporal perspective, in order to maintain the consistency of objects, the augmented objects are pasted synchronously in the temporal dimension, significantly improving the detection performance of single-frame point clouds. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0033] Figure 1 It is a flowchart of temporal fusion point cloud 3D object detection in an embodiment of the present invention;

[0034] Figure 2 It is a flowchart of temporal fusion point cloud 3D object detection in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0036] Referring to Figure 1 As shown, it is a flowchart of the method in an embodiment of the present invention, wherein the temporal fusion point cloud 3D object detection method includes:

[0037] S100, Obtain a time series of point cloud data;

[0038] S200, Align the point cloud data to the same coordinate system;

[0039] In this step, through the alignment of point cloud data, the point cloud data at past moments is transformed to the current ego-vehicle coordinate system through a parameter matrix, aligning the data at different moments.

[0040] S300. During the training process of the target detection model, use a temporally consistent data augmentation method to augment the point cloud data;

[0041] In this step, temporally consistent data augmentation is adopted. During the training process, single-frame detectors often use data augmentation methods such as pasting and copying. From the temporal perspective, in order to maintain the consistency of objects, the augmented objects are pasted synchronously in the temporal dimension.

[0042] S400. After encoding the augmented point cloud data into a bird's-eye view feature map, use the deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame;

[0043] In this step, after encoding the point cloud into a bird's-eye view feature map, it is preferable to use the method of deformable attention to dynamically fuse the features of past moments.

[0044] S500. Send the fused feature map into the detection head to predict objects.

[0045] In this step, the feature map can be sent into any detection head for prediction.

[0046] The temporally fused point cloud 3D object detection method provided in the above embodiments of the present invention significantly improves the performance of the single-frame detection algorithm by using the information collected in the past moments to enhance the bird's-eye view feature map of the current moment. It solves the problem of few objects and unbalanced distribution in the 3D scene, and considers the position movement caused by the movement of dynamic objects in the multi-frame temporal fusion.

[0047] Refer to Figure 2 As shown, in a preferred embodiment of the present invention, the temporally fused point cloud 3D object detection method includes the following steps:

[0048] First step, obtain a point cloud data p = {p t-n , …, p t-1 , p t} with a time series of n, where p t is the point cloud data of the current frame, and the remaining n - 1 frames are the point cloud data of past moments.

[0049] Second step, unify the point cloud of this time series to the lidar coordinate system of the current frame.

[0050] As the vehicle moves, the position of the lidar sensor is also constantly moving, and the coordinate origin of the point cloud is also changing. Different coordinate systems are not conducive to the network to learn the relationship between time series. Therefore, use a parameter matrix to transform the point cloud data of past moments to the lidar coordinate system of the current frame, so that the network focuses on learning the correlation of objects during temporal movement. The specific method is (taking the transformation from the p t-1 frame to pt For example:

[0051]

[0052] Among them, is the transformation matrix that transforms the point cloud from the lidar coordinate system to the ego vehicle coordinate system at the t-1 frame moment, is the transformation matrix that transforms the point cloud from the ego vehicle coordinate system to the global coordinate system at the t-1 frame moment; conversely, is the transformation matrix that transforms the point cloud from the global coordinate system to the current ego vehicle coordinate system at the t frame moment, is the transformation that transforms the point cloud from the ego vehicle coordinate system to the lidar coordinate system at the t frame moment, p t is the point cloud data of the current frame, p t-1 is the point cloud data of the past moment

[0053] Step 3: Adopt temporally consistent data augmentation during training.

[0054] Different from images, the 3D spatial range has a much larger scale, but there are far fewer objects in each scene, which greatly limits the convergence speed and final performance of 3D detection networks. To solve this problem, data augmentation methods can be adopted, such as pasting additional objects into the current scene during training.

[0055] In some embodiments, pasting additional objects into the current scene during training can be performed according to the following steps:

[0056] First, generate a database from the training dataset (any labeled dataset), which contains all the manually labeled tags in the training dataset and the points within their manually labeled 3D bounding boxes;

[0057] Then, during the training process of the object detection model, randomly select some manually labeled tags and the points within their manually labeled 3D bounding boxes from this database for each category, and introduce them into the current training point cloud and its manually labeled tags by splicing; using this method can greatly increase the number of labels for each point cloud and simulate objects existing in different environments. At the same time, to avoid physically impossible situations, this method will perform a collision test and delete any sampled objects that collide with other objects.

[0058] The object detection model is a point cloud detection network, and existing detection networks or detection models can be used to implement it.

[0059] Finally, extend the data augmentation of a single frame in the temporal dimension.

[0060] Under the timing setting of this embodiment, the above data augmentation operation will destroy the data consistency. To solve this problem, in this embodiment, the data augmentation of a single frame is further extended in the time dimension, and the specific implementation is described as follows.

[0061] The data augmentation method for a single frame randomly selects a target object O t′ from the point cloud p t′ , and then adds it to the current point cloud p t . Under the timing setting of the embodiment of the present invention, the training scene sequence is {p t-Δt , Δt = 0, 1, 2…n}, correspondingly, an object sequence {O t′-Δt} also needs to be selected from {p t′-Δt}. However, directly adding this object sequence to the training scene will cause great noise interference because the relative motion between the object sequences is inconsistent with the relative motion in the training scene, making it impossible for the network to learn well. Therefore, it is also necessary to transform this object sequence to the current training scene sequence:

[0062] O′ t′-Δt = T t→(t-Δt) × T (t′-Δt)→t′ × O t′-Δt

[0063] In the above formula, T (t′-Δt)→t′ is to transform the pasted object from the t′ - Δt moment in the source point cloud sequence to the t′ moment, and T t→(t-Δt) means to transform the pasted object from the t moment of the current training point cloud sequence to the t - Δt moment, and O′ t′-Δt refers to the object finally pasted into the training point cloud sequence.

[0064] In this embodiment, first use T (t′-Δt)→t′ to transform the historical objects in the object sequence {O t′-Δt} to the t′ moment, and then use T t→(t-Δt) to transform these historical objects to the historical frames corresponding to the current training scene. Using the above design, the relative motion in the training scene is maintained.

[0065] In this step, in order to achieve data augmentation with consistent timing during training, when pasting additional objects for the current scene during training, the data augmentation of a single frame is extended in the time dimension, and the pasting of additional objects in a single frame is extended to the pasting of objects in a time series, maintaining data consistency and maintaining the relative motion of objects in the time series.

[0066] Fourth step, encode the point cloud into a bird's-eye view feature map {F t-Δt = B N×C×X×Y}.

[0067] This process can use any existing point cloud encoding method.

[0068] Step 5: Dynamically fuse the features of historical frames for the bird's-eye view feature map of the current frame.

[0069] Transformer can use the attention mechanism to adaptively fuse features, but it will cause a large amount of computation and is not suitable for large-sized feature maps. Therefore, in this embodiment, a deformable attention mechanism is adopted for temporal feature fusion. Specifically, for the feature map F at time t t at a certain position q=(x, y), its feature is z q , and the position is l q ,

[0070]

[0071] For position q, a corresponding sampling offset Δp mΔtqk and a corresponding weight A mΔtqk are generated for each historical feature map through a linear layer. Finally, the fused feature is obtained through weighted summation; K refers to the number of sampling points for each position, M refers to the number of attention heads of the multi-head attention mechanism, and F t-Δt (l q +Δp mΔtqk ) refers to taking the feature at the corresponding position on the feature map F t-Δt , and W m and W' m are both linear layers.

[0072] Step 6: Send the fused bird's-eye view feature map into the detection head to obtain the final detection result.

[0073] In this embodiment, various objects in the current training scenario are increased through a temporally consistent data augmentation scheme, and the relative motion relationship between time series is preserved, which is very beneficial to the training of the model. And the deformable attention mechanism used generates motion offsets dynamically for each position and adaptively obtains corresponding features on the temporal feature map, which is more suitable for the temporal fusion of dynamic and static objects.

[0074] Most of the existing point cloud detection algorithms focus on the input of single-frame data and rarely involve the fusion of temporal information. In the above embodiment of the present invention, a point cloud temporal fusion method is proposed. Through the method of the deformable attention mechanism, features of past moments are dynamically extracted for the current bird's-eye view feature map, and it is easy to extend to longer time series. The introduction of temporal information improves the detection performance of the algorithm for occluded objects and small moving objects, which is crucial for safe driving.

[0075] Based on the same technical concept, in another embodiment of the present invention, a point cloud 3D object detection system for temporal fusion is further provided, including:

[0076] Data acquisition module: Acquire a time series of point cloud data;

[0077] Alignment module: Align the point cloud data to the same coordinate system;

[0078] Data augmentation module: During training, use temporally consistent data augmentation for training to address the uneven distribution of objects;

[0079] Feature fusion module: After encoding the point cloud into a bird's-eye view feature map, use a deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame;

[0080] Detection module: Feed the fused feature map into the detection head to predict objects.

[0081] In another embodiment of the present invention, a terminal is further provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it is used to execute the above-mentioned temporal fusion point cloud 3D object detection method, or, run the above-mentioned temporal fusion point cloud 3D object detection system.

[0082] In another embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored. When the program is executed by a processor, it is used to execute the temporal fusion point cloud 3D object detection method in any of the above embodiments, or, run the temporal fusion point cloud 3D object detection system in any of the above embodiments.

[0083] To better understand the above embodiments of the present invention, the following is described in conjunction with a specific application:

[0084] Taking the PillarNet point cloud detector as an example, PillarNet is a detection algorithm that divides the point cloud into pillar representations and achieves excellent detection performance at real-time speed. Figure 1 It is the overall flowchart of the method in an embodiment of the present invention.

[0085] Specifically, the detection method in this embodiment includes the following steps:

[0086] First step, acquire a point cloud data p = {p t-n , …, p t-1 , p t} with a time series of n, where p t is the point cloud data of the current frame, and the remaining n - 1 frames are the point cloud data of past moments.

[0087] Second step, unify the point cloud of this time series to the lidar coordinate system of the current frame.

[0088] Step 3: During training, use temporally consistent data augmentation.

[0089] Step 4: Encode the point cloud into a bird's-eye view feature map {F t-Δt = B N×C×X×Y}.

[0090] In this embodiment, PillarNet first divides the 3D space into equal-sized pillars (with infinite height along the z-axis) on the x, y plane according to a set size, then calculates the relationship between the points and the pillars based on the coordinates of each point, and then uses a network similar to PointNet to encode the points inside the pillars into feature vectors of equal length. However, due to the sparsity of the point cloud, not all positions have non-empty pillars, so the finally encoded feature vectors are sparse. After that, to further extract features, sparse 2D convolution is used to process the encoded vectors to obtain the bird's-eye view feature map.

[0091] Step 5: Use the deformable attention mechanism to dynamically fuse the features of historical frames for the bird's-eye view feature map of the current frame.

[0092] Step 6: Use a feature pyramid network to extract multi-scale features of the fused features to facilitate the detection of objects of different sizes. Then, the feature map is sent into the detection head to obtain the corresponding detection results.

[0093] This embodiment uses a detection head that does not require predefined anchor boxes. This detection head directly predicts possible center point offsets, lengths, widths, and other object attributes for each position, and finally obtains the final detection results by using non-maximum suppression.

[0094] Implementation effect:

[0095] According to the above steps, corresponding tests were carried out on the commonly used autonomous driving dataset nuScenes. The official evaluation metrics used mAP and NDS to evaluate the performance. mAP is the average accuracy of detections for each category, based on the weighted sum of the distances from the center of the bird's-eye view. NDS is a custom metric that combines attributes such as the size, rotation, and speed of the detection box. Table 1 shows the test results of the temporal fusion method (pillarnet_temporal) of this invention and the original PillarNet (single-frame detector) on nuScenes. From various evaluation metrics, the temporal fusion method proposed in this embodiment of the invention has achieved significant improvement compared to single-frame input.

[0096] Table 1

[0097] Method mAP NDS pillarnet 60.95 67.77 pillarnet_temporal 62.84 69.29 pillarnet_fade15 62.45 68.66 pillarnet_temporal_fade15 64.08 69.76

[0098] Note: Here, "fade" means canceling the data augmentation strategy in the last five rounds of training. The data augmentation method is beneficial and can improve the performance of the model on almost all categories. Using "fade" in the last few rounds will achieve better results because during the augmentation process, the pasting positions of the data are random. For example, a car may be placed inside a building, resulting in incorrect data distribution. Since the model learns the data distribution during the learning process, this incorrect object distribution is also learned by the model, leading to incorrect detection results. Therefore, canceling the data augmentation strategy in the last few rounds enables the model to learn the true scene distribution and further improve the model's performance.

[0099] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the system to implement the step flow of the method. That is, the embodiments in the system can be understood as preferred examples for implementing the method, and will not be elaborated here.

[0100] Those skilled in the art know that in addition to implementing the system and its various devices provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices provided by the present invention can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices for implementing various functions can also be regarded as either software modules for implementing the method or the structure within the hardware component.

[0101] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific implementation manners. Those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be combined arbitrarily without conflict.

Claims

1. A method for 3D object detection of point cloud with temporal fusion, characterized in that, Including: Obtain point cloud data of a time series; Align the point cloud data to the same coordinate system; During the training process of the object detection model, use a temporally consistent data augmentation method to augment the point cloud data; After encoding the augmented point cloud data into a bird's-eye view feature map, use a deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame; Send the fused feature map into the detection head to predict objects; During the training process of the object detection model, using a temporally consistent data augmentation method to augment the point cloud data means: when training the object detection model, paste additional objects into the current scene, and use the augmented point cloud data as the training dataset; Pasting additional objects into the current scene during training includes: First, generate a database from the training dataset, which contains all manually annotated labels and the points within their manually annotated 3D bounding boxes; Then, during the training process of the object detection model, randomly select some manually annotated labels and the points within their manually annotated 3D bounding boxes from this database for each category, and introduce them into the current training point cloud and its manually annotated labels by splicing; Finally, extend the data augmentation of a single frame in the temporal dimension; The extension of the data augmentation of a single frame in the temporal dimension is specifically: Under the setting of time sequence, the training scenario sequence is {p t-Δt , Δt = 0, 1, 2…n}, and an object sequence {O t-Δt} is selected from {p t′-Δt}; among them, convert this object sequence to the current training scenario sequence: O′ t′-Δt = T t→(t-Δt) × T (t′-Δt)→t′ × O t′-Δt In the above formula, T (t′-Δt)→t′ is to transform the pasted object from the t′ - Δt moment in the source point cloud sequence to the t′ moment, and T t→(t-Δt) refers to transforming the pasted object from the t moment in the current training point cloud sequence to the t - Δt moment, and O′ t′-Δt refers to the object finally pasted into the training point cloud sequence; first use T (t′-Δt)→t′ to transform the historical objects in the object sequence {O t′-Δt} to the t′ moment, and then use T t→(t-Δt) to transform these historical objects into the historical frames corresponding to the current training scene.

2. The method for 3D object detection of point cloud with temporal fusion according to claim 1, characterized in that, Aligning the point cloud data to the same coordinate system includes: Use a parameter matrix to transform the point cloud data of past moments to the lidar coordinate system of the current frame, so that the object detection model focuses on learning the correlation of objects during temporal movement.

3. The method for 3D object detection of point cloud with temporal fusion according to claim 1, characterized in that, Specifically, aligning the point cloud data to the same coordinate system is as follows: Among them, is the transformation matrix that transforms the point cloud from the lidar coordinate system to the ego-vehicle coordinate system at the t-1 frame moment, is the transformation matrix that transforms the point cloud from the ego-vehicle coordinate system to the global coordinate system at the t-1 frame moment; conversely, is the transformation matrix that transforms the point cloud from the global coordinate system to the current ego-vehicle coordinate system at the t moment, is the transformation that transforms the point cloud from the ego-vehicle coordinate system to the lidar coordinate system at the t moment, p t is the point cloud data of the current frame, p t-1 is the point cloud data of the past moment.

4. The method for 3D object detection of point cloud with temporal fusion according to claim 1, characterized in that, Using the deformable attention mechanism to dynamically fuse the features of past moments for the feature map of the current frame includes: For the feature map F at time t t At a certain position q=(x,y), its feature is z q , the position is l q , For position q, generate the corresponding sampling offset Δp for each historical feature map through a linear layer mΔtqk and the corresponding weight A mΔtqk , and finally obtain the fused feature through weighted sum; K refers to the number of sampling points for each position, M refers to the number of attention heads in the multi-head attention mechanism, F t-Δt (l q +Δp mΔtqk ) refers to taking the feature at the corresponding position on the feature map F t-Δt , and W m and W′ m are both linear layers.

5. A system for 3D object detection of point cloud with temporal fusion, characterized in that, Including: Data acquisition module: Obtain point cloud data of a time series; Alignment module: Align the point cloud data to the same coordinate system; Data augmentation module: During the training process of the object detection model, use a temporally consistent data augmentation method to augment the point cloud data; Feature fusion module: After encoding the augmented point cloud data into a bird's-eye view feature map, use a deformable attention mechanism to dynamically fuse the feature maps of past moments for the feature map of the current frame; Detection module: Send the fused feature map into the detection head to predict objects; The data augmentation module, during the training process of the object detection model, using a temporally consistent data augmentation method to augment the point cloud data means: when training the object detection model, paste additional objects into the current scene, and use the augmented point cloud data as the training dataset; Pasting additional objects into the current scene during training includes: First, generate a database from the training dataset, which contains all manually annotated labels and the points within their manually annotated 3D bounding boxes; Then, during the training process of the object detection model, randomly select some manually annotated labels and the points within their manually annotated 3D bounding boxes from this database for each category, and introduce them into the current training point cloud and its manually annotated labels by splicing; Finally, expand the data augmentation of a single frame in the temporal dimension; The data augmentation of expanding a single frame in the temporal dimension is specifically as follows: Under the setting of time series, the training scenario sequence is {p t-Δt , Δt = 0, 1, 2…n}, and an object sequence {O t-Δt} is selected from {p t′-Δt}; among them, convert this object sequence to the current training scenario sequence: O′ t′-Δt = T t→(t-Δt) × T (t′-Δt)→t′ × O t′-Δt In the above formula, T (t′-Δt)→t′ is to convert the pasted object from the t′-Δt moment in the source point cloud sequence to the t′ moment, and T t→(t-Δt) refers to converting the pasted object from the t moment in the current training point cloud sequence to the t-Δt moment, and O′ t′-Δt refers to the object finally pasted into the training point cloud sequence; first use T (t′-Δt)→t′ to convert the historical objects in the object sequence {O t′-Δt} to the t′ moment, and then use T t→(t-Δt) to convert these historical objects to the historical frames corresponding to the current training scene.

6. A terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it can be used to execute the method described in any one of claims 1-4, or to run the system described in claim 5.

7. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it can be used to execute the method described in any one of claims 1-4, or to run the system described in claim 5.

Citation Information

Patent Citations

  • Laser radar 3D real-time target detection method fusing multi-frame time sequence point cloud

    CN111429514A

  • Three-dimensional space model reconstruction method and device and storage medium

    CN113570721A