Sparse time fusion framework and method for detecting three-dimensional object based on point cloud, electronic equipment and storage medium
By adopting a sparse time fusion framework in the processing of lidar point cloud data, the features of current and historical frames are extracted and fused, the problem of insufficient information in long-distance detection of point cloud data is solved, and higher target detection accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510142251.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-20
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-10
AI Technical Summary
The point cloud data acquired by lidar contains insufficient effective information in long-distance detection, which makes it difficult to accurately identify and locate targets in complex occlusion scenarios and when detecting long-distance targets.
A sparse time fusion framework based on point cloud detection of three-dimensional objects is adopted. Through the feature extraction module, multi-historical feature alignment module, sparse feature extraction module and timing fusion module, feature extraction, alignment, sparse representation and timing fusion module of the current frame and historical frame are used to generate high-value sparse representations to improve the performance and accuracy of object detection.
Through the sparse time fusion framework, the response of non-target areas is effectively suppressed, potential target locations are highlighted, the performance and accuracy of three-dimensional object detection are improved, and the understanding and processing capabilities of dynamic scenarios are enhanced.
Smart Images

Figure CN120126089A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular, to a sparse temporal fusion framework and method for detecting three-dimensional objects based on point clouds, an electronic device, and a storage medium. Background Art
[0002] LiDAR (Light Detection and Ranging) technology has been widely used in the field of three-dimensional object detection because it can generate high-precision three-dimensional point cloud data by scanning the surrounding environment with laser beams. Especially in the application of vehicle autonomous driving, three-dimensional object detection based on LiDAR has become a core technology in the autonomous driving perception system, which relies on high-precision data from LiDAR. This enables the vehicle to accurately and comprehensively perceive the surrounding environment, thus ultimately ensuring driving safety. However, the point cloud data obtained by LiDAR is inherently sparse and distributed over a vast space. This characteristic significantly increases the difficulty of accurately identifying and locating objects in complex occlusion scenarios and when detecting distant objects. In addition, single-frame point cloud data can only provide a partial view of the scene and cannot fully capture the environmental context and dynamic changes. This information loss is mainly caused by object occlusion, occlusion by other objects, and the limitation of the sensor's perspective. Therefore, a technical solution is needed to increase the effective information contained in the point cloud data to improve the accuracy of LiDAR in long-distance detection. Summary of the Invention
[0003] Embodiments of this application provide a sparse temporal fusion framework and method for detecting three-dimensional objects based on point clouds, an electronic device, and a storage medium, to solve the defect that the effective information contained in point cloud data is insufficient in long-distance detection in the prior art.
[0004] To achieve the above object, embodiments of this application provide a sparse temporal fusion framework for detecting three-dimensional objects based on point clouds, and the framework includes:
[0005] A feature extraction module, configured to extract features from the point cloud data of the current frame and historical frames falling into each three-dimensional voxel in a pre-divided three-dimensional voxel network, so as to generate a current frame feature and a historical frame feature;
[0006] A multi-historical feature alignment module, which performs a series of affine transformations on the historical frame features by using the vehicle chassis information of the target vehicle, and converts them into the coordinate system of the current moment;
[0007] A sparse feature extraction module, which is used to extract bird's-eye view features from the current frame features and historical frame features, extract a dense heat map from the bird's-eye view features, only retain the maximum value in the local area of the heat map to generate the current frame sparse representation and the historical frame sparse representation, sort the feature vectors in the current frame sparse representation and the historical frame sparse representation, and select the first predetermined number of features to create a high-value sparse representation;
[0008] A temporal fusion module, which is used to extract the current sparse features and historical sparse features from the current frame high-value sparse representation and the historical frame sparse representation respectively, perform projection processing on the current sparse features and the historical sparse features to obtain the projected current features and the projected historical features, use the projected current features to extract relevant information from the projected historical features, and extract relevant information from the bird's-eye view features based on the current sparse features, the historical sparse features, the position embedding vector for the sparse features, and the position embedding vector for the bird's-eye view features to generate the current features.
[0009] An embodiment of the present application further provides a sparse temporal fusion method for detecting three-dimensional objects based on point clouds. The method includes:
[0010] Feature extraction is performed on the point cloud data of the current frame and the historical frame falling into each three-dimensional voxel in a pre-divided three-dimensional voxel network to generate the current frame features and the historical frame features;
[0011] A series of affine transformations are performed on the historical frame features using the vehicle chassis information of the target vehicle to transform them into the coordinate system of the current moment;
[0012] Bird's-eye view features are extracted from the current frame features and the historical frame features, a dense heat map is extracted from the bird's-eye view features, only the maximum value is retained in the local area of the heat map to generate the current frame sparse representation and the historical frame sparse representation, the feature vectors in the current frame sparse representation and the historical frame sparse representation are sorted, and the first predetermined number of features are selected to create a high-value sparse representation;
[0013] The current sparse features and the historical sparse features are extracted from the current frame high-value sparse representation and the historical frame sparse representation respectively, the current sparse features and the historical sparse features are subjected to projection processing to obtain the projected current features and the projected historical features, the projected current features are used to extract relevant information from the projected historical features, and relevant information is extracted from the bird's-eye view features based on the current sparse features, the historical sparse features, the position embedding vector for the sparse features, and the position embedding vector for the bird's-eye view features to generate the current features.
[0014] An embodiment of the present application further provides an electronic device, including:
[0015] A memory, which is used to store programs;
[0016] A processor for running the program stored in the memory, and when the program runs, it executes the sparse temporal fusion method for detecting three-dimensional objects based on point cloud provided in the embodiments of the present application.
[0017] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program executable by a processor is stored. When the program is executed by the processor, it implements the sparse temporal fusion method for detecting three-dimensional objects based on point cloud provided in the embodiments of the present application.
[0018] The sparse temporal fusion framework and method for detecting three-dimensional objects based on point cloud, electronic device and storage medium provided in the embodiments of the present application adopt a temporal partitioning strategy, divide the features extracted by the feature extraction module into current frame features and historical frame features, and the sparse feature extraction module only retains the maximum response value in the local area of each position in the generated dense heat map and aggregates it to form a new sparse representation, effectively suppressing the response of non-target areas, highlighting potential target positions and improving the performance and accuracy in three-dimensional object detection. In the temporal fusion module, the current sparse feature and historical sparse feature are used as queries to extract relevant information from the original bird's-eye view features, enabling the network to effectively capture the current features, enhancing the understanding and processing ability for dynamic scenes, and realizing efficient information fusion through the interaction between the current frame and the historical frame.
[0019] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically illustrates the specific embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0021] Figure 1 It is a schematic structural diagram of the sparse temporal fusion framework for detecting three-dimensional objects based on point cloud provided by the present application;
[0022] Figure 2 It is a schematic diagram of the direction and position information of vehicles in the bird's-eye view used in the embodiments of the present application;
[0023] Figure 3 It is a schematic diagram of the alignment of historical frame features in the embodiments of the present application;
[0024] Figure 4Schematic diagram of the sparse feature extraction module of the sparse temporal fusion framework for three-dimensional object detection based on point cloud provided by an embodiment of the present application;
[0025] Figure 5 Schematic diagram of the temporal fusion module of the sparse temporal fusion framework for three-dimensional object detection based on point cloud provided by an embodiment of the present application;
[0026] Figure 6 Flowchart of an embodiment of the sparse temporal fusion method for three-dimensional object detection based on point cloud provided by the present application;
[0027] Figure 7 Schematic diagram of the structure of an embodiment of the electronic device provided by the present application. Detailed implementation manners
[0028] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0029] Embodiment 1
[0030] Due to its ability to generate high-precision three-dimensional point cloud data by scanning laser beams around the environment, Light Detection and Ranging (LiDAR) technology has been widely used in the field of three-dimensional object detection. Especially in the application of vehicle autonomous driving, three-dimensional object detection based on LiDAR has become a core technology in the autonomous driving perception system, which relies on high-precision data from LiDAR. This enables the vehicle to accurately and comprehensively perceive the surrounding environment, thus ultimately ensuring driving safety. However, the point cloud data obtained by LiDAR is inherently sparse and distributed over a vast spatial range. This characteristic significantly increases the difficulty of accurately identifying and locating targets in complex occlusion scenarios and when detecting distant targets. In addition, single-frame point cloud data can only provide a partial view of the scene and cannot fully capture the environmental context and dynamic changes. This information loss is mainly caused by target occlusion, occlusion by other objects, and limitations of the sensor's perspective.
[0031] In the prior art, it has been proposed to use multi-frame lidar data as input, expand, aggregate and distinguish the input data, thereby improving the detection performance of the network. The direct temporal fusion method proposed in the prior art utilizes the point cloud sequences captured when the vehicle moves in the real scene. Compared with single-frame information, these multi-frame sequences provide more comprehensive information and represent the surrounding environment in a denser manner. Therefore, they exhibit better performance than single-frame lidar networks. However, when additional frames are simply stacked without precisely modeling the inter-frame relationship, the bottleneck of performance improvement quickly emerges. In addition, when each frame is stacked with different adjacent frames, it must be processed multiple times, which significantly increases the complexity of the work. It should be noted that in the embodiments of the present invention, taking lidar as the input device for generating point clouds as an example, in fact, the framework and method of the embodiments of the present invention can also be applied to scenarios based on other input devices that can directly or indirectly generate point cloud data. For example, using a depth camera as a visual sensor to form three-dimensional point cloud data by extracting feature points in the image.
[0032] The prior art has also proposed to use a Region Proposal Network (RPN) to obtain multiple 3D proposals, and then extract foreground object features through 3D ROI processing to achieve temporal fusion. This prior art utilizes the learned historical embeddings to iteratively fuse the latent embeddings between consecutive frames in deeper layers of the model.
[0033] The above-mentioned solutions in the prior art avoid the rough fusion of input frames and avoid the repeated processing of data. However, the effectiveness of temporal fusion is limited by the quality of the generated 3D proposals.
[0034] For this reason, this application proposes a structure of a sparse temporal fusion framework for detecting three-dimensional objects based on point clouds as shown in Figure 1 This framework is designed specifically for temporal fusion in object detection algorithms. It adopts a sparse fusion method and focuses on merging the key features containing objects from consecutive point cloud sequences. The framework solution proposed in the embodiments of this application enhances the importance of feature extraction and improves the network's ability to capture key information. The core idea of this framework is to guide the network to focus on the focal information in each frame of the point cloud. By integrating the foreground point information of the previous and subsequent frames, the network's sensitivity to the change of the target position is improved. This reduces the size of the fused features and enhances the fusion efficiency. Therefore, this method can more effectively process the feature fusion of multi-frame temporal point clouds and make the network pay more attention to the foreground features and the change of the target position, thereby improving the robustness and accuracy in complex scenes.
[0035] Specifically, as shown in Figure 1As shown in the figure, the sparse temporal fusion framework for detecting three-dimensional objects based on point cloud provided by this application may include: a feature extraction module 11, a multi-historical feature alignment module 12, a sparse feature extraction module 13, and a temporal fusion module 14.
[0036] The feature extraction module 11 can be used to extract features from the point cloud data of the current frame and historical frames falling into each three-dimensional voxel in a pre-divided three-dimensional voxel network, so as to generate current frame features and historical frame features.
[0037] The multi-historical feature alignment module 12 can perform a series of affine transformations on the historical frame features by using the vehicle chassis information of the target vehicle, and transform them into the coordinate system of the current moment;
[0038] The sparse feature extraction module 13 can be used to extract bird's-eye view features from the current frame features and historical frame features, extract a dense heat map from the bird's-eye view features, only retain the maximum value in the local area of the heat map to generate a current frame sparse representation and a historical frame sparse representation, sort the feature vectors in the current frame sparse representation and the historical frame sparse representation and select the first predetermined number of features to create a high-value sparse representation;
[0039] The temporal fusion module 14 can be used to extract current sparse features and historical sparse features from the current frame high-value sparse representation and the historical frame sparse representation respectively, perform projection processing on the current sparse features and the historical sparse features to obtain projected current features and projected historical features, use the projected current features to extract relevant information from the projected historical features, and extract relevant information from the bird's-eye view features based on the current sparse features, the historical sparse features, the position embedding vector for sparse features, and the position embedding vector of the bird's-eye view features, so as to generate current features.
[0040] In the embodiments of this application, the feature extraction module 11 can use a regular grid structure (such as voxels or cylinders) to organize irregular point clouds, and adopt a 2D or 3D convolutional neural network (CNN) to extract dense features. Specifically, both the cylinder and voxel methods discretize 3D point cloud data and achieve object detection by extracting features from specific regions in the grid. However, the voxel method can usually learn richer feature representations because they can capture more spatial relationships in the point cloud data. In the embodiments of this application, the feature extraction module 11 can use the voxel method for feature extraction, and may include a voxel feature encoding layer for extracting features from the point cloud data falling into each voxel to generate the current frame features and historical frame features.
[0041] Assume the current moment is t 0 , and the historical time series can be correspondingly expressed as {t -1 , t -2 , t-3 , …, t -m}. The multi-historical feature alignment module 12 can apply pose transformation to the historical frame features extracted by the feature extraction module 11 based on the position and orientation of the vehicle. First, the relative position (x, y) of the target vehicle and the ego-vehicle direction angle can be utilized in the bird's-eye view ego to calculate the vehicle's direction angle angle based on the following formulas (1)-(3) bev and displacement T bev :
[0042]
[0043]
[0044] angle bev = angle ego - angle tra (3)
[0045] As Figure 2 shown in Figure 2 is a schematic diagram of the direction and position information of the vehicle in the bird's-eye view used in the embodiments of this application. In the above formula, T bev is the displacement amount of the vehicle in the bird's-eye view, angle tra is the moving angle, and angle bev is the direction angle of the vehicle in the bird's-eye view.
[0046] Then, the multi-historical feature alignment module 12 can use the angle, displacement, and grid coordinates (grid 1 , grid h ) of the vehicle in the bird's-eye view to calculate the offsets shift x and shift y of the historical frame features in the horizontal and vertical directions in the grid coordinates through the following formulas (4) and (5):
[0047]
[0048] Based on the offsets shift x and shift y of the historical frame features in the bird's-eye view and the rotation angle angle rot of the vehicle, an affine transformation is performed on the historical frame features based on the following formula (6) to obtain the transformed historical frame features pre-feat ∈ R (B,C,h,w) :
[0049]
[0050] In the above formula, feat ∈ R (B,C,h,w)denote the original features in the input shared network, B denotes the number of samples in one input, C denotes the number of channels of the input features, h denotes the height of the input features, and w denotes the width of the input features. Then, the transformed historical frame features are combined with the features at the last time point in the historical time series {t -1 ,t -2 ,t -3 ,…,t -m} to form a fused feature representation, as shown in formula (7) below.
[0051]
[0052] In formula (7), denotes the concatenation operation.
[0053] In addition, a schematic diagram of the alignment of the historical frame features obtained according to the above alignment process is shown in Figure 3 .
[0054] In existing object detection tasks based on point cloud data, generating dense feature representations is considered to improve object detection performance. However, there are many problems in practical applications, such as noise generation and low efficiency. Specifically, generating dense features requires dense prediction of the entire point cloud, which will generate a large amount of background noise during the processing. These noises may interfere with the accurate recognition of the model, especially when dealing with complex scenes. In the embodiments of the present application, a time data partitioning strategy is adopted to divide the input data into current frame information and historical frame information. In order to fully utilize the potential information in these data, the sparse feature extraction module 13 performs sparse feature extraction on each part separately.
[0055] As shown in Figure 4 , Figure 4 is a schematic structural diagram of the sparse feature extraction module of the sparse time fusion framework for detecting three-dimensional objects based on point clouds provided by the embodiments of the present application. In the embodiments of the present application, a shared network composed of, for example, 3×3 convolutional layers can be constructed. This network is used to learn BEV (bird's-eye view) features from the current frame and historical frames and extract relevant information. The original feature input to the shared network is denoted as feat∈R (B,C,h,w) .
[0056] feat bev =Shared-Conv(feat) (8)
[0057] The sparse feature extraction module 13 can output a series of predicted heatmaps using the BEV feature maps from the current frame and historical frames. Each predicted heatmap corresponds to a specific target category. The peak in the heatmap represents the center position of the targets of each category in the BEV feature map. In the embodiment of this application, in order to generate the ground-truth heatmap, the Gaussian distribution mechanism is used to "draw" the center of each target into a region, creating a two-dimensional ground-truth heatmap with a size of h×w. Since each target may belong to different categories, a ground-truth heatmap needs to be generated for each category. These heatmaps are stacked according to the number of categories, forming a three-dimensional ground-truth heatmap, where k represents the number of different categories. During this process, the center position of each target affects the heatmap according to the Gaussian function, which can be expressed by the following formula (9):
[0058]
[0059] where k i represents the category of the i-th target, and σ i is the standard deviation related to the target scale.
[0060] The ground-truth heatmap generated by the Gaussian distribution function can be used as the supervision information during the training process in the embodiment of this application, guiding the predicted heatmap to gradually align with the ground-truth heatmap, so as to generate a dense heatmap DH∈R (B,K,h×w) .
[0061] In order to reduce the interference of irrelevant features in the spatial domain, in the embodiment of this application, only the maximum response value within the local region of each position in the generated dense heatmap can be retained. Thus, the sparse feature extraction module 13 can extract these local maximum elements and aggregate them based on the following formula (10) to form a new sparse representation map SH∈R (B,K,h×w) . This sparse representation effectively suppresses the response of non-target regions, thereby highlighting potential target positions and improving the performance and accuracy of subsequent detection tasks:
[0062] SH(x,y,k i )= max(DH(x,y,k i )) (10)
[0063] To extract the most likely target candidate regions from the sparse representation and assign class information to these regions, the sparse feature extraction module 13 can first sort the elements in the sparse representation and extract the top N elements to create a high-value sparse representation. Then, each class in the high-value sparse representation is converted into a one-hot vector (One-Hot Encoding). These one-hot vectors can be further converted into learnable class embedding vectors based on the following formula (11). These embeddings capture the semantic information of the classes and provide valuable object prior knowledge during the prediction process, enabling the network to focus more on the changes in the internal relationships of the classes.
[0064]
[0065] In the above formula (11), represents the learnable class embedding vector, represents each class in the high-value sparse representation, and M is the number of classes.
[0066] Furthermore, in the embodiments of this application, the high-value sparse representation H T ∈R (B,N) can be mapped to the BEV feature feat bev ∈R (B,C,h×w) and the features at the corresponding positions are collected table to obtain the high-value feature feat T ∈R (B,C,N) . Then, these class embedding vectors can be integrated into the high-value features to enhance the feature representation. Finally, this fusion can generate a set of sparse features feat S ∈R (B,C,N) , which can contain geometric and class information and can be used for the subsequent temporal fusion module 14.
[0067]
[0068] As Figure 5 shown in Figure 5 is the structural schematic diagram of the temporal fusion module of the sparse temporal fusion framework for 3D object detection based on point cloud provided by the embodiments of this application. In the embodiments of this application, the temporal fusion module 14 divides the time series into the current frame and the historical frames, and thus processes the features extracted from these two types of frames to efficiently capture the continuous information at different times. Specifically, assuming that the total frame length is T and the current time is t. The current sparse feature can be M c ={f t} and the historical sparse features can be represented as M h ={f t-1 , f t-2 , f t-3 , …, f t-m}, where m represents the number of historical frames.
[0069] In the embodiment of the present application, the current sparse feature M c ∈R (B,C,N) and the historical sparse feature M h ∈R (B,C,N) , successfully retain the key information and filter the noise, but some clues crucial for object detection may still be excluded. To make full use of these potentially valuable details, in the embodiment of the present application, the temporal fusion module 14 also adopts a self-attention mechanism for encoding. In this mechanism, the current sparse feature M c ∈R (B,C,N) or the historical sparse feature M h ∈R (B,C,N) is used as a query to extract relevant information from the original bird's-eye view feature feat bev ∈R (B,C,h×w) . Therefore, the framework of the embodiment of the present application can adaptively determine the location and content of the information extracted from the original frame data. This significantly improves the efficiency and accuracy of temporal data processing.
[0070] In the embodiment of the present application, two position embedding vectors can be further used in the temporal fusion module 14: the position embedding vector P for sparse features s ∈R (B,N,2) and the position embedding vector P for the original BEV features bev ∈R (B ,h×w,2) . This ensures that the framework of the embodiment of the present application not only focuses on the content of each feature but also considers their spatial relationships in the input sequence. Finally, it can effectively capture the current feature M" c ∈R (B,C,N) and the historical feature M" h ∈R (B,C,N) , improving the understanding and processing ability of dynamic scenes.
[0071] In addition, a cross-attention mechanism can also be used in the temporal fusion module 14. For example, a projection layer can be applied to the current feature M" c ∈R (B,C,N) and the historical feature M" h ∈R (B,C,N) . Layer normalization can be further performed on the projected features to enhance the model's representation ability. Finally, the projected current feature J c ∈R (B,C,N) or the projected historical feature J h ∈R(B,C,N) .
[0072]
[0073] PRO = LayerNorm[Linear(·)] (14)
[0074] In the temporal fusion module 14, a cross-attention mechanism can be adopted to introduce the interaction between the current frame and the historical frames, so as to achieve efficient information fusion. For example, a cross-attention network is applied to the projected features, using the projected current feature J c ∈R (B,C,N) as the query (Query) to extract relevant information from the projected historical feature J h ∈R (B,C,N) .
[0075] In addition, the temporal fusion module 14 can also introduce the current position vector P c ∈R (B,N,2) and the historical position vector P h ∈R (B,N,2) , enabling the model to better understand and capture the sequential and structural relationships between the current and historical features. This design improves the accuracy of temporal information fusion and enhances the model's ability to understand dynamic scenes.
[0076] Therefore, the framework provided in this application extracts the bird's-eye view (BEV) features of the historical frames and the current frame. In the historical feature extraction stage, the gradient does not backpropagate to reduce memory consumption. Subsequently, using the vehicle motion information, the historical features are aligned to the coordinate system of the current moment through an affine transformation (multi-frame historical feature alignment module, MFAM). Then, sparse feature extraction is performed to obtain the sparse heat features at each timestamp (sparse feature extraction module, SFEM). Through these processes, the foreground sparse features of the current frame and the historical frames can be obtained, and these features have an inherent relationship in time and space. Finally, a temporal attention mechanism is used for temporal feature fusion to obtain the final current fusion feature for object recognition and prediction.
[0077] Embodiment 2
[0078] The embodiment of this application also provides a sparse temporal fusion method for detecting three-dimensional objects based on point clouds. As Figure 6 shown Figure 6 is the flowchart of the embodiment of the sparse temporal fusion method for detecting three-dimensional objects based on point clouds provided in this application. Figure 6 The sparse temporal fusion method for detecting three-dimensional objects based on point clouds shown in
[0079] S601. Extract features from the point cloud data of the current frame and the historical frames that fall into each 3D voxel in a pre-divided 3D voxel network to generate the current frame features and the historical frame features.
[0080] In step S601, a regular grid structure (such as voxels or cylinders) can be used to organize the irregular point cloud, and a 2D or 3D convolutional neural network (CNN) can be employed to extract dense features. Specifically, both the cylinder and voxel methods discretize the 3D point cloud data and achieve object detection by extracting features from specific regions in the grid. However, the voxel method can usually learn a richer feature representation because it can capture more spatial relationships in the point cloud data. In the embodiments of the present application, in step S601, the voxel method can be used for feature extraction, and it can include a voxel feature encoding layer for extracting features from the point cloud data that falls into each voxel to generate the current frame features and the historical frame features. S602. Perform a series of affine transformations on the historical frame features using the vehicle chassis information of the target vehicle and transform them into the coordinate system of the current moment.
[0081] In step S602, a pose transformation can be applied to the historical frame features extracted in step S601 based on the position and orientation of the vehicle. Assume the current moment is t 0 , and the historical time series can be correspondingly represented as {t -1 , t -2 , t -3 , …, t -m}. First, the relative position (x, y) of the target vehicle and the ego vehicle direction angle in the bird's-eye view can be used to calculate the vehicle's direction angle angle ego and displacement T bev based on the following formulas (1)-(3): bev :
[0082]
[0083]
[0084] angle bev = angle ego - angle tra (3)
[0085] As Figure 2 shown, Figure 2 is a schematic diagram of the vehicle's direction and position information in the bird's-eye view used in the embodiments of the present application. In the above formulas, T bev is the displacement amount of the vehicle in the bird's-eye view, angle tra is the moving angle, and angle bev is the direction angle of the vehicle in the bird's-eye view.
[0086] Then, the angle, displacement, and grid coordinates (grid 1 , grid h ) of the vehicle in the bird's-eye view can be used to calculate the offsets shift x and shift y of the historical frame features in the horizontal and vertical directions in the grid coordinates through the following formulas (4) and (5):
[0087]
[0088] Based on the offsets shift x and shift y of the historical frame features in the bird's-eye view and the rotation angle angle rot of the vehicle, an affine transformation is performed on the historical frame features based on the following formula (6) to obtain the transformed historical frame features pre-feat ∈ R (B,C,h,w) :
[0089]
[0090] In the above formula, feat ∈ R (B,C,h,w) represents the original features in the input shared network, B represents the number of samples input at one time, C represents the number of channels of the input features, h represents the height of the input features, and w represents the width of the input features. Then, the transformed historical frame features are combined with the features at the last time point in the historical time series {t -1 , t -2 , t -3 , …, t -m} to form a fused feature representation, as shown in the following formula (7).
[0091]
[0092] In formula (7), represents the concatenation operation.
[0093] S603. Extract bird's-eye view features from the current frame features and historical frame features, extract a dense heat map from the bird's-eye view features, retain only the maximum value in the local area of the heat map to generate the current frame sparse representation and historical frame sparse representation, sort the feature vectors in the current frame sparse representation and historical frame sparse representation, and select the top predetermined number of features to create a high-value sparse representation.
[0094] In existing object detection tasks based on point cloud data, generating dense feature representations is considered to improve object detection performance. However, there are many problems in practical applications, such as noise generation and low efficiency. Specifically, generating dense features requires dense prediction of the entire point cloud, which generates a large amount of background noise during the processing. These noises may interfere with the accurate recognition of the model, especially when dealing with complex scenes. In the embodiments of the present application, a time data partitioning strategy is adopted to divide the input data into current frame information and historical frame information. To fully utilize the potential information in these data, in step S603, sparse feature extraction can be performed on each part separately.
[0095] In the embodiments of the present application, a shared network composed of, for example, 3×3 convolutional layers can be constructed. This network is used to learn BEV (Bird's Eye View) features from the current frame and historical frames and extract relevant information. The original feature input to the shared network is feat ∈ R (B,C,h,w) 。
[0096] feat bev =Shared_Conv(feat) (8)
[0097] In step S603, a series of predicted heatmaps can be output using the BEV feature maps from the current frame and historical frames. Each predicted heatmap corresponds to a specific object category. The peak in the heatmap represents the center position of each category of objects in the BEV feature map. To generate the ground truth heatmap, in the embodiments of the present application, the Gaussian distribution mechanism is used to "draw" the center of each object into a region, creating a two-dimensional ground truth heatmap with a size of h×w. Since each object may belong to different categories, a ground truth heatmap needs to be generated for each category. These heatmaps are stacked according to the number of categories to form a three-dimensional ground truth heatmap, where k represents the number of different categories. During this process, the center position of each object affects the heatmap according to the Gaussian function, which can be expressed by the following formula (9):
[0098]
[0099] where k i represents the category of the i-th object, and σ i is the standard deviation related to the object scale.
[0100] The ground truth heatmap generated by the Gaussian distribution function can be used as the supervision information during the training process in the embodiments of the present application, guiding the predicted heatmap to gradually align with the ground truth heatmap, so as to generate a dense heatmap DH ∈ R (B,K,h×w) 。
[0101] To reduce the interference of irrelevant features in the spatial domain, in the embodiments of the present application, only the maximum response value within the local region of each position in the generated dense heat map can be retained. Thus, these local maximum elements can be extracted and aggregated based on the following formula (10) to form a new sparse representation map SH ∈ R (B,K,h×w) . This sparse representation effectively suppresses the responses in non-target regions, thereby highlighting potential target positions and improving the performance and accuracy of subsequent detection tasks:
[0102] SH(x, y, k i ) = max(DH(x, y, k i )) (10)
[0103] To extract the most likely target candidate regions from the sparse representation and assign class information to these regions, in step S603, the elements in the sparse representation can first be sorted, and the top N elements can be extracted to create a high-value sparse representation. Then, each class in the high-value sparse representation is converted into a one-hot vector (One-Hot Encoding). These one-hot vectors can be further converted into learnable class embedding vectors based on the following formula (11). These embeddings capture the semantic information of the classes and provide valuable target prior knowledge during the prediction process, enabling the network to focus more on the changes in the internal relationships of the classes.
[0104]
[0105] In the above formula (11), represents the learnable class embedding vector, represents each class in the high-value sparse representation, and M is the number of classes.
[0106] In addition, in the embodiments of the present application, the high-value sparse representation H T ∈ R (B,N) can be mapped to the BEV feature feat bev ∈ R (B,C,h×w) and the features at the corresponding positions are collected table to obtain the high-value feature feat T ∈ R (B,C,N) . Then, these class embedding vectors can be integrated into the high-value features to enhance the feature representation. Finally, this fusion can generate a set of sparse features feat S ∈ R (B,C,N) , which can contain geometric and class information and can be used for subsequent steps.
[0107] S604. For the current frame high-value sparse representation and the historical frame sparse representation, extract the current sparse feature and the historical sparse feature respectively, perform projection processing on the current sparse feature and the historical sparse feature to obtain the projected current feature and the projected historical feature, use the projected current feature to extract relevant information from the projected historical feature, and extract relevant information from the bird's-eye view feature based on the current sparse feature, the historical sparse feature, the position embedding vector for the sparse feature, and the position embedding vector for the bird's-eye view feature to generate the current feature.
[0108] In the embodiment of the present application, in step S604, the time series can be divided into the current frame and the historical frame, so as to process the features extracted from these two types of frames to efficiently capture the continuous information at different moments. Specifically, assume that the total frame length is T and the current moment is t. The current sparse feature can be M c ={f t} and the historical sparse feature can be represented as M h ={f t-1 , f t-2 , f t-3 , …, f t-m}, where m represents the number of historical frames.
[0109] In the embodiment of the present application, the current sparse feature M c ∈R (B,C,N) and the historical sparse feature M h ∈R (B,C,N) , have successfully retained the key information and filtered the noise, but some clues that are crucial for object detection may still be excluded. To make full use of these potentially valuable details, in the embodiment of the present application, a self-attention mechanism is also adopted in step S604 for encoding. In this mechanism, the current sparse feature M c ∈R (B,C,N) or the historical sparse feature M h ∈R (B,C,N) , is used as a query (Query) to extract relevant information from the original bird's-eye view feature feat bev ∈R (B,C,h×w) . Therefore, the method of the embodiment of the present application can adaptively determine the position and content of extracting information from the original frame data. This significantly improves the efficiency and accuracy of time series data processing.
[0110] In the embodiment of the present application, in step S604, two types of position embedding vectors can be further used: the position embedding vector P s ∈R (B,N,2) for the sparse feature and the position embedding vector P bev ∈R (B,h×w,2) This ensures that the framework of the embodiments of the present application not only focuses on the content of each feature but also considers their spatial relationships in the input sequence. Ultimately, it can effectively capture the current feature M" c ∈R (B,C,N) and the historical feature M" h ∈R (B,C,N) , enhancing the ability to understand and process dynamic scenarios.
[0111] In addition, a cross-attention mechanism can also be used in step S604. For example, the projection layer can be applied to the current feature M" c ∈R (B ,C,N) and the historical feature M" h ∈R (B,C,N) . Layer normalization (LayerNormalization) can further be performed on the projected features to enhance the representational ability of the model. Ultimately, the projected current feature J c ∈R (B,C,N) or the projected historical feature J h ∈R (B,C,N) can be obtained.
[0112] J c = PRO(M″ c ) J h = PRO(M″ h ) (13)
[0113] PRO = LayerNorm[Linear(·)] (14)
[0114] In step S604, a cross-attention mechanism can be adopted to introduce the interaction between the current frame and the historical frame to achieve efficient information fusion. For example, the cross-attention network is applied to the projected features, using the projected current feature J c ∈R (B ,C,N) as the query to extract relevant information from the projected historical feature J h ∈R (B,C,N) .
[0115] In addition, in the embodiments of the present application, the current position vector P c ∈R (B,N,2) and the historical position vector P h ∈R (B,N,2), enabling the model to better understand and capture the sequential and structural relationships between current and historical features. This design improves the accuracy of temporal information fusion and enhances the model's ability to understand dynamic scenarios. Therefore, the method provided in this application extracts the bird's-eye view (BEV) features of historical and current frames. During the historical feature extraction stage, the gradient does not backpropagate to reduce memory consumption. Subsequently, using vehicle motion information, the historical features are aligned to the coordinate system at the current moment through affine transformation (Multi-Frame Alignment Module, MFAM). Then, sparse feature extraction is performed to obtain sparse heat features at each timestamp (Sparse Feature Extraction Module, SFEM). Through these processes, foreground sparse features of the current and historical frames are obtained, which have inherent relationships in time and space. Finally, a temporal attention mechanism is used for temporal feature fusion to obtain the final current fusion features for object recognition and prediction.
[0116] Embodiment III
[0117] The internal functions and structures of the sparse temporal fusion framework for detecting three-dimensional objects based on point clouds are described above, which can be implemented as an electronic device. Figure 7 The following is a schematic structural diagram of an embodiment of the electronic device provided in this application. As Figure 7 shown, the electronic device includes a memory 71 and a processor 72.
[0118] The memory 71 is used to store programs. In addition to the above programs, the memory 71 can also be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0119] The memory 71 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0120] The processor 72 is not limited to a processor (CPU), but may also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip and other processing chips. The processor 72 is coupled to the memory 71 and executes the programs stored in the memory 71 to perform the sparse temporal fusion method for detecting three-dimensional objects based on point clouds in the above embodiments.
[0121] Further, as Figure 7As shown, the electronic device may further include: other components such as a communication component 73, a power supply component 74, an audio component 75, and a display 76. Figure 7 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 7 the components shown.
[0122] The communication component 73 is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device may access a wireless network based on a communication standard, such as WiFi, 3G, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 73 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 73 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0123] The power supply component 74 provides power for various components of the electronic device. The power supply component 74 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.
[0124] The audio component 75 is configured to output and / or input audio signals. For example, the audio component 75 includes a microphone (MIC). When the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals may be further stored in the memory 71 or transmitted via the communication component 73. In some embodiments, the audio component 75 further includes a speaker for outputting audio signals.
[0125] The display 76 includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations.
[0126] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A sparse temporal fusion framework for detecting three-dimensional objects based on point clouds, characterized in that: The framework includes: A feature extraction module is used to extract features of the point cloud data of the current frame and the historical frame falling into each three-dimensional voxel in the pre-divided three-dimensional voxel network to generate current frame features and historical frame features; A multi-historical feature alignment module uses the vehicle chassis information of the target vehicle to perform a series of affine transformations on the historical frame features and converts them into the coordinate system of the current moment; A sparse feature extraction module is used to extract bird's-eye view features from current frame features and historical frame features, extract a dense heat map from the bird's-eye view features, retain only the maximum value in a local area of the heat map to generate a current frame sparse representation and a historical frame sparse representation, sort the feature vectors in the current frame sparse representation and the historical frame sparse representation and select a predetermined number of features to create a high-value sparse representation; A temporal fusion module is used to extract current sparse features and historical sparse features for the current frame high-value sparse representation and the historical frame sparse representation respectively, project the current sparse features and the historical sparse features to obtain projected current features and projected historical features, use the projected current features to extract relevant information from the projected historical features, and extract relevant information from the bird's-eye view features based on the current sparse features and the historical sparse features as well as the position embedding vector for the sparse features and the position embedding vector for the bird's-eye view features to generate current features.
2. The sparse temporal fusion framework for detecting three-dimensional objects based on point clouds according to claim 1, characterized in that: The feature extraction module includes a voxel feature encoding layer, which is used to extract features from point cloud data falling into each voxel to generate the current frame features and historical frame features.
3. The sparse temporal fusion framework for detecting three-dimensional objects based on point clouds according to claim 1, characterized in that: The sparse feature extraction module is further used to: use one-hot encoding to convert each category in the high-value sparse representation into a one-hot vector, convert these one-hot vectors into machine-learnable category embedding vectors, and aggregate the category embedding vectors into the features of the high-value sparse representation to generate sparse features.
4. The sparse temporal fusion framework for detecting three-dimensional objects based on point clouds according to claim 1, characterized in that: The multi-history feature alignment module is used to: Using the relative position (x, y) of the target vehicle and the direction angle of the vehicle in the bird's-eye view ego Calculate the vehicle's direction angle bev and displacement T bev : angle bev =angle ego -angle tra Among them, T bev is the displacement of the vehicle in the bird's-eye view, angle tra is the moving angle, angle bev is the direction angle of the vehicle in the bird's-eye view, Then, the vehicle's angle, displacement, and grid coordinates (grid1, grid h ), calculate the horizontal and vertical shift of the historical frame features in the grid coordinates x and shift y : Offset shift in bird's-eye view based on historical frame features x and shift y And the vehicle's rotation angle angle rot , perform affine transformation on the historical frame features to obtain the transformed historical frame features pre_feat∈R (B,C,h,w) : Among them, feat∈R (B,C,h,w) represents the original features in the input shared network, B represents the number of samples input at a time, C represents the number of channels of the input features, h represents the height of the input features, and w represents the width of the input features. The converted historical frame features are compared with the historical time series {t -1 ,t -2 ,t -3 ,…,t -m } to form a fused feature representation in, Represents a connection operation, and t0 in the historical time series represents the current time.
5. A sparse temporal fusion method for detecting three-dimensional objects based on point clouds, characterized in that: The method comprises: In the pre-divided three-dimensional voxel network, feature extraction is performed on the point cloud data of the current frame and the historical frame falling into each three-dimensional voxel to generate current frame features and historical frame features; Using the chassis information of the target vehicle, a series of affine transformations are performed on the historical frame features to convert them into a coordinate system at the current moment; Extracting bird's-eye view features from current frame features and historical frame features, extracting a dense heat map from the bird's-eye view features, retaining only the maximum value in a local area of the heat map to generate a current frame sparse representation and a historical frame sparse representation, sorting feature vectors in the current frame sparse representation and the historical frame sparse representation and selecting a predetermined number of features to create a high-value sparse representation; For the current frame high-value sparse representation and the historical frame sparse representation, current sparse features and historical sparse features are extracted respectively, and the current sparse features and historical sparse features are projected to obtain projected current features and projected historical features, and relevant information is extracted from the projected historical features using the projected current features, and relevant information is extracted from the bird's-eye view features based on the current sparse features and the historical sparse features as well as the position embedding vector for the sparse features and the position embedding vector for the bird's-eye view features to generate the current features.
6. An electronic device, characterized in that: include: Memory, used to store programs; A processor is used to run the program stored in the memory to execute the sparse time fusion method for detecting three-dimensional objects based on point cloud as described in claim 5.
7. A computer-readable storage medium having a computer program executable by a processor stored thereon, characterized in that: When the program is executed by the processor, the sparse time fusion method for detecting three-dimensional objects in the laser radar point cloud as described in claim 5 is implemented.
Citation Information
Cited By
BEV dense and sparse hybrid multi-task sensing method and device, electronic equipment and storage medium
CN121582759A