Automatic driving target detection system and method
By fusing camera and radar data in the autonomous driving system, extracting and interacting BEV features, the problem of poor single-modal target detection is solved, and more accurate and robust automatic driving target detection is achieved, especially in an occlusion environment, which can effectively estimate the target speed.
Patent Information
- Application Number
- CN202411892841.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-13
AI Technical Summary
In existing autonomous driving systems, single modal object detection is poor, especially in complex and dynamic driving environments, it is difficult to accurately detect the position and speed of the obstructed object.
An autonomous driving target detection system is adopted to obtain circumferential image data and radar point cloud data through camera network branches and radar network branches, and extract camera BEV features and radar BEV features respectively. Then, through the historical feature timing fusion module and the radar-camera feature interaction attention module, these features are fused and interacted, and finally regression prediction and classification prediction are performed through the fusion detection head network to achieve target detection.
It effectively solves the problem of poor single-modal object detection, improves the accuracy and robustness of target detection in autonomous driving, especially in severe occlusion environments, which can better estimate the speed of the target.
Smart Images

Figure CN119992265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to an automatic driving target detection system and method. Background Art
[0002] In recent years, autonomous driving technology has attracted much attention due to its ability to reduce the burden of driving and improve driving safety. The perception system is an integral part of the autonomous driving algorithm, which aims to accurately estimate the state of the environment and provide reliable data for prediction and planning. 3D object detection can intelligently predict the position, size and category of key 3D objects near the autonomous driving vehicle, and is an important part of the perception system. However, due to the complex and dynamic driving environment, it is still a difficult task to achieve fully autonomous driving. Commonly used sensors in current autonomous driving systems include cameras, lidars, and millimeter-wave radars, each with its own advantages and disadvantages. The 2D image data generated by the camera has high accuracy in perceiving the shape and category of objects, but is easily affected by weather and light. Lidar can make up for its shortcomings and provide fine-grained 3D structural information. It is suitable for detecting 3D objects and is not affected by weather and time, but has fewer reflection points for long-distance or small objects. By fusing the data of cameras and lidars, their respective shortcomings can be compensated to achieve accurate and reliable perception in autonomous driving systems.
[0003] The current multimodal sensor data fusion technology is divided into: pre-fusion, feature-level fusion and post-fusion. Pre-fusion allows data to be fused earlier, but the direct fusion performance of heterogeneous data is poor; the post-fusion method is simple, but a lot of information is lost in the middle, which affects the detection results; feature-level fusion extracts intermediate features from each sensor through different networks, which can more fully mine the information of each modality, and thus is more likely to achieve the best effect. However, the image data of the camera is in a perspective perspective, while the point cloud data of the lidar is in 3D space. The two are essentially heterogeneous, and the direct fusion of the intermediate features of the two is not effective. In addition, the existing technical solutions are poor in estimating the moving speed of occluded objects and cannot solve the problem of severe occlusions well. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide an autonomous driving target detection system and method to solve at least one problem existing in the prior art.
[0005] In order to achieve the above-mentioned purpose, a first aspect of an embodiment of the present invention provides an autonomous driving target detection system, the system comprising: a camera network branch, used to obtain surround view image data of a vehicle, and extract camera BEV features based on the surround view image data; a radar network branch, used to obtain radar point cloud data of the vehicle, and extract radar BEV features based on the radar point cloud data; a historical feature temporal fusion module, used to extract temporal fusion camera BEV features based on the camera BEV features, and to extract temporal fusion radar BEV features based on the radar BEV features; a radar-camera feature interactive attention module, used to interact the temporal fusion camera BEV features with the temporal fusion radar BEV features to obtain the interacted camera BEV features; and a fusion detection head network, used to perform regression prediction and classification prediction based on the interacted camera BEV features and the temporal fusion radar BEV features to obtain an autonomous driving target detection result, wherein the surround view image data comprises a historical frame image and a current frame image, and the radar point cloud data comprises a historical frame radar point cloud and a current frame radar point cloud.
[0006] Optionally, the camera network branch includes: a first acquisition module, used to acquire the surround view image data of the vehicle and the depth information of the radar; an image segmentation module, used to segment the surround view image data to obtain an image list, each element of the image list contains an image of one frame; an image front view feature extraction network, used to extract multi-scale front view features for each element of the image list; a feature pyramid network, used to perform multi-scale fusion of the multi-scale front view features corresponding to each of the elements to obtain multi-scale fused image front view features corresponding to each of the elements; and a view conversion module, used to obtain the camera BEV feature based on the multi-scale fused image front view features corresponding to each of the elements and the depth information of the radar.
[0007] Optionally, the camera network branch also includes: an auxiliary classification head, used to calculate the auxiliary classification head loss based on the front view features of the multi-scale fusion image, wherein the auxiliary classification head loss is used to update the camera network branch when training the system.
[0008] Optionally, the radar network branch includes: a second acquisition module, used to acquire radar point cloud data of the vehicle; a voxel grid division module, used to divide the 3D space of the radar point cloud data into voxels of a preset size; a grouping operation module, used to distribute the point cloud into each voxel in sequence; a random sampling module, used to determine whether the number of point clouds of each voxel exceeds a preset first number. If the number of point clouds of a voxel exceeds the first number, the number of point clouds of the voxel is reduced to the first number through random sampling; and a voxel feature encoding module, used to encode, extract features and multi-scale fuse the voxels to obtain radar BEV features.
[0009] Optionally, the historical feature time series fusion module includes: a short-term time series fusion module, used to perform time series fusion on the BEV features of each two adjacent frames to obtain a short-term time series fusion feature; and a long-term time series fusion module, used to perform time series fusion on the short-term time series fusion features to obtain a long- and short-term time series fusion feature, wherein the BEV feature is a camera BEV feature or a radar BEV feature. When the BEV feature is a camera BEV feature, the long- and short-term time series fusion feature is a time series fusion camera BEV feature; when the BEV feature is a radar BEV feature, the long- and short-term time series fusion feature is a time series fusion radar BEV feature.
[0010] Optionally, performing temporal fusion on the BEV features of each two adjacent frames to obtain short-term temporal fusion features includes: performing spatiotemporal aligning on the BEV features of each two adjacent frames; and performing feature fusion on the BEV features of each two adjacent frames after spatiotemporal alignment to obtain short-term temporal fusion features.
[0011] Optionally, the radar-camera feature interactive attention module includes: a self-attention enhancement module, used to enhance the time-series fusion camera BEV feature, and perform a first processing to obtain a query vector; a deformable attention module, used to use the time-series fusion radar BEV feature as a key vector and a numerical vector to interact with the query vector to obtain the interacted BEV feature; and an interactive feature output module, used to perform a first processing, a forward propagation and a first processing on the interacted BEV feature in sequence to obtain the interacted camera BEV feature, wherein the first processing includes addition and normalization.
[0012] Optionally, the fusion detection head network includes: a dynamic hybrid feature fusion module, used to dynamically hybridize the interacted camera BEV features and the time series fusion radar BEV features to obtain dynamic hybrid fusion features; a heat map generator, used to extract features from the dynamic hybrid fusion features to obtain a dense heat map; and a decoupling detection head, used to perform classification prediction based on the dense heat map and the dynamic hybrid fusion features to obtain a category prediction result, and perform regression prediction based on the dense heat map and the time series fusion radar BEV features to obtain a regression prediction result, and the category prediction result and the regression prediction result together constitute the target detection result.
[0013] Optionally, the interactive camera BEV features and the time-series fusion radar BEV features are dynamically mixed and fused to obtain dynamic mixed fusion features, including: preliminarily fusing the interactive camera BEV features and the time-series fusion radar BEV features to obtain preliminary fusion features; sending the preliminary fusion features to the fine channel attention branch and the coarse channel attention branch respectively to obtain fine channel attention features and coarse channel attention features; merging the fine channel attention features and the coarse channel attention features into coarse and fine channel attention fusion features, and sending them to the fine space attention branch and the coarse space attention branch respectively to obtain fine space attention features and coarse space attention features; and fusing the fine space attention features with the coarse space attention features to obtain the dynamic mixed fusion features.
[0014] On the other hand, the present invention provides an autonomous driving target detection method, the method comprising: acquiring surround view image data of a vehicle, and extracting camera BEV features based on the surround view image data; acquiring radar point cloud data of the vehicle, and extracting radar BEV features based on the radar point cloud data; extracting time-series fusion camera BEV features based on the camera BEV features, and extracting time-series fusion radar BEV features based on the radar BEV features; interacting the time-series fusion camera BEV features with the time-series fusion radar BEV features to obtain interacted camera BEV features; and performing regression prediction and classification prediction based on the interacted camera BEV features and the time-series fusion radar BEV features to obtain autonomous driving target detection results, wherein the surround view image data includes historical frame images and current frame images, and the radar point cloud data includes historical frame radar point clouds and current frame radar point clouds.
[0015] The system proposed in the present invention fuses the camera BEV features and the radar BEV features when performing target detection, effectively solving the problem of poor single-modal target detection effect. The historical feature temporal fusion module integrates historical frame information in both the camera BEV features and the radar BEV features, and incorporates implicit time dimension information, which is not only beneficial for target speed estimation, but also helps solve the problem of severe occlusion obstacles. The radar-camera feature interactive attention module interacts the temporal fusion camera BEV features and the temporal fusion radar BEV features, improving the performance degradation of the camera BEV features due to the lack of explicit deep supervision. Through the above technical solution, the effect of target detection in autonomous driving is greatly improved, which helps to improve the accuracy and robustness of the autonomous driving system.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0018] Figure 1 It is a schematic diagram of the network structure of an automatic driving target detection system provided by one embodiment of the present invention.
[0019] Figure 2 It is a schematic diagram of the structure of a camera network branch provided by an embodiment of the present invention.
[0020] Figure 3 It is a schematic diagram of the structure of an image front view feature extraction network provided by an embodiment of the present invention.
[0021] Figure 4 It is a schematic diagram of the structure of a feature pyramid network provided by an embodiment of the present invention.
[0022] Figure 5 It is a structural diagram of a view conversion module provided by an embodiment of the present invention.
[0023] Figure 6 It is a schematic diagram of the structure of a radar network branch provided by an embodiment of the present invention.
[0024] Figure 7 It is a structural diagram of a historical feature time series fusion module provided by an embodiment of the present invention.
[0025] Figure 8 It is a structural diagram of a radar-camera feature interactive attention module provided by an embodiment of the present invention.
[0026] Fig. 9 It is a schematic diagram of the structure of a fusion detection head network provided by an embodiment of the present invention.
[0027] Fig.10 It is a structural diagram of a dynamic hybrid feature fusion module provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0029] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, and models may be mentioned, which should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0030] Figure 1 FIG. 1 is a schematic diagram of a network structure of an automatic driving target detection system provided by an embodiment of the present invention. Figure 1 As shown, the system provided by the present invention includes the following structure.
[0031] The camera network branch is used to obtain surround view image data of the vehicle and extract camera BEV features based on the surround view image data.
[0032] The radar network branch is used to obtain radar point cloud data of the vehicle and extract radar BEV features based on the radar point cloud data.
[0033] The historical feature time series fusion module is used to extract and fuse camera BEV features based on the camera BEV features, and to extract and fuse radar BEV features based on the radar BEV features.
[0034] The radar-camera feature interactive attention module is used to interact the time-series fusion camera BEV feature and the time-series fusion radar BEV feature to obtain the interacted camera BEV feature.
[0035] The fusion detection head network is used to perform regression prediction and classification prediction based on the interacted camera BEV features and the time-series fusion radar BEV features to obtain an autonomous driving target detection result.
[0036] The surround view image data includes a historical frame image and a current frame image, and the radar point cloud data includes a historical frame radar point cloud and a current frame radar point cloud.
[0037] First, the camera and radar network branches proposed in the present invention extract dual-modal BEV features, namely camera BEV features and radar BEV features, respectively, which solves the problem of poor target detection effect under a single modality. Secondly, the proposed historical time series fusion module and radar-camera feature interactive attention module introduce historical frames, among which the historical time series fusion module additionally introduces implicit time information, which is not only conducive to the target speed estimation, but also helps to solve the problem of severe occlusion obstacles, while the radar-camera feature interactive attention module interacts the time series fusion camera BEV features and the time series fusion radar BEV features, which improves the performance degradation of the camera BEV features due to the lack of explicit deep supervision. Finally, the fusion detection head network performs regression prediction and classification prediction based on the two types of BEV features, realizing the effective fusion of dual-modal BEV features.
[0038] Figure 2 is a schematic diagram of the structure of a camera network branch provided by an embodiment of the present invention, such as Figure 2 As shown, the camera network branch includes the following structure.
[0039] The first acquisition module is used to acquire the surround view image data of the vehicle and the depth information of the radar. The surround view image data size of the vehicle is 40×3×384×1056, where 40 means that 7 historical frames plus 1 current frame are input, and one surround view image is randomly discarded for each frame of 6, so there are 40 images. 3 represents the original RGB channel of the image, and 384 and 1056 represent the size of the image after image preprocessing.
[0040] The image splitting module is used to split the surround image data to obtain an image list, each element of which contains an image of one frame. The module splits multiple frames of images with a size of 40×3×384×1056 into a list, each list element is a single frame surround image with a size of 5×3×384×1056, and the list length is 8.
[0041] The image front view feature extraction network is used to extract multi-scale front view features for each element of the image list. Figure 3 FIG. 1 is a schematic diagram of the structure of a front view feature extraction network for an image provided by an embodiment of the present invention. Figure 3As shown in Figure 1, the processing of the network includes five stages, including Stage 0 to Stage 4, and the convolution operations of each stage are shown in Table 1, Table 2, Table 3 and Table 4. The network uses the ResNet50 network to extract the front view features of the surround view image. Specifically, the input surround view image is processed by STAGE 2, STAGE 3, and STAGE 4 in the ResNet50 network to obtain three features of different sizes, 5×512×48×132, 5×1024×24×66, and 5×2048×12×33, as multi-scale front view features.
[0042] Table 1 Stage 0 & Stage 1 of the image front view feature extraction network
[0043]
[0044] Table 2 Stage 2 of the image front view feature extraction network
[0045]
[0046] Table 3 Stage 3 of the image front view feature extraction network
[0047]
[0048] Table 4 Stage 4 of the image front view feature extraction network
[0049]
[0050] The feature pyramid network is used to perform multi-scale fusion on the multi-scale front view features corresponding to each of the elements to obtain the multi-scale fused image front view features corresponding to each of the elements. Figure 4 is a schematic diagram of the structure of a feature pyramid network provided by an embodiment of the present invention. The specific convolution operation parameters are shown in Table 5. Please refer to Figure 4 As shown in Table 5, the network fuses the front view features of three scales through upsampling and convolution. Finally, a large-scale feature with a size of 5×128×48×132 is obtained, which fuses the medium- and small-scale feature information of the feature map size of 24×66 and 12×33, and outputs it as the front view feature of the multi-scale fused image.
[0051] Table 5 Feature pyramid network, radar deep feature extraction network, deep network
[0052]
[0053] A view conversion module is used to obtain the camera BEV feature based on the multi-scale fusion image front view feature corresponding to each element and the depth information of the radar. Figure 5is a schematic diagram of the structure of a view conversion module provided by an embodiment of the present invention. Figure 5 As shown in the figure, the module first inputs the radar depth information of size 5×1×384×1024 into the radar depth feature extraction network and outputs the radar depth feature. The radar depth feature and the front view feature of the surround image of size 5×128×48×132, that is, the front view feature of the multi-scale fusion image, are concatenated in the channel dimension and added to the depth extraction network to obtain a feature of size 5×150×48×132, of which the first 118 dimensions are depth features and the last 32 dimensions are image features. The feature is sent to the spilt module, and the coordinates of each front view feature are transformed to obtain an image pseudo point cloud feature of size 5×118×32×48×132. The feature is then sent to the Pooling module, which is assigned to a 256×256 BEV grid, and finally sum-pooling is performed to obtain a camera BEV feature of size 32×256×256.
[0054] It should be noted that the radar depth information is true value information, which is obtained from a pre-established data set.
[0055] For further information, please refer to Figure 2 , the camera network branch also includes: an auxiliary classification head, used to calculate the auxiliary classification head loss based on the front view features of the multi-scale fusion image, wherein the auxiliary classification head loss is used to update the camera network branch when training the system.
[0056] Similar to the classification head of 3D object detection in the prior art, the system proposed in the present invention uses a simple feedforward neural network to implement an auxiliary classification head, whose input is a multi-scale fused image front view feature with a size of 5×128×48×132. The sparse feature size of each labeled object is extracted from the image front view feature according to the centroid of the labeled true value, with a size of 128×150, where 128 represents the number of channels of the feature and 150 represents the maximum number of instances. The feature is then subjected to a set of conv1d+BN+RELU operations and conv1d to obtain a classification feature with a size of 10×150, where 10 represents the number of categories. Finally, the auxiliary classification head loss is calculated based on the classification feature. The auxiliary classification head loss participates in the camera network branch, for example, it is added to the loss calculation with a certain weight, and the camera network branch is optimized and updated together with the loss of other parts. Since the system process proposed in the present invention is relatively long, during network training, if the loss is calculated only by relying on the subsequent fusion detection head, the gradient will not be sufficiently optimized for the camera network branch in the step-by-step reverse conduction. However, after adding the auxiliary classification head, it is helpful to directly optimize the camera network branch, thereby enhancing the ability of the camera network branch in searching for candidate objects and distinguishing different categories.
[0057] Figure 6 This is a schematic structural diagram of a radar network branch provided by an embodiment of the present invention. As Figure 6 shown, the radar network branch includes the following structures.
[0058] A second acquisition module for acquiring the radar point cloud data of the vehicle. The acquired radar point cloud data has a size of N×5, where N indicates that there are N 3D point clouds in the range of -51.2 meters < x < 51.2 meters, -51.2 meters < y < 51.2 meters, and -3 meters < z < 5 meters for this frame of point cloud, and the feature dimension of the point cloud is 5.
[0059] A voxel grid division module for dividing the 3D space of the radar point cloud data into voxels of a preset size. In this module, the 3D space is divided according to voxel sizes of 0.05 meters, 0.05 meters, and 0.2 meters. In specific operations, the max_voxels function is used to set the maximum number of voxels to 120,000. When the number of divided voxels exceeds this maximum number of voxels, this function can automatically discard some voxels to reduce the computational amount.
[0060] A grouping operation module for sequentially distributing the point clouds into each voxel. In the previous step, the 3D space was divided into voxels of a fixed size. In the grouping operation module, the N sparse and unevenly distributed 3D point clouds are sequentially distributed into each voxel to achieve grouping.
[0061] A random sampling module for determining whether the number of point clouds in each voxel exceeds a preset first number. If the number of point clouds in a certain voxel exceeds the first number, then through random sampling, the number of point clouds in this voxel is reduced to the first number. Since the number of 3D point clouds is often more than 100,000, if feature extraction is directly performed at this order of magnitude, it is very computationally resource-consuming. Therefore, through this module, T points are randomly sampled from the voxels containing more than T points, and the resulting voxel size is {N′, T, C}, where N′ is the number of non-empty voxels and C is the number of channels of the feature. For example, the voxels with more than 10 point clouds are reduced to 10. Assuming there are N′ non-empty voxels in total, after being processed by the random sampling module, the voxel size becomes N′×10×5, thus saving the computational amount.
[0062] A voxel feature encoding module for encoding, feature extraction, and multi-scale fusion of the voxels to obtain radar BEV features. The voxel feature encoding module includes three parts, namely a point cloud intermediate encoding module, a radar backbone feature extraction network, and a radar FPN network. The specific convolution parameters of each part are shown in Tables 6, 7, and 8. When the voxel feature encoding module works, first, the average value of each point in the input voxels is taken, and the resulting size is N′×5. Subsequently, it is input into the point cloud intermediate encoding module to obtain point cloud sparse features, with a size of The extracted point cloud sparse features are converted into two-dimensional dense features, and the size of the preliminary point cloud BEV features is 512×256×256. The features are then sent to the radar backbone feature extraction network to extract deep features. The feature dimensions output by three convolution modules (blocks0, 1, 2) are 128×256×256, 256×128×128, and 256×64×64, respectively. These three multi-scale features are then sent to the radar FPN network to obtain multi-scale fused radar BEV features with a size of 384×256×256.
[0063] Table 6 conv_input & encoder_layer1 & encoder_layer2 of the point cloud intermediate encoding module
[0064]
[0065] Table 7 encoder_layer3 & encoder_layer4 & conv_output of the point cloud intermediate encoding module
[0066]
[0067] Table 8 Radar backbone feature extraction network and radar FPN network
[0068]
[0069]
[0070] Figure 7 FIG. 1 is a schematic diagram of the structure of a historical feature time series fusion module provided by an embodiment of the present invention. Figure 7 As shown, the historical feature time series fusion module includes the following structure.
[0071] The short-term time series fusion module is used to perform time series fusion on the BEV features of two adjacent frames to obtain short-term time series fusion features.
[0072] The long-term time series fusion module is used to perform time series fusion on the short-term time series fusion features to obtain long-term and short-term time series fusion features. The specific convolution operation of the long-term time series fusion module is shown in Table 9.
[0073] Among them, the BEV feature is a camera BEV feature or a radar BEV feature. When the BEV feature is a camera BEV feature, the long-term and short-term time series fusion feature is a time series fusion camera BEV feature. When the BEV feature is a radar BEV feature, the long-term and short-term time series fusion feature is a time series fusion radar BEV feature.
[0074] Table 9 Long-term time series fusion module
[0075]
[0076] from Figure 7 It can be seen that the four short-term time series fusion modules fuse the BEV features of two adjacent frames to obtain four short-term time series fusion features, and then embed the time step information to perform long-term time series fusion on the four short-term time series features, and finally obtain the long-term and short-term time series fusion features.
[0077] In the specific implementation, the short-term time series fusion module first performs spatiotemporal alignment on the BEV features of two adjacent frames, and then splices them in the channel dimension. Specifically, due to the movement of the ego-vehicle, a stationary object in the global coordinate system will become a moving object in the ego-vehicle coordinate system. The error in the ego-vehicle coordinate system at time T and time T-1 can be expressed using the following formula (1):
[0078]
[0079] in, is the position of the vehicle in the vehicle coordinate system at time T. It is transformed from the vehicle coordinate system to the world coordinate system, that is, multiplied by a transformation matrix on the left, and is expressed by the following formula (2):
[0080]
[0081] in, is the transformation matrix between the ego-vehicle coordinate system and the global coordinate system at time T. The world coordinate system at time T-1 is transformed into the ego-vehicle coordinate system again, and it is transformed into the transformation matrix of the ego-vehicle coordinate system from time T to time T-1 multiplied by the transformation matrix between the ego-vehicle coordinate system and the global coordinate system at time T, which is expressed by the following formula (3):
[0082]
[0083] in, is the change matrix of the transformation matrix of the ego-vehicle coordinate system at time T and time T-1. It can be seen that the target position translation in the two features is related to the ego-vehicle motion. Therefore, in order to eliminate the influence of the vehicle's motion, we can multiply formula (3) by one before time T-1. As shown in formula (4). and When the multiplication is 0, the influence of the ego-vehicle motion is eliminated, so that the target position translation in the two features is changed from being related to the ego-vehicle motion to the ego-vehicle coordinate system at the same time, thus achieving feature alignment.
[0084]
[0085] After completing the time series alignment, the two features are added together through the Concatenate operation to obtain the short-term time series fusion feature.
[0086] When performing temporal fusion, if two BEV features are directly concatenated and added, it will cause feature misalignment, thus affecting the final result. Therefore, spatiotemporal feature alignment is very important. Through ego-vehicle motion correction, the module learning target can be changed from being related to ego-vehicle motion to being related only to the target motion in the ego-vehicle coordinate system of the current frame, thereby eliminating the influence of ego-vehicle motion.
[0087] In the long-term temporal fusion module, when the BEV feature is a camera BEV feature, each short-term temporal fusion feature of size 64×256×256 is passed through four camera stride convolution blocks, including conv+BN+RELU operations, to obtain four time step embeddings of size 4×256×256, which are added to each short-term temporal fusion feature to obtain a feature size of 68×256×256, and then the Concatenate operation is performed. After the camera long-term temporal convolution module, the camera's long-term and short-term temporal fusion feature size is 32×256×256. When the BEV feature is a radar BEV feature, similar to the camera BEV feature, the size of the output of the four radar stride convolution blocks is 4×256×256, and the size of the output of the radar long-term temporal convolution module is 384×256×256.
[0088] The existing technology is usually limited to introducing short-term temporal information for features, while the present invention adopts a long-term and short-term temporal fusion method. The BEV features of two adjacent frames are fused through a short-term temporal fusion module, and then long-term temporal fusion is performed by embedding time step information. Finally, a BEV fusion feature that combines long-term and short-term is obtained, thereby introducing richer temporal information.
[0089] Figure 8 is a schematic diagram of the structure of a radar-camera feature interactive attention module provided by an embodiment of the present invention, such as Figure 8 As shown, the radar-camera feature interaction attention module includes the following structure.
[0090] The self-attention enhancement module is used to enhance the BEV features of the temporal fusion camera and perform a first process to obtain a query vector, wherein the first process includes addition and normalization.
[0091] The deformable attention module is used to use the time series fusion radar BEV feature as a key vector and a numerical vector to interact with the query vector to obtain the interacted BEV feature.
[0092] The interactive feature output module is used to sequentially perform the first processing, forward propagation and the first processing on the BEV features after the interaction to obtain the camera BEV features after the interaction.
[0093] Please refer to Figure 8 The radar-camera feature interactive attention module first enhances the weaker temporal fusion camera BEV features through a self-attention, and then after the first processing through an Add&Norm layer, it is input into the deformable attention module as the query vector Q, and the temporal fusion radar BEV features are used as the key vector K and the numerical vector V of the deformable attention, so that the radar BEV features can effectively interact with the camera BEV features, and then the interacted BEV features are passed through the Add&Norm layer, and then input into the feedforward network for forward propagation, and finally the interacted camera BEV features can be obtained through an Add&Norm layer. In addition, the deformable attention module and the interactive feature output module can be repeated 3 times, so that the camera BEV features can be effectively enhanced.
[0094] In the specific implementation, the radar-camera feature interactive attention module first inputs the time-series fused camera BEV features and the time-series fused radar BEV features into a conv+BN layer+RELU module to change the number of channels, so that the size of the radar and camera BEV features is 256×256×256. Then the camera and radar BEV features are flattened as vectors. The flattened size of the camera BEV features is 1×65536×256. The radar BEV features are superimposed twice as stronger BEV features and then flattened to 2×65536×256. The flattened camera BEV features are used as the query vector Q, and the flattened radar BEV features are used as the key vector K and the value vector V. The camera BEV features are first self-attention enhanced, and the feature vector size obtained is 1×65536×256. Then, they are added and normalized with the camera BEV features before self-attention enhancement to obtain a feature size of 1×65536×256. Then, the radar BEV features are effectively interacted with using deformable attention, and the feature size obtained is 1×65536×256. Then, the feature size obtained by adding and normalizing with the camera BEV features before deformable attention is 1×65536×256. Then, the feature is sent to the feedforward network for nonlinear transformation, and the feature size obtained is 1×65536×256. Then, the feature size obtained by adding and normalizing with the camera BEV features before sending to the feedforward network is 1×65536×256. After repeating the deformable attention module and the interactive feature output module 3 times, the final feature is reshaped to 256×256×256 as the output result of the module.
[0095] Fig. 9is a schematic diagram of the structure of a fusion detection head network provided by an embodiment of the present invention. Fig. 9 As shown, the fusion detection head network includes the following structure.
[0096] The dynamic hybrid feature fusion module is used to dynamically hybridize the interacted camera BEV feature and the time-series fusion radar BEV feature to obtain a dynamic hybrid fusion feature.
[0097] Specifically, the method for dynamically mixing and fusing the interacted camera BEV features and the time-series fusion radar BEV features to obtain dynamic mixed fusion features includes: preliminarily fusing the interacted camera BEV features and the time-series fusion radar BEV features to obtain preliminary fusion features; sending the preliminary fusion features to the fine channel attention branch and the coarse channel attention branch respectively to obtain fine channel attention features and coarse channel attention features; merging the fine channel attention features and the coarse channel attention features into coarse and fine channel attention fusion features, and sending them to the fine space attention branch and the coarse space attention branch respectively to obtain fine space attention features and coarse space attention features; and fusing the fine space attention features with the coarse space attention features to obtain the dynamic mixed fusion features.
[0098] Fig.10 is a schematic diagram of the structure of a dynamic hybrid feature fusion module provided by an embodiment of the present invention. The specific convolution operations of each branch are shown in Table 10 and Table 11. Please refer to Fig.10, Table 10 and Table 11, the input of this module is the interactive camera BEV features and the time-series fused radar BEV features, both of which are 256×256×256 in size. First, the two are concatenated and added, and then input into the preliminary fusion module (fusion_input) to reduce the number of channels to obtain the preliminary fusion features, with a size of 160×256×256. The preliminary fusion features are input into the fine channel attention branch, which encodes the preliminary fusion features into channel-by-channel weighted features to obtain the fine channel attention features, with a size of 160×256×256. The preliminary fusion features are input into the coarse channel attention branch (c_rough), which encodes the preliminary fusion features into four channel weighted features, and then merges them into the final weighted features to obtain the coarse channel attention features, with a size of 160×256×256. The shared channel attention module is similar to the above-mentioned fine channel attention branch. The two are concatenated and added, and then input into the channel attention fusion module (fusion_channel) to extract features, so that the coarse and fine channel attention fusion feature size is 160×256×256. The feature is input into the fine spatial attention branch (s_refined), and the fine spatial attention branch encodes the feature into a channel-by-channel spatial weighted feature, and the fine spatial attention feature size is 160×256×256. Then the coarse and fine channel attention fusion feature is input into the coarse spatial attention branch (s_rough), and the coarse spatial attention branch encodes the feature into four spatial weighted features, and the coarse spatial attention feature size is 160×256×256. The shared spatial attention module consists of global maximum pooling, global average pooling and 7*7 convolution. The two are concatenated and added, and then input into the spatial attention fusion module (fusion_output) to obtain the final dynamic hybrid fusion feature size of 160×256×256.
[0099] The radar-camera feature interactive attention module proposed in the present invention ensures the effective interaction between the radar BEV and the camera BEV by introducing self-attention and deformable attention, thereby realizing effective supervision and guidance of the camera BEV features, while enhancing the camera BEV features while fully retaining the camera's semantic information, enhancing the semantic expression ability of the camera BEV features, and improving the performance degradation problem of the camera BEV features due to the lack of explicit deep supervision.
[0100] Table 10 fusion_input&c_refined&c_rough&fusion_channel of dynamic hybrid fusion module
[0101]
[0102] Table 11 s_refined&s_rough&fusion_output of dynamic hybrid fusion module
[0103]
[0104]
[0105] The heat map generator is used to extract the dynamic mixed fusion features to obtain a dense heat map. The specific convolution operation of the heat map generator is shown in Table 12. First, the dynamic mixed fusion features are input to the dense heat map encoding module to output intermediate features of size 128×256×256. The feature is then input to the heat map generation module and the output feature size is 10×256×256. Finally, the final dense heat map is obtained after the sigmoid activation function. The size of the obtained dense heat map is {10, H, W}, where 10 represents the number of categories, and H and W are the height and width of the feature, respectively.
[0106] Table 12 Heatmap generator
[0107]
[0108] The decoupled detection head is used to perform classification prediction based on the dense heat map and the dynamic hybrid fusion feature to obtain a category prediction result, and to perform regression prediction based on the dense heat map and the time-series fusion radar BEV feature to obtain a regression prediction result. The category prediction result and the regression prediction result together constitute the target detection result. This module selects the first K, for example, 200 candidate objects from the dense heat map generated in the previous stage, and then samples the time-series fused radar BEV features and the dynamic hybrid fusion features respectively. According to the features collected by the candidate objects in the time-series fused radar BEV features, a feedforward network is used to predict the final 5 regression targets: the center feature size is 2×200, the height feature size is 1×200, the dim feature size is 3×200, the rot feature size is 2×200, and the vel feature size is 2×200. According to the features collected by the candidate objects in the dynamic hybrid fusion features, a feedforward network is used to predict the final category prediction: the heatmap feature size is 10×200. The final 6 results together constitute the prediction result of the BEV multimodal target detection network.
[0109] The system proposed in the present invention can be trained with the following settings: the network is trained from scratch in an end-to-end manner using the AdamW optimizer. The number of training iterations is set to 20, the initial learning rate is 0.0001, the weight-decay is set to 0.01, and the moment parameter is set to 0.9. For the nuScenes dataset, the original point cloud is cropped to a space of [-51.2m, 51.2m]×[-51.2m, 51.2m]×[-3m, 5m] along the X, Y, and Z axes, and the image is preprocessed to a size of 384×1056. The voxel size is (0.05, 0.05, 0.2)m. The experimental environment is a server with two GeForce RTX 4090s (CUDA 11.8, CUDNN8.9.5). All experiments were conducted based on Pytoch 2.1.0 and mmdet3d 1.0.0rc4.
[0110] An embodiment of the present invention also provides an autonomous driving target detection method, the method comprising: acquiring surround view image data of a vehicle, and extracting camera BEV features based on the surround view image data; acquiring radar point cloud data of the vehicle, and extracting radar BEV features based on the radar point cloud data; extracting time-series fusion camera BEV features based on the camera BEV features, and extracting time-series fusion radar BEV features based on the radar BEV features; interacting the time-series fusion camera BEV features with the time-series fusion radar BEV features to obtain interacted camera BEV features; and performing regression prediction and classification prediction based on the interacted camera BEV features and the time-series fusion radar BEV features to obtain autonomous driving target detection results, wherein the surround view image data includes historical frame images and current frame images, and the radar point cloud data includes historical frame radar point clouds and current frame radar point clouds.
[0111] The specific working principle and benefits of the autonomous driving target detection method provided by the embodiment of the present invention are similar to the specific working principle and benefits of the autonomous driving target detection system provided by the embodiment of the present invention, and will not be repeated here.
[0112] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0116] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0117] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0118] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0119] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0120] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. An automatic driving target detection system, characterized in that: The system comprises: A camera network branch, for acquiring surround view image data of the vehicle and extracting camera BEV features based on the surround view image data; A radar network branch, for acquiring radar point cloud data of the vehicle, and extracting radar BEV features based on the radar point cloud data; A historical feature time series fusion module, used for extracting and fusing camera BEV features based on the camera BEV features, and extracting and fusing radar BEV features based on the radar BEV features; A radar-camera feature interactive attention module, configured to interact the temporal fusion camera BEV feature with the temporal fusion radar BEV feature to obtain an interacted camera BEV feature; and The fusion detection head network is used to perform regression prediction and classification prediction based on the interacted camera BEV features and the time series fusion radar BEV features to obtain the autonomous driving target detection result. The surround view image data includes a historical frame image and a current frame image, and the radar point cloud data includes a historical frame radar point cloud and a current frame radar point cloud.
2. The system according to claim 1, characterized in that The camera network branch includes: A first acquisition module, used to acquire surround view image data of the vehicle and depth information of the radar; An image splitting module, used for splitting the surround view image data to obtain an image list, each element of the image list contains an image of one frame; An image front view feature extraction network, used for extracting multi-scale front view features for each element of the image list; A feature pyramid network is used to perform multi-scale fusion on the multi-scale front view features corresponding to each of the elements to obtain multi-scale fused image front view features corresponding to each of the elements; and A view conversion module is used to obtain the camera BEV feature based on the multi-scale fusion image front view feature corresponding to each element and the depth information of the radar.
3. The system according to claim 2, characterized in that The camera network branch also includes: an auxiliary classification head, used to calculate the auxiliary classification head loss based on the front view features of the multi-scale fused image, The auxiliary classification head loss is used to update the camera network branch when training the system.
4. The system according to claim 1, characterized in that The radar network branches include: A second acquisition module, used to acquire radar point cloud data of the vehicle; A voxel grid partitioning module, used for partitioning the 3D space of the radar point cloud data into voxels of a preset size; A grouping operation module, used for distributing the point cloud into various voxels in sequence; A random sampling module, used to determine whether the number of point clouds of each voxel exceeds a preset first number, and if the number of point clouds of a certain voxel exceeds the first number, reduce the number of point clouds of the voxel to the first number through random sampling; and The voxel feature encoding module is used to encode, extract features and perform multi-scale fusion on the voxels to obtain radar BEV features.
5. The system according to claim 1, characterized in that The historical feature time series fusion module includes: A short-term time series fusion module is used to perform time series fusion on the BEV features of two adjacent frames to obtain short-term time series fusion features; and The long-term time series fusion module is used to perform time series fusion on the short-term time series fusion features to obtain long-term and short-term time series fusion features. Among them, the BEV feature is a camera BEV feature or a radar BEV feature. When the BEV feature is a camera BEV feature, the long-term and short-term time series fusion feature is a time series fusion camera BEV feature. When the BEV feature is a radar BEV feature, the long-term and short-term time series fusion feature is a time series fusion radar BEV feature.
6. The system according to claim 5, characterized in that The BEV features of each two adjacent frames are temporally fused to obtain short-term temporal fusion features including: Aligning the BEV features of two adjacent frames in time and space; and The BEV features of each two adjacent frames after spatiotemporal alignment are fused to obtain short-term temporal fusion features.
7. The system according to claim 1, characterized in that The radar-camera feature interaction attention module includes: A self-attention enhancement module, used to enhance the BEV feature of the temporal fusion camera and perform a first processing to obtain a query vector; A deformable attention module, configured to use the time series fusion radar BEV feature as a key vector and a numerical vector to interact with the query vector to obtain an interacted BEV feature; and The interactive feature output module is used to sequentially perform the first processing, forward propagation and the first processing on the BEV feature after the interaction to obtain the camera BEV feature after the interaction. The first processing includes addition and normalization.
8. The system according to claim 1, characterized in that The fusion detection head network includes: A dynamic hybrid feature fusion module, used for dynamically hybridizing and fusing the interacted camera BEV feature and the time series fusion radar BEV feature to obtain a dynamic hybrid fusion feature; A heat map generator, used for extracting features from the dynamic hybrid fusion features to obtain a dense heat map; and The decoupled detection head is used to perform classification prediction based on the dense heat map and the dynamic hybrid fusion feature to obtain a category prediction result, and to perform regression prediction based on the dense heat map and the time series fusion radar BEV feature to obtain a regression prediction result. The category prediction result and the regression prediction result together constitute the target detection result.
9. The system according to claim 8, characterized in that The interactive camera BEV feature and the time series fusion radar BEV feature are dynamically mixed and fused to obtain the dynamic mixed fusion feature including: Preliminarily fusing the interacted camera BEV feature and the time-series fusion radar BEV feature to obtain a preliminary fusion feature; Sending the preliminary fusion features to the fine channel attention branch and the coarse channel attention branch respectively to obtain fine channel attention features and coarse channel attention features; Merging the fine channel attention feature and the coarse channel attention feature into a coarse and fine channel attention fusion feature, and sending them to the fine spatial attention branch and the coarse spatial attention branch respectively to obtain a fine spatial attention feature and a coarse spatial attention feature; and The fine spatial attention feature and the coarse spatial attention feature are fused to obtain the dynamic mixed fusion feature.
10. An automatic driving target detection method, characterized in that: The method comprises: Acquire surround view image data of the vehicle, and extract camera BEV features based on the surround view image data; Acquire radar point cloud data of the vehicle, and extract radar BEV features based on the radar point cloud data; Extracting and fusing camera BEV features in a time series based on the camera BEV features, and extracting and fusing radar BEV features in a time series based on the radar BEV features; Interacting the time-series fusion camera BEV feature with the time-series fusion radar BEV feature to obtain an interacted camera BEV feature; and Based on the interactive camera BEV features and the time series fusion radar BEV features, regression prediction and classification prediction are performed to obtain the autonomous driving target detection result. The surround view image data includes a historical frame image and a current frame image, and the radar point cloud data includes a historical frame radar point cloud and a current frame radar point cloud.