Method for detecting three-dimensional vehicle in road scene based on multi-modal fusion network
By using a multimodal fusion network method in highway scenarios, fusing the data of lidar and surveillance cameras, the MMFN-PVA-VDHS network model is constructed, which solves the problem that a single data source is difficult to accurately detect vehicles under harsh conditions, and achieves high-precision three-dimensional vehicle detection.
Patent Information
- Application Number
- CN202510009902.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
In highway scenarios, a single data source such as lidar and surveillance cameras are difficult to provide sufficient information to accurately detect vehicles in severe weather and severe light changes, resulting in a decrease in vehicle detection accuracy.
Using a multimodal fusion network method, the MMFN-PVA-VDHS network model is constructed, and three-dimensional vehicle detection is performed by fusing the data of the lidar and the monitoring camera, combining the point-voxel attention multimodal fusion network. The model includes cross-modal GT enhancement, cross-modal multi-head attention and point-voxel global alignment modules, which can effectively process multi-modal data and improve vehicle detection accuracy.
It effectively improves the accuracy of three-dimensional vehicle detection in highway scenarios, and can accurately detect vehicles under vibration and severe weather conditions, providing a basis for subsequent vehicle hazard behavior analysis.
Smart Images

Figure CN119942472A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the research field of smart highways and smart perception, and specifically relates to a three-dimensional vehicle detection method in highway scenes based on a multimodal fusion network. Background Art
[0002] Vehicle detection is crucial for highway scene perception. Highways are characterized by high vehicle speeds, dramatic lighting changes, and harsh weather conditions. LiDAR and surveillance cameras, the primary data sources for perception, significantly degrade LiDAR performance in adverse weather conditions like rain, snow, and fog, and their high cost limits their widespread adoption. Surveillance cameras experience image quality and recognition accuracy degradation when lighting conditions fluctuate dramatically, and their field of view is easily obstructed by vehicles ahead. A single data source often fails to provide sufficient information for accurate vehicle detection. Therefore, integrating multimodal perception information from LiDAR and surveillance cameras to improve vehicle detection performance is crucial. Summary of the Invention
[0003] Purpose of the invention: In order to overcome the shortcomings of the existing technology, a three-dimensional vehicle detection method in highway scenes based on a multimodal fusion network is provided, and an MMFN-PVA-VDHS network model for vehicle tracking in highway scenes is constructed. After model training, parameter optimization and model comparison, it can effectively perform three-dimensional detection of vehicles in vibration conditions in highway scenes, improve the three-dimensional detection accuracy of vehicles, and provide a basis for subsequent analysis of dangerous vehicle behaviors.
[0004] Technical solution: To achieve the above objectives, the present invention provides a three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network, comprising the following steps:
[0005] S1: Construct a 3D vehicle detection dataset for highway scenarios, including complex traffic conditions such as intersections, lane markings, and traffic lights, congestion scenarios, and different weather conditions (e.g., sunny, rainy, and nighttime). Specifically, it includes gantry vibration in highway scenarios affected by alignment.
[0006] S2: Based on the characteristics of the vehicle detection dataset, a MMFN-PVA-VDHS (Vehicle Detection in Highway Scenarios Based on Point-Voxel Attention Multimodal Fusion Network) network model is constructed for highway vehicle detection.
[0007] S3: Train the MMFN-PVA-VDHS network model and optimize its parameters, and compare it with other detection models for highway vehicle detection.
[0008] Furthermore, the specific steps of step S1 are as follows:
[0009] S1-1: Construct a 3D vehicle detection dataset in highway scenes using a mixture of Nuscenes and Nuscenes-C. The training set includes 750 scenes and 28,730 frames of data, and the validation set includes 150 scenes and 6,019 frames of data.
[0010] Furthermore, the specific steps of step S2 are as follows:
[0011] S2-1: We built a Cross-Modal GT-AUG module that distinguishes foreground and background, and simultaneously considers instance depth and occlusion information, to more realistically simulate the cross-modal data augmentation process in real-world scenarios.
[0012] S2-2: We built a cross-modal multi-head attention module for images and radar bird’s-eye view images, effectively compensating for the fine-grained information in high-resolution features while enhancing the information representation in the image.
[0013] S2-3: Constructed a Point-Voxel fusion with Global-Align module, combining projection hard connections with attention soft connections to aggregate point and voxel features.
[0014] Furthermore, in step S2-1, the instance object used for enhancement in the Cross Modal Multi-Head-Attention module does not have spatial collision with the instance object in the current scene, and no complex point cloud 3D perspective filtering operation and image occlusion mask operation are performed during the pasting of the point cloud and image instances; during the image fusion process, for instance objects whose foreground and background of the fusion area correspond to different scale values, and whose depth values are large and occluded multiple times, the fusion pixel method in PuzzleMix is used, and the transparency will become smaller and smaller. Considering that the instance object with a larger depth value occupies fewer pixels, the instance object is pixel compensated.
[0015] In step S2-2, in order to further associate the feature relationship between the point cloud and the image, and taking advantage of the high resolution of the image for detecting small targets and the spatial position of the point cloud detection, a radar bird's-eye view (Lidar BEV) guided query initialization strategy is adopted, and the radar bird's-eye view information is used to select the object query.
[0016] In step S2-3, the process of combining point-image and voxel-image feature aggregation is carried out. If the cross-modal hard projection association method is completely abandoned and the adaptive soft connection based on the cross-attention mechanism is adopted, the computational cost is too high. Therefore, the Point-Voxel Fusion with Global-Align Module adopts hard projection and soft connection methods for point-image and voxel-image respectively, overcoming the misalignment of hard connection features while reducing the computational cost.
[0017] The voxel-image fusion features are computed using both hard projection and soft connection methods. Deformable Cross-Attention is used to establish feature associations between voxels and images. Unlike operations that associate voxels with all pixels in an image, deformable operations selectively determine the keypoint regions of each voxel query feature in the image plane.
[0018] Combined with V through GAM (Global-Align Module) t and V s The Lidar-Camera features are self-learned and globally aligned, and PVF (Point-Voxel Fusion) fuses Point-Image and Voxel-Image features.
[0019] Furthermore, the specific steps of step S3 are as follows:
[0020] S3-1: Analyze the effectiveness of the improved strategies in the MMFN-PVA-VDHS model, compare the indicators of the AutoAlignV2 and PointAugmenting models, and comprehensively compare multimodal vehicle detection methods.
[0021] The beneficial effects of the present invention are: it integrates multimodal detection algorithms and improvement strategies for corresponding scenario problems, constructs an MMFN-PVA-VDHS network model for three-dimensional detection of vehicles in highway scenarios, can effectively perform three-dimensional detection of vehicles in highway vibration scenarios, improve vehicle detection accuracy, and provide a basis for subsequent analysis of vehicle dangerous behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Partial visualization of the highway vibration dataset.
[0023] Figure 2 This is the MMFN-PVA-VDHS model structure.
[0024] Figure 3This is the Cross Modal GT-AUG example effect picture.
[0025] Figure 4 This is the Cross Modal Multi-Head-Attention flow chart.
[0026] Figure 5 This is the flowchart of Point-Voxel fusion with Global-Align.
[0027] Figure 6 The detection effect diagram of MMFN-PVA-VDHS in the test set, where (a) is scene 1, (b) is scene 2, (c) is scene 3, and (d) is scene 4. DETAILED DESCRIPTION
[0028] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.
[0029] The present invention provides a three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network, which specifically includes the following steps:
[0030] S1: We use a mixture of Nuscenes and Nuscenes-C as the experimental dataset, including complex traffic conditions such as intersections, lane markings, and traffic lights, congestion scenarios, and different weather conditions (e.g., sunny, rainy, and nighttime). We specifically include gantry vibration in highway scenarios affected by alignment. The training set includes 750 scenes and 28,730 frames of data, and the validation set includes 150 scenes and 6,019 frames of data. Partial visualization of the scene data (CAM_FRONT and LIDAR_TOP) is available. Some visualizations of the vibration scene dataset are shown in the figure below. Figure 1 shown.
[0031] S2: Based on the characteristics of the dataset, a MMFN-PVA-VDHS network model is constructed for highway scene vehicle detection. The structure is as follows Figure 2 shown.
[0032] S2-1: Using the Cross Modal GT-AUG method, the filtered point cloud instance object set P is known 3D , the corresponding image instance object set P 2D , the corresponding instance depth set D, the current scene point cloud P and image I to be enhanced, the corresponding point cloud real set and image reality Corresponding to the real depth set D ′ Specific steps for the Cross Modal GT-AUG method:
[0033] (1) According to the depth value of D, the P3D and P 2D Sort.
[0034] (2) Traverse the instance objects in order according to the sorting, initialize the number of times the current instance object is blocked n = 0, and for the current scene point cloud P to be enhanced, Paste into P, where i is the index of the instance object sorted by depth value.
[0035] (3) For the current scene image I to be enhanced, construct correspond Mask, according to D ′ Depth value from large to small traverses the corresponding calculate and The IOU value, if the IOU value is greater than the threshold Then n=n+1, and combined with Calculate the pixel superposition of the overlapping area and update Where j is the index of GT sorted by depth value, and the calculation formula is as follows:
[0036]
[0037] (4) Calculate the pixels in the non-overlapping area, where α f and α b They represent the foreground and background fusion ratios respectively, and the calculation formula is as follows:
[0038]
[0039] (5) Get the final image I Figure 3 This is the visualization of the Cross Modal GT-AUG example operation. Here we take α f =0.2,α b =0.6,
[0040] S2-2: The Cross Modal Multi-Head-Attention module adopts a Lidar BEV-guided query initialization strategy and uses Lidar BEV information to select object queries.
[0041] Specifically, for the Lidar BEV pseudo image, the dimension is scaled from 1024×1024×3 to 640×640×3, and multi-scale features are obtained through the feature extraction network, which are 8 times down-sampled to 80×80×64, 16 times down-sampled to 40×40×64, and 32 times down-sampled to 20×20×64. In order to ensure the overall parameter amount of the fusion model, ResNet18 is selected as the feature extraction network; for the camera image, multi-scale image features are obtained through CSPDarknet, which are 8 times down-sampled to 80×144×64, 16 times down-sampled to 40×72×64, and 32 times down-sampled to 20×36×64. By performing the image feature F C BEV with Lidar Features F L Perform cross-head attention to obtain Lidar-Camera feature F LC , define the attention sequence Taking 8 times downsampling as an example, N C and N L Respectively represent F C and F L The length of the tokens of the sequence, after flattening, is 11520 and 6400 respectively. c and D L Respectively represent the dimensional information of the corresponding token. Cross Modal Attention uses linear projection to calculate query, keys and values. The calculation formula is as follows:
[0042]
[0043] in, Corresponding weight matrix, tokens output dimension D q =D k =D v .
[0044] Cross Modal Attention is extended to Nheads, and the calculation formula of Cross Modal Multi-Head-Attention is obtained:
[0045] MultiHead(Q,K,V)=Concat[head1,head2,head3,…,head N ]W O (4)
[0046] in, W O is a fully connected layer.
[0047] Taking 8 times downsampling as an example, N C and N L The sizes are 11520 and 6400 respectively. Directly performing cross attention operation will result in too many operation parameters. L The features folded along the width and height directions are used as the key value sequence of the attention mechanism, such as Figure 4 Folding along the height direction can significantly reduce the amount of computation without losing key information. Considering the operation of Lidar BEV, it is necessary to compensate for information in the height direction and fold in the width direction. At the same time, Cross Modal Multi-Head-Attention is performed on the folded features in both directions. Finally, through experiments, it is found that feature addition can achieve a good modal interaction effect, effectively compensating for the fine-grained information in the high-resolution features.
[0048] S2-3: Hard projection and soft connection are used for Point-Image and Voxel-Image respectively to overcome the misalignment of hard connection features while reducing the computational cost. The details are as follows:
[0049] (1) Calculate the Voxel-Image fusion features of hard projection and soft connection. The feature association between voxels and images is established through Deformable Cross-Attention. Unlike all pixel operations associated with voxels, Deformable operations selectively determine the key point area of each voxel query feature in the image plane. Given an image feature Non-empty voxels Voxel points can be created by hard projection With image reference point Connect to get hard projection features The calculation formula is as follows:
[0050]
[0051] Where D represents the depth of the voxel projected onto the image, ξ represents the intrinsic parameters of the camera, R and T represent the rotation and translation matrices of the lidar relative to the camera, respectively, and h represents the downsampling coefficient.
[0052] R i Perform bilinear interpolation to obtain the corresponding F i Image features, query vector Q i By F i The computational complexity of feature aggregation is reduced from O(NWH) to O(NK) by performing element-by-element dot product with voxel features. 2 ), get the soft connection feature The formula for calculating Deformable Cross-Attention is as follows:
[0053]
[0054] Among them, W m and W m ′ is a learnable weight, M corresponds to the number of detection heads (here 8), K represents the number of sampling points around the reference point (here 4), A mqk Indicates the attention score of the sampling points around the reference point obtained by MLP operation, ΔR mqk It is through Q i The predicted offset of the sampling points around the reference point.
[0055] (2) Combine V with GAM (Global-Align Module) t and V s The Lidar-Camera features are self-learned and globally aligned, and PVF (Point-Voxel Fusion) fuses Point-Image and Voxel-Image features. Figure 5 As shown, the GAM process first calculates V t and V s Perform difference operation to obtain Lidar-Camera feature offset and use it as a priori for network learning. In the Offset_Encoder stage, the sparse point cloud features are densified through submanifold sparse convolution, and then the obtained 3D features are compressed along the Z axis to form 2D BEV features. In the Dense_Encoder stage, the compressed 2D features are further extracted through maximum pooling and convolution operations, and then average pooling operations are performed. In the Dense_Linear stage, the rotation and translation parameters of the projection matrix are regressed through full connection operations. Finally, based on the initial projection matrix, the rotation matrix is regressed through the matrix multiplication @ operation, the translation matrix is regressed through the Add operation, and the updated hard projection matrix PVF process, using M l Get the Point-Image feature V after hard projection alignment t ′ , Gate_Fusion stage, V t ′ and V s Feature adjustment is performed through Linear_1 and Linear_2 respectively, and after Cat operation, Linear_3 is used to obtain the Point-Voxel fusion feature V pv .
[0056] S3: Comparison of the detection accuracy of MMFN-PVA-VDHS with other common state-of-the-art multimodal fusion methods on mixed datasets. Table 1 shows the comparison of various model metrics for training and validation on the 1 / 10 dataset, and Table 2 shows the comparison of various model metrics for training and validation on the full dataset. The comparison models in this paper include those based on Lidar-Camera and Lidar. Our MMFN-PVA-VDHS model achieves improvements on the Car, Bus, and Truck classes. Specifically, the best comparison model, the BEVFusion (MIT) model, achieves improvements of 0.94%, 3.21%, and 0.93% in AP for Car, 1.96% for Bus, and 0.96% for Truck on the 1 / 10 dataset. Improvements of 0.80%, 1.50%, and 0.68% in AP, 1.16% for mAP, and 1.09% for NDS on the full dataset. The Lidar-Camera-based detection model outperforms the Lidar-based model in these metrics.
[0057] Table 1 Performance comparison of different models on the Nuscenes and Nuscenes-c 1 / 10 validation sets
[0058]
[0059] Table 2 Performance comparison of different models on the Nuscenes and Nuscenes-c validation sets
[0060]
[0061] Comparing the detection visualization results based on Lidar-Camera and Lidar, such as Figure 6 (a)-6(d), the advantage of MMFN-PVA-VDHS compared to other methods is that it misses fewer vehicles, especially for vehicle targets with large depth values and in an occluded state. This is mainly because they contain less point cloud information and are more affected by vibration during the multimodal feature alignment process. Vehicle targets with small depth values are less affected by vibration. Even if the vibration causes some information to be misaligned, the image information can still make up for the semantic information corresponding to the point cloud in some sparse positions.
[0062] The technical means disclosed in the solutions of the present invention are not limited to those disclosed in the above-mentioned embodiments, but also include technical solutions composed of any combination of the above-mentioned technical features. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network, characterized by: The steps include: S1: Construct a 3D vehicle detection dataset in highway scenarios, including complex traffic conditions such as intersections, lane markings, and traffic lights, congestion scenarios, and different weather conditions, especially including the gantry vibration in highway scenarios affected by alignment; S2: Based on the characteristics of the vehicle detection dataset, the MMFN-PVA-VDHS vehicle detection network model for highway scene vehicle detection based on the point-voxel attention multimodal fusion network is constructed; S3: Train the MMFN-PVA-VDHS network model and optimize its parameters, and compare it with other detection models for highway scene vehicle detection.
2. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 1 is characterized in that: The step S1 is specifically as follows: S1-1: Construct a 3D vehicle detection dataset in highway scenes using a mixture of Nuscenes and Nuscenes-C. The training set includes 750 scenes and 28,730 frames of data, and the validation set includes 150 scenes and 6,019 frames of data.
3. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 1 is characterized in that: The step S2 is specifically as follows: S2-1: Construct a cross-modal GT enhancement Cross ModalGT-AUG module that distinguishes the foreground and background and considers instance depth and occlusion information at the same time, which more realistically simulates the cross-modal data enhancement process of real scenes; S2-2: Construct a cross-modal multi-head-attention module for images and radar bird's-eye views, which can compensate for the fine-grained information in high-resolution features and enhance the information representation in the image; S2-3: Construct a Point-Voxel fusion with Global-Align module, combining projection hard connections with attention soft connections to aggregate point and voxel features.
4. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 3 is characterized in that: In the step S2-1, there is no spatial collision between the instance object used for enhancement in the cross-modal multi-head attention module and the instance object in the current scene, and no complex point cloud 3D perspective filtering operation and image occlusion mask operation are performed during the pasting of the point cloud and image instances; during the image fusion process, for instance objects with different scale values corresponding to the foreground and background of the fusion area, and with large depth values and occluded multiple times, the fusion pixel method in PuzzleMix is used, and the transparency will become smaller and smaller. Considering that the instance object with a larger depth value occupies fewer pixels, pixel compensation is performed on the instance object.
5. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 3 is characterized in that: In step S2-2, in order to further associate the feature relationship between the point cloud and the image, the advantages of high-resolution image detection of small targets and point cloud detection in spatial position are utilized, and an initialization strategy of radar bird's-eye view guided query is adopted, and radar bird's-eye view information is used to select object query.
6. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 3 is characterized in that: In the step S2-3, the Point-Voxel fusion with Global-Align Module module adopts hard projection and soft connection methods for point-image and voxel-image respectively, which overcomes the misalignment of hard connection features and reduces the calculation cost.
7. The method for detecting three-dimensional vehicles in highway scenes based on a multimodal fusion network according to claim 6, characterized in that: The voxel-image fusion features of hard projection and soft connection are calculated, and the feature association between voxels and images is established through deformable cross-attention. Different from all pixel operations of voxel association image, the variable operation selectively determines the key point area of each voxel query feature in the image plane.
8. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 6 is characterized in that: Combined with V through the global alignment module t and V s Self-learning global alignment of radar-camera features, point-to-voxel fusion of point-image and voxel-to-image features.
9. The three-dimensional vehicle detection method in a highway scene based on a multimodal fusion network according to claim 1, characterized in that: The step S3 is specifically as follows: S3-1: Analyze the effect of the improved strategy in the MMFN-PVA-VDHS model, compare the indicators of the AutoAlignV2 and PointAugmenting models, and comprehensively compare multimodal vehicle detection methods.