Three-dimensional vehicle detection method fusing dual-scale channel attention and depth estimation enhancement

By integrating the dual-scale channel attention and depth estimation enhancement method, the existing three-dimensional vehicle detection method has solved the problem of performance degradation in complex scenarios, achieving higher detection accuracy and robustness, and meeting the high safety requirements of autonomous driving.

CN120164176APending Publication Date: 2025-06-17GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510253112.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing three-dimensional vehicle detection methods have degraded performance in complex scenarios, which is difficult to meet the high safety requirements of autonomous driving, and a single sensor cannot fully capture environmental information, resulting in insufficient detection accuracy and robustness.

Method used

The three-dimensional vehicle detection method that combines the attention and depth estimation enhancement of the dual-scale channel, enhances multimodal features through the dual-scale channel attention module, improves the accuracy of the depth information, and uses the Transformer encoder-decoder to generate candidate target boxes, and optimizes the model training through the improved loss function.

Benefits of technology

It improves the accuracy and robustness of three-dimensional vehicle detection, especially in complex scenarios, meets the high safety requirements of autonomous driving, and realizes efficient fusion of multimodal data, reduces computing costs, and meets the real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164176A_ABST
    Figure CN120164176A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional vehicle detection method fusing dual-scale channel attention and depth estimation enhancement, and aims to improve the detection precision and real-time performance of a three-dimensional detection algorithm. The method comprises the following steps: S1, obtaining original point cloud data and an RGB image from a radar and a camera sensor, and carrying out denoising and filtering on the point cloud data; s2, extracting 2D and 3D features, and enhancing the features through a dual-scale channel attention module; s3, improving the depth information precision by using a depth estimation enhancement module; s4, the multi-modal features are input into a Transform encoder-decoder, and a candidate target frame is generated; s5, optimizing model training through an improved loss function; and S6, training the model, outputting a three-dimensional vehicle detection result, realizing deployment, and visualizing the actual application effect of the method. The method is mainly used for realizing efficient and accurate three-dimensional vehicle detection, can accurately identify the position of the vehicle in a complex scene, and provides key technical support for applications such as automatic driving and intelligent traffic systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vehicle detection in computer vision, and particularly to a three-dimensional vehicle detection method that fuses dual-scale channel attention and enhanced depth estimation. Background Art

[0002] Three-dimensional vehicle detection is one of the core technologies in autonomous driving and intelligent transportation systems. Its goal is to accurately identify and locate the three-dimensional position, size, and orientation of vehicles from sensor data (such as LiDAR point clouds and RGB images). However, existing three-dimensional vehicle detection methods still face many challenges in practical applications and urgently need improvement and optimization. Autonomous vehicles need to perceive the surrounding environment in real time and accurately, especially the positions and motion states of other vehicles. Existing detection methods experience a decline in performance in complex scenarios (such as occlusion and lighting changes) and are difficult to meet the high safety requirements of autonomous driving. In addition, a single sensor (such as LiDAR or camera) has limitations and cannot comprehensively capture environmental information. Therefore, it is necessary to fuse multi-modal data (LiDAR point clouds and RGB images) to improve detection accuracy and robustness.

[0003] The significance of this research is to provide more efficient and reliable technical support for autonomous driving, intelligent transportation systems, and industrial applications. In the field of autonomous driving, the ability of high-precision three-dimensional vehicle detection can help vehicles perceive the surrounding environment in real time and improve the safety and reliability of autonomous driving. In intelligent transportation systems, real-time monitoring of the vehicle distribution and motion states on the road can optimize traffic flow management, detect traffic violations, and improve road safety. In industrial applications, detecting vehicles and goods in warehouses or factories can optimize logistics management and improve the efficiency and safety of industrial automation. Summary of the Invention

[0004] The present invention provides a three-dimensional vehicle detection method that fuses dual-scale channel attention and enhanced depth estimation, aiming to solve the problem that traditional methods in the prior art usually directly use raw point cloud or image features and lack effective modeling of multi-scale information and inter-channel relationships, resulting in a decline in detection performance in complex scenarios. The present invention adopts the following technical solutions:

[0005] A three-dimensional vehicle detection method that fuses dual-scale channel attention and enhanced depth estimation, characterized by including the following steps:

[0006] Step S1: Obtain the raw point cloud data and RGB images from radar and camera sensors, and denoise and filter the point cloud data.

[0007] Step S2: Extract 2D and 3D features, and enhance the features through a dual-scale channel attention module.

[0008] Step S3: Use the depth estimation enhancement module to improve the accuracy of depth information;

[0009] Step S4: Input the multi-modal features into the Transformer encoder-decoder to generate candidate target boxes;

[0010] Step S5: Optimize the model training through an improved loss function;

[0011] Step S6: Output the 3D vehicle detection results, implement the deployment, and visualize the actual application effect of this method.

[0012] Furthermore, in the step S1, obtaining the original point cloud data and performing preprocessing includes the following steps:

[0013] Generate point cloud data and RGB images through the radar and camera, align the point cloud with the RGB image, remove the outlier points from the collected data through statistical filtering or radius filtering; use RANSAC or height filtering to remove the ground points; save the processed point cloud data.

[0014] Furthermore, in the step S2, designing the dual-scale channel attention module includes the following steps:

[0015] Step S2.1: In the global channel attention block, use the average pooling layer and the max pooling layer to reduce the computational cost. In order to aggregate the global context across channels, fuse and reduce the dimension of the input feature map to 2C / R through a 1×1 convolutional layer, where C is the dimension of the input x∈R C×H×W and r is the reduction rate. Then, this block generates the global channel attention vector g(x)∈R through another 1×1 convolutional layer C×1×1

[0016] Step S2.2: In the local channel attention block, place two 1×1 convolutional layers to extract the cross-channel local information and generate the local channel attention map l(x)∈R C×H×W . Before calibration, fuse the local channel attention map and the global channel attention vector into the channel attention map A C (x)∈R C×H×W as shown, and the specific formula is:

[0017]

[0018] Step S2.3: Further calibrate the original channel attention map to improve the feature representation, and calibrate the new channel-attention map

[0019] Furthermore, in the step S3, designing the depth estimation enhancement module includes the following steps:

[0020] Step S3.1: Input the feature map F ∈ R C×H×W Predict the discrete depth probability D ∈ R D×H×W through two convolutional layers, where D is the number of depth categories. The depth value of each pixel is discretized into D categories, and the probability represents the confidence that the depth value of the pixel belongs to a certain category.

[0021] Step S3.2: Introduce the central representation (depth prototype) of depth categories to enhance features. The depth prototype F d is calculated by aggregating the depth-aware features of each pixel belonging to the specified category. The specific formula is:

[0022]

[0023] where X i ' is the feature of the i-th pixel, Ω is the set of pixels in the feature map, and F di is the normalized probability of the d-th depth category.

[0024] Step S3.3: Reconstruct the new depth-aware feature F′ based on the depth prototype representation, concatenate the initial depth-aware feature X with the reconstructed feature F′, and obtain the enhanced depth feature through a 1×1 convolutional layer. The specific formula is:

[0025]

[0026] Step S3.4: Use group convolution to reduce the number of depth categories from D to D′ = D / r, where r is the reduction ratio. Group convolution can share similar depth cues, reduce the computational amount, and at the same time retain the expressive ability of depth features.

[0027] Furthermore, in step S4, utilize the global modeling ability of the Transformer to capture the long-range dependencies between targets. Input the multi-modal features into the Transformer encoder, and extract the global context information through the self-attention mechanism. Use the Transformer decoder to generate candidate target boxes (Proposals) and their corresponding features. Perform non-maximum suppression (NMS) on the candidate target boxes to remove redundant boxes.

[0028] Furthermore, in step S5, optimize the model training through the loss function to improve the detection accuracy. Solve the class imbalance problem by using the Focal Loss. Use the Smooth L1 Loss to optimize the position and size of the target boxes. Introduce the depth consistency loss to constrain the consistency between the point cloud depth and the predicted depth. Weight and sum the classification loss, regression loss, and depth consistency loss as the final loss function.

[0029] Furthermore, in step S6, it includes the following steps:

[0030] Step S6.1: Train the dataset to obtain an efficient detection model after training with the improved method; and output 3D vehicle detection results, including the target category, 3D bounding box (center point, size, orientation angle), and confidence score.

[0031] Step S6.2: Deploy this method on the vehicle monitoring system, visualize the detection results of the vehicle, and apply this vehicle detection method to actual traffic.

[0032] Compared with the prior art, the present invention has the following advantages and beneficial effects: By introducing a dual-scale channel attention module, the expression ability of multi-modal features is enhanced, and the detection accuracy is improved. Through the depth estimation enhancement module, the depth estimation problem is transformed into a sequential classification problem, the discrete depth probability distribution is predicted, and the depth prototype is used to enhance the depth perception features, improving the depth estimation accuracy. At the same time, an efficient multi-modal fusion strategy is designed to reduce the computational cost and meet the real-time requirement in practical applications. Brief Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art.

[0034] Figure 1 It is a flow diagram of steps S1 to S6 in the present invention.

[0035] Figure 2 It is the network structure of the dual-scale channel attention module in step S3 of the present invention.

[0036] Figure 3 It is the network structure of the depth estimation enhancement module in step S4 of the present invention. Detailed Embodiments

[0037] The following will describe the present invention in detail with specific embodiments.

[0038] As Figure 1 shown, a three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement, characterized by including the following steps:

[0039] Step S1: Obtain the original point cloud data and RGB images from radar and camera sensors, and denoise and filter the point cloud data;

[0040] Step S2: Extract 2D and 3D features, and enhance the features through a dual-scale channel attention module;

[0041] Step S3: Use the depth estimation enhancement module to improve the accuracy of depth information;

[0042] Step S4: Input the multi-modal features into the Transformer encoder-decoder to generate candidate target boxes;

[0043] Step S5: Optimize the model training through an improved loss function;

[0044] Step S6: Output the 3D vehicle detection results, implement the deployment, and visualize the actual application effect of this method.

[0045] Further, in the step S1, obtaining the original point cloud data and performing preprocessing includes the following steps:

[0046] Generate point cloud data and RGB images through radar and cameras, align the point cloud with the RGB image, and remove outliers from the collected data through statistical filtering or radius filtering; use RANSAC or height filtering to remove ground points; save the processed point cloud data.

[0047] Further, in the step S2, as Figure 2 shown, design a dual-scale channel attention module, including the following steps:

[0048] Step S2.1: In the global channel attention block, use average pooling layer and max pooling layer to reduce the computational cost. To aggregate the global context across channels, fuse and reduce the dimension of the input feature map to 2C / R through a 1×1 convolutional layer, where C is the dimension of the input x∈R C×H×W and r is the reduction rate. Then, this block generates a global channel attention vector g(x)∈R C×1×1

[0049] Step S2.2: In the local channel attention block, place two 1×1 convolutional layers to extract cross-channel local information and generate a local channel attention map l(x)∈R C×H×W . Before calibration, fuse the local channel attention map and the global channel attention vector into the channel attention map A C (x)∈R C×H×W shown, and the specific formula is:

[0050]

[0051] Step S2.3: Further calibrate the original channel attention map to improve the feature representation, and calibrate the new channel-attention map

[0052] Further, in the step S3, as Figure 3 shown, design a depth estimation enhancement module, including the following steps:

[0053] Step S3.1: Input feature map F∈RC×H×W Predict the discrete depth probability \(D\in\mathbb{R}\) through two convolutional layers D×H×W , where \(D\) is the number of depth categories. The depth value of each pixel is discretized into \(D\) categories, and the probability represents the confidence that the depth value of the pixel belongs to a certain category.

[0054] Step S3.2: Introduce the central representation (depth prototype) of depth categories to enhance features. The depth prototype \(F\) d is calculated by aggregating the depth-aware features of each pixel belonging to the specified category. The specific formula is as follows:

[0055]

[0056] where \(X\)' i is the feature of the \(i\)-th pixel, \(\Omega\) is the set of pixels in the feature map, and \(F\) di is the normalized probability of the \(d\)-th depth category.

[0057] Step S3.3: Reconstruct the new depth-aware feature \(F'\) based on the depth prototype representation, connect the initial depth-aware feature \(X\) with the reconstructed feature \(F'\), and obtain the enhanced depth feature through a \(1\times1\) convolutional layer. The specific formula is as follows:

[0058]

[0059] Step S3.4: Use grouped convolution to reduce the number of depth categories from \(D\) to \(D' = D / r\), where \(r\) is the reduction ratio. Grouped convolution can share similar depth cues, reduce the computational amount, and at the same time retain the expressive ability of depth features.

[0060] Furthermore, in step S4, utilize the global modeling ability of the Transformer to capture the long-range dependencies between targets. Input the multi-modal features into the Transformer encoder, and extract the global context information through the self-attention mechanism. Use the Transformer decoder to generate candidate target boxes (Proposals) and their corresponding features. Perform non-maximum suppression (NMS) on the candidate target boxes to remove redundant boxes.

[0061] Furthermore, in step S5, optimize the model training through the loss function to improve the detection accuracy. Solve the class imbalance problem by using the Focal Loss. Use the Smooth L1 Loss to optimize the position and size of the target boxes. Introduce the depth consistency loss to constrain the consistency between the point cloud depth and the predicted depth. Weightedly sum the classification loss, regression loss, and depth consistency loss as the final loss function.

[0062] Furthermore, in step S6, it includes the following steps:

[0063] Step S6.1: Train the dataset to obtain an efficient detection model after training with the improved method; and output 3D vehicle detection results, including the target category, 3D bounding box (center point, size, orientation angle), and confidence score.

[0064] Step S6.2: Deploy this method on the vehicle monitoring system, visualize the detection results of the vehicle, and apply this vehicle detection method to actual traffic.

[0065] The above has described the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and these should also be regarded as the protection scope of the present invention.

Claims

1. A three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement, characterized in that , including the following steps: Step S1: De-noising and filtering the original point cloud data and RGB images obtained from the radar and camera sensors; Step S2: extract 2D and 3D features and enhance the features through a dual-scale channel attention module; Step S3: using a depth estimation enhancement module to improve the accuracy of depth information; Step S4: Input the multimodal features into the Transformer encoder-decoder to generate candidate target boxes; Step S5: Optimizing model training by improving the loss function; Step S6: Output the three-dimensional vehicle detection results and implement deployment to visualize the actual application effect of the method.

2. A three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: In step S1, the original point cloud data is obtained and preprocessed, including the following steps: Generate point cloud data and RGB images through radar and camera, align the point cloud with the RGB image, and remove outliers through statistical filtering or radius filtering on the collected data; Use RANSAC or height filtering to remove ground points; save the processed point cloud data.

3. The three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: In step S2, a dual-scale channel attention module is designed, comprising the following steps: Step S2.1: In the global channel attention block, average pooling layer and max pooling layer are used to reduce computational cost. In order to aggregate global context across channels, the dimension of the input feature map is fused and reduced to 2C / R through a 1×1 convolutional layer, where C is the input x∈R C×H×W The block then passes through another 1×1 convolutional layer to produce a global channel attention vector g(x)∈R C×1×1 Step S2.2: In the local channel attention block, two 1×1 convolutional layers are placed to extract local information across channels and generate a local channel attention map l(x)∈R C×H×W Before calibration, the local channel attention map and the global channel attention vector are fused into the channel attention map A C (x)∈R C×H×W As shown, the specific formula is: Step S2.3: The original channel attention map is further calibrated to improve feature representation, and the new channel-attention map is calibrated .

4. The three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: In step S3, a depth estimation enhancement module is designed, which includes the following steps: Step S3.1: Input feature map F∈R C×H×W Predict discrete depth probabilities D∈R through two convolutional layers D×H×W , where D is the number of depth categories. The depth value of each pixel is discretized into D categories, and the probability represents the confidence that the pixel depth value belongs to a certain category. Step S3.2: Introduce the central representation of the deep category (deep prototype) to enhance the features. Deep prototype F d It is calculated by aggregating the depth perception features of each pixel belonging to the specified category. The specific formula is: Where X i ' is the feature of the i-th pixel, Ω is the set of pixels in the feature map, F di is the normalized probability of the d-th depth class. Step S3.3: Reconstruct a new depth-aware feature F′ based on the deep prototype representation, connect the initial depth-aware feature X with the reconstructed feature F′, and obtain the enhanced depth feature through a 1×1 convolutional layer. The specific formula is: Step S3.4: Use group convolution to reduce the number of depth categories from D to D′=D / r, where r is the reduction ratio. Group convolution can share similar depth clues, reduce the amount of computation, and retain the expressive power of deep features.

5. The three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: In step S4, the global modeling capability of Transformer is used to capture the long-distance dependencies between objects, and the multimodal features are input into the Transformer encoder to extract the global context information through the self-attention mechanism. The Transformer decoder is used to generate candidate object boxes (Proposals) and their corresponding features. Perform non-maximum suppression (NMS) on the candidate target boxes to remove redundant boxes.

6. The three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: In step S5, the model training is optimized by the loss function to improve the detection accuracy. The class imbalance problem is solved by using Focal Loss. The position and size of the target box are optimized by using Smooth L1 Loss. The depth consistency loss is introduced to constrain the consistency of the point cloud depth and the predicted depth. The classification loss, regression loss and depth consistency loss are weighted and summed as the final loss function.

7. The three-dimensional vehicle detection method integrating dual-scale channel attention and depth estimation enhancement according to claim 1, characterized in that: The step S6 includes the following steps: Step S6.1: Train the data set to obtain a trained efficient detection model of the improved method; and output the three-dimensional vehicle detection results, including the target category, 3D bounding box (center point, size, direction angle) and confidence score. Step S6.2: Deploy the method on a vehicle monitoring system, visualize the vehicle detection results, and apply the vehicle detection method to actual traffic.

Citation Information

Cited By

  • Real-time vehicle collision prediction method based on multi-modal depth fusion and time sequence modeling

    CN121564670A