A 3D object detection method fusing residual sub-flow manifold convolution and double attention
Patent Information
- Application Number
- CN202610698016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]然而,尽管现有的3D目标检测算法(如Voxel-RCNN)在天气晴朗、目标清晰的标准场景下取得了优异的成果,但在面对复杂和极端的自动驾驶场景时,尤其是在冰雪等恶劣天气环境下,其检测性能面临着严峻的挑战,存在以下几个等待解决的瓶颈问题:
1、本发明创造性地提出了“先净化、后增强”的特征处理范式,极大提升了模型在复杂雪天背景下的抗噪声干扰能力。 在传统的CNN网络中,雪花点云的特征响应会被盲目放大;本发明首先通过SCAC模块,利用组归一化的方差特性分离出低信息量的冗余特征并予以抑制,显著降低了计算复杂度和模型处理负担。随后串联的CBAM模块能够在无噪声干扰的纯净特征图上,精准地利用大感受野空间注意力机制凸显目标主体轮廓。这种联合设计使得模型在含有暴风雪干扰的极度恶劣场景中依然能够准确锁定目标,有效降低了误检率。
Smart Images

Figure CN122780934A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, specifically to a 3D object detection method that integrates residual submanifold convolution and dual attention. Background Technology
[0002] With the rapid development of intelligent manufacturing and autonomous driving technologies, the ability of vehicles to perceive their surroundings in high precision in three dimensions has become a core foundation for achieving high-level autonomous driving (L3 and above). Among numerous environmental perception sensors, LiDAR (Light Detection and Ranging) is increasingly widely used in 3D target detection tasks due to its ability to provide accurate depth information in all weather conditions, precise geometry, and its immunity to severe lighting conditions.
[0003] Currently, point cloud-based 3D object detection methods are mainly divided into three categories: point-based methods, voxel-based methods, and methods combining point and voxel approaches. Point-based methods (such as PointNet and PointNet++ series) directly utilize raw point cloud data for feature extraction, effectively preserving geometric structure and spatial details. However, their purely point-based architecture consumes significant computational resources, making it difficult to efficiently scale to large-scale and dense outdoor autonomous driving point cloud data. Voxel-based methods (such as VoxelNet and Voxel-RCNN series) divide disordered 3D point cloud data into regular 3D voxel or columnar meshes, and then use 3D sparse convolutional neural networks (3D Sparse CNNs) to extract local features in the regular 3D space. This method balances the fine-grained features of point clouds with the global expressive power of voxels, offering significant advantages in computational efficiency and parallel computing capabilities, thus becoming the mainstream solution in both academia and industry.
[0004] However, although existing 3D object detection algorithms (such as Voxel-RCNN) have achieved excellent results in standard scenarios with clear weather and well-defined targets, their detection performance faces severe challenges when dealing with complex and extreme autonomous driving scenarios, especially in harsh weather conditions such as ice and snow. Several bottlenecks remain to be addressed: (1) Environmental noise interference leads to the submersion of effective features: In snowy weather, the laser beam of the lidar hits the snowflakes suspended in the air and reflects, generating a large number of "snowflake noise" floating in the air. Existing models such as Voxel-RCNN, after compressing 3D voxel features into two-dimensional bird's-eye view (BEV) features along the height direction, usually use ordinary 2DCNN convolutional layers for deep feature extraction. These ordinary convolutional kernels do not have the ability to distinguish noise, so the snowflake point cloud will produce a strong activation response after being processed by the network, resulting in a large number of redundant features being extracted and infinitely amplified. This not only seriously consumes computing resources, but also causes the network to incorrectly focus on snowflake noise, covering the weak signals of real targets (especially small vehicles in the distance), ultimately leading to serious missed detections and a decrease in detection accuracy.
[0005] (2) Insufficient multi-scale feature extraction capability leads to diffuse information of small targets: In voxel feature extraction networks (such as 3D backbone networks), in order to expand the receptive field and extract deep semantics, the network usually stacks multiple downsampling layers sequentially (such as 1x, 2x, 4x, 8x downsampling). Since the point cloud of distant targets in snowy weather is very sparse, after these sparse features are downsampled multiple times, the fine geometric details and boundary information preserved in the high-resolution layer are often directly blurred or lost, which makes it impossible for the network to construct an effective target representation in the deep layer, greatly weakening the model's ability to detect small targets or distant occluded targets.
[0006] (3) Multi-task conflicts caused by coupled detection heads: Most mainstream networks currently use coupled detection heads, which output the target's category (classification), center point offset, length and width dimensions, and yaw angle simultaneously through a single path and a simple 1×1 convolution. In practical applications, classification tasks require the network to focus on the target's global semantic features and texture, while bounding box regression tasks heavily rely on the target's edges and local geometric features. In snowy scenes with sparse features and complex backgrounds, this coupled design can lead to mutual interference and conflicts among multiple tasks during optimization, making it difficult for the network to focus accurately and further resulting in low-quality candidate box generation.
[0007] (4) Feature forgetting phenomenon in cascaded networks: When refining the initial candidate region using a cascaded architecture, existing methods often rely solely on the features of the current stage for fine-tuning, or simply concatenate features from multiple stages. This crude fusion approach cannot adaptively learn the mutual influence and importance between features from different stages, resulting in the inability of the high-definition detail features extracted in the early stages to be effectively passed to the final prediction stage, thus limiting the upper limit of detection performance.
[0008] In summary, overcoming the sparsity and strong noise interference of snow point cloud data, optimizing the voxel feature extraction process, and decoupling the detection task to improve the quality of candidate regions are major technical challenges that need to be solved in this field. Summary of the Invention
[0009] This invention aims to provide a 3D target detection method that integrates residual submanifold convolution and dual attention. While retaining the advantages of efficient parallel computing of voxelization methods, this method significantly improves the network's comprehensive reconstruction and innovative design of multiple purification networks, detection heads, and cascaded refinement networks by addressing issues such as feature extraction backbone, high data noise in feature dimensionality reduction, feature diffusion of small targets, and task coupling conflicts. It effectively solves the detection performance problems of point clouds, small targets, and occluded targets in snowy scenes.
[0010] The 3D object detection method that integrates residual submanifold convolution and dual attention includes the following steps: A. Construct a 3D object detection network model, which includes: a 3D backbone network, a 2D bird's-eye view feature network, a decoupled multi-branch detection head, and a cascaded attention network. B. Train the 3D object detection network model to obtain a trained 3D object detection network model; C. Preprocess the point cloud data of the lidar to be tested to obtain preprocessed point cloud data; D. Input the preprocessed point cloud data into the trained 3D object detection network model. First, it enters the 3D backbone network for voxel feature extraction, obtaining comprehensive and complementary feature information at different resolution levels. The 2D bird's-eye view feature network compresses the feature information into bird's-eye view BEV features along the height direction, and enhances the bird's-eye view BEV features to obtain enhanced features. After the enhanced feature solution is predicted by the decoupled multi-branch detection head, it is input into the cascaded attention network. The region proposal is gradually optimized and refined using multi-stage target region features, and the final 3D target bounding box is output, which is the final 3D object detection result.
[0011] The preprocessing process for the point cloud data of the lidar under test includes the following steps: a. Divide the original irregular 3D point cloud data into three-dimensional voxel meshes of equal size; b. Model the space using voxel structures to obtain a three-dimensional voxel model, which is the preprocessed point cloud data.
[0012] The 3D backbone network is a multi-branch feature fusion residual submanifold network (RMFN), and its processing includes the following steps: The input results are sequentially processed by Submanifold convolution (SubM) to expand the number of feature channels to 16 while maintaining the resolution by 1x, resulting in the SubM processing result. The SubM processing result is then processed by residual subflow modules three times to obtain three residual subflow module processing results. The three residual subflow module processing results and the SubM processing result are then summed element-wise to obtain the fused feature. The fused features are processed by the first SRBSubM module. The results of the first SRBSubM module are then processed by the second SRBSubM module four times. The four results are then summed element-wise. Finally, the results are processed by the third SRBSubM module and the first Spareconv module to obtain the output.
[0013] The first, second, and third SRBSubM modules have the same structure, differing only in the number of input and output channels. Specifically, the first SRBSubM module has 16 input channels and 32 output channels; the second SRBSubM module has 32 input channels and 64 output channels; the third SRBSubM module has 64 input channels and 128 output channels; and the first Spareconv module has 128 input and output channels. The processing procedure includes the following steps: The input result is processed sequentially through the second Spareconv module, the BatchNorm function, and the ReLU function to obtain intermediate processing results. The intermediate processing results are then processed sequentially through the first SubMConv module, the BatchNorm function, the ReLU function, the second SubMConv module, and the BatchNorm function. The resulting product is multiplied by the intermediate processing results and then processed by the ReLU function to obtain the output result. The number of input and output channels of the second Spareconv module is the same as that of its corresponding SRBSubM module; the number of input and output channels of both the first and second SubMConv modules is the same as that of their corresponding SRBSubM modules.
[0014] The processing steps in the 2D bird's-eye view feature network include the following: The input result is first processed by transport linear projection to change the number of channels from 128 to 512, and then processed by convolutional layer to change the number of channels to 64. Then it is processed by convolutional layer, SCAC module, SCAC module, CBAM module and convolutional layer in sequence to obtain intermediate results. The number of channels 64 is kept unchanged throughout these processes. The intermediate results are divided into two paths. The first path is processed by a convolutional layer to change the number of channels to 128, resulting in the first result. The second path is processed by a convolutional layer to change the number of channels to 128, and then sequentially processed by a convolutional layer, a SCAC module, a CBAM module, a convolutional layer, and another convolutional layer to obtain the second result. The number of channels remains unchanged at 128 throughout these processes. After element-wise summation of the first and second results, an output with 256 channels is obtained.
[0015] The processing procedure in the SCAC module includes the following steps: The input result is split into informative and low-information features using the GroupNorm function. The informative and low-information features are then processed by the Sigmoid function to obtain the informative features. and redundancy features Then it splits into two branches, the first branch consisting of feature-rich... and redundancy features After element-wise summation, the first branch result is obtained; the second branch consists of feature-rich... and redundancy features After element-wise summation, the second branch result is obtained; the first and second branch results are then concatenated using the Concat function to obtain the output result. ; Output Channels are divided into characteristic αC and characteristic (1-α)C There are two parts, where α is a learnable parameter that is randomly set during deep learning and then automatically tuned as training progresses; Feature αC After 1×1 convolution processing, it is divided into two paths: one path is processed by GWC convolution, and the other path is processed by PWC convolution. The two convolution results are summed element-wise to obtain the output result Y1, and then processed by average pooling to obtain the first average pooling result. Feature (1-α)C After 1×1 convolution, followed by PWC convolution, the result is similar to the feature (1-α)C. After element-wise summation, the output result Y2 is obtained. Then, after average pooling, the second average pooling result is obtained. After the first and second average pooling results are processed by the SoftMax function, the results are divided into two paths. The first path is multiplied by the output result Y1 to obtain the first path result. The first path is multiplied by the output result Y2 to obtain the second path result. The first path result and the second path result are summed element-wise to obtain the output result.
[0016] The processing steps in the CBAM module are as follows: the input results are processed by average pooling and max pooling respectively. The two pooling results are processed by the Share MLP module respectively. The two processed results are then summed element-wise and then processed by the Sigmoid function to obtain the intermediate result. The intermediate results are processed by average pooling and max pooling respectively. The two pooling results are concatenated by the Concat function. The concatenated result is then processed by the Conv layer module and the Sigmoid function. The result is multiplied by the intermediate results to obtain the output result.
[0017] The prediction process in the decoupled multi-branch detection head includes the following steps: The input results are first filtered by a masking mechanism to identify regions that may contain the target. Then, they are processed sequentially by 3×3 convolution, BatchNorm function, and ReLU function. The 64-dimensional shared features of the regions are then extracted and split into classification branch, height branch, offset branch, size branch, and angle branch. The height branch, offset branch, size branch, and angle branch are concatenated by the Concat function. The resulting output is obtained by element-wise summation of the output, the classification branch, and the input results.
[0018] The processing steps in the cascaded attention network include the following: The input result is processed by the Region Generation Network (RPN) to obtain the RPN processing result. The RPN processing result is then processed sequentially by the POOL module, the first Encoder module, and the first multi-head attention mechanism module, and then detected by the first 3D detection head. The detection result of the first 3D detection head is then processed sequentially by the POOL module and the second Encoder module. The result obtained from the first 3D detection head is then processed together with the result obtained from the first multi-head attention mechanism module and then processed by the second multi-head attention mechanism module. Finally, the detection result of the second 3D detection head is processed sequentially by the POOL module and the third Encoder module. The result obtained from the second 3D detection head is then processed together with the results obtained from the first and second multi-head attention mechanism modules and then processed by the third multi-head attention mechanism module. Finally, the result obtained from the third 3D detection head is detected to obtain the output result.
[0019] The first, second, and third multi-head attention mechanism modules have the same structure, but differ in their input results. The processing steps include the following: In the first multi-head attention mechanism module, the output of the first Encoder module is used as the input. This input is divided into three paths, each of which is processed by a linear layer to obtain three results, defined as a query matrix, a key matrix, and a value matrix. In the second multi-head attention mechanism module, the output of the first multi-head attention mechanism module is copied into two results, which are defined as the query matrix Query and the key matrix Key, respectively. The output of the second Encoder module is defined as the value matrix Value. In the third multi-head attention mechanism module, the output of the first multi-head attention mechanism module is defined as the query matrix Query, the output of the second multi-head attention mechanism module is defined as the key matrix Key, and the output of the third Encoder module is defined as the value matrix Value. These three matrices are combined to form the input features. The input features are copied multiple times, and each set of input features is fed into the scaling dot product attention module for processing, resulting in multiple scaling dot product attention module processing results. These results are concatenated by the Concat function and then processed by the linear layer to obtain the output result. The processing procedure in the scaling dot product attention module is as follows: The query matrix and key matrix are processed by the MatMul function. The resulting matrix multiplication is then processed by the Scale function, Mask (opt.), and SoftMax function. The resulting matrix and value matrix are then processed by the MatMul function to obtain the output.
[0020] The beneficial effects of this invention are as follows: 1. This invention creatively proposes a "clean first, enhance later" feature processing paradigm, which greatly improves the model's resistance to noise interference in complex snowy weather backgrounds. In traditional CNN networks, the feature responses of snowflake point clouds are blindly amplified; this invention first uses the SCAC module to separate and suppress redundant features with low information content by utilizing the variance characteristics of group normalization, significantly reducing computational complexity and model processing burden. Subsequently, the cascaded CBAM module can accurately highlight the outline of the target subject on the clean feature map without noise interference by utilizing a large receptive field spatial attention mechanism. This joint design enables the model to accurately lock onto the target even in extremely harsh scenes with blizzard interference, effectively reducing the false detection rate.
[0021] 2. This invention introduces a multi-branch feature fusion residual submanifold network (RMFN), which effectively overcomes the problem of information diffusion in small targets and significantly improves the recall rate of distant targets. Addressing the information scarcity caused by sparse point clouds, RMFN uses a multi-path extraction and fusion method at the highest resolution level (1x downsampling), enabling the lower-level feature maps to incorporate information from more different receptive fields. This provides higher-quality, high-fidelity input features for subsequent compression and proposal generation stages, fundamentally ensuring the localization accuracy of small targets.
[0022] 3. This invention designs a decoupled detection head, DMAHead, combined with a cascaded attention network, breaking through the performance ceiling of traditional voxel detection models. DMAHead effectively coordinates the inherent contradictions between classification and geometric regression tasks through feature sharing and branch independence, significantly improving the quality of candidate boxes. The cascaded attention network cleverly borrows the self-attention mechanism from natural language processing, dynamically allocating the weights of cross-stage features based on occlusion and target shape, effectively fusing previous region features. Experiments show that this combination not only accelerates the model's convergence speed but also achieves a qualitative leap in the compactness and accuracy of the final detection boxes.
[0023] 4. Demonstrates outstanding performance on authoritative test sets, possessing extremely high potential for clinical and industrial applications. Based on rigorous evaluation on the Canadian Adverse Driving Conditions Dataset (CADC), the 3D object detection network model constructed in this invention achieves, compared to the baseline model Voxel-RCNN, an absolute accuracy improvement of 4.41%, 3.48%, and 5.14% respectively in the highly challenging vehicle 3D detection (IoU=0.7) at the Easy, Mod, and Hard difficulty levels; the BEV detection accuracy reaches 95.61%, 92.81%, and 92.04% respectively. Its overall performance significantly surpasses existing mainstream 3D object detection algorithms such as PointPillars, Point-RCNN, and PV-RCNN, providing more reliable support for the decision-making and planning modules of intelligent driving systems. Attached Figure Description
[0024] Figure 1 This is a diagram illustrating the overall system architecture and workflow of the 3D object detection network model in Example 1. Figure 2 This is a schematic diagram of the multi-branch feature fusion residual submanifold network (RMFN) of Example 1; Figure 3 This is a schematic diagram of the structure of the 2D bird's-eye view feature network in Example 1; Figure 4 This is a schematic diagram of the SCAC module in Example 1; Figure 5 This is a schematic diagram of the CBAM module in Example 1; Figure 6 This is a schematic diagram of the decoupled multi-branch detection head in Example 1; Figure 7 This is a schematic diagram of the cascaded attention network in Example 1; Figure 8 This is a schematic diagram of the multi-head attention mechanism module in Example 1; Figure 9 is an example image of the detection result of the method of the present invention in Example 2; Figure 10 This is a comparison chart of the detection performance of the method of the present invention in Example 2 and Voxel-RCNN in occluded scenes; Figure 11 This is a comparison chart of the detection performance of the method of the present invention in Example 2 and Voxel-RCNN in long-distance scenes. Detailed Implementation
[0025] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Example 1
[0026] The 3D object detection method that integrates residual submanifold convolution and dual attention in this embodiment includes the following steps: A. Construct a 3D object detection network model, such as Figure 1 As shown, the 3D object detection network model includes: a 3D backbone network, a 2D bird's-eye view feature network, a decoupled multi-branch detection head, and a cascaded attention network; B. Train the 3D object detection network model to obtain a trained 3D object detection network model; C. Preprocess the point cloud data of the lidar to be tested to obtain preprocessed point cloud data; The preprocessing process for the point cloud data of the lidar under test includes the following steps: a. Divide the original irregular 3D point cloud data into three-dimensional voxel meshes of equal size; b. Model the space using voxel structures to obtain a three-dimensional voxel model, which is the preprocessed point cloud data.
[0027] D. Input the preprocessed point cloud data into the trained 3D object detection network model. First, it enters the 3D backbone network for voxel feature extraction, obtaining comprehensive and complementary feature information at different resolution levels. The 2D bird's-eye view feature network compresses the feature information into bird's-eye view BEV features along the height direction, and enhances the bird's-eye view BEV features to obtain enhanced features. After the enhanced feature solution is predicted by the decoupled multi-branch detection head, it is input into the cascaded attention network. The region proposal is gradually optimized and refined using multi-stage target region features, and the final 3D target bounding box is output, which is the final 3D object detection result.
[0028] like Figure 2 As shown, the 3D backbone network is a multi-branch feature fusion residual submanifold network (RMFN), and its processing includes the following steps: The input results are sequentially processed by Submanifold convolution (SubM) to expand the number of feature channels to 16 while maintaining the resolution by 1x, resulting in the SubM processing result. The SubM processing result is then processed by residual subflow modules three times to obtain three residual subflow module processing results. The three residual subflow module processing results and the SubM processing result are then summed element-wise to obtain the fused feature. The fused features are processed by the first SRBSubM module. The results of the first SRBSubM module are then processed by the second SRBSubM module four times. The four results are then summed element-wise. Finally, the results are processed by the third SRBSubM module and the first Spareconv module to obtain the output.
[0029] The first, second, and third SRBSubM modules have the same structure, differing only in the number of input and output channels. Specifically, the first SRBSubM module has 16 input channels and 32 output channels; the second SRBSubM module has 32 input channels and 64 output channels; the third SRBSubM module has 64 input channels and 128 output channels; and the first Spareconv module has 128 input and output channels. The processing procedure includes the following steps: The input result is processed sequentially through the second Spareconv module, the BatchNorm function, and the ReLU function to obtain intermediate processing results. The intermediate processing results are then processed sequentially through the first SubMConv module, the BatchNorm function, the ReLU function, the second SubMConv module, and the BatchNorm function. The resulting product is multiplied by the intermediate processing results and then processed by the ReLU function to obtain the output result. The number of input and output channels of the second Spareconv module is the same as that of its corresponding SRBSubM module; the number of input and output channels of both the first and second SubMConv modules is the same as that of their corresponding SRBSubM modules.
[0030] like Figure 3 As shown, the processing steps in the 2D bird's-eye view feature network include the following: The input result is first processed by transport linear projection to change the number of channels from 128 to 512, and then processed by convolutional layer to change the number of channels to 64. Then it is processed by convolutional layer, SCAC module, SCAC module, CBAM module and convolutional layer in sequence to obtain intermediate results. The number of channels 64 is kept unchanged throughout these processes. The intermediate results are divided into two paths. The first path is processed by a convolutional layer to change the number of channels to 128, resulting in the first result. The second path is processed by a convolutional layer to change the number of channels to 128, and then sequentially processed by a convolutional layer, a SCAC module, a CBAM module, a convolutional layer, and another convolutional layer to obtain the second result. The number of channels remains unchanged at 128 throughout these processes. After element-wise summation of the first and second results, an output with 256 channels is obtained.
[0031] like Figure 4 As shown, the processing steps in the SCAC module include the following: The input result is split into informative and low-information features using the GroupNorm function. The informative and low-information features are then processed by the Sigmoid function to obtain the informative features. and redundancy features Then it splits into two branches, the first branch consisting of feature-rich... and redundancy features After element-wise summation, the first branch result is obtained; the second branch consists of feature-rich... and redundancy features After element-wise summation, the second branch result is obtained; the first and second branch results are then concatenated using the Concat function to obtain the output result. ; Output Channels are divided into characteristic αC and characteristic (1-α)C There are two parts, where α is a learnable parameter that is randomly set during deep learning and then automatically tuned as training progresses; Feature αC After 1×1 convolution processing, it is divided into two paths: one path is processed by GWC convolution, and the other path is processed by PWC convolution. The two convolution results are summed element-wise to obtain the output result Y1, and then processed by average pooling to obtain the first average pooling result. Feature (1-α)C After 1×1 convolution, followed by PWC convolution, the result is similar to the feature (1-α)C. After element-wise summation, the output result Y2 is obtained. Then, after average pooling, the second average pooling result is obtained. After the first and second average pooling results are processed by the SoftMax function, the results are divided into two paths. The first path is multiplied by the output result Y1 to obtain the first path result. The first path is multiplied by the output result Y2 to obtain the second path result. The first path result and the second path result are summed element-wise to obtain the output result.
[0032] like Figure 5 As shown, the processing in the CBAM module includes the following steps: the input results are processed by average pooling and max pooling respectively. After the two pooling results are processed by the Share MLP module, the two processed results are summed element-wise and then processed by the Sigmoid function to obtain the intermediate result. The intermediate results are processed by average pooling and max pooling respectively. The two pooling results are concatenated by the Concat function. The concatenated result is then processed by the Conv layer module and the Sigmoid function. The result is multiplied by the intermediate results to obtain the output result.
[0033] like Figure 6 As shown, the prediction process in the decoupled multi-branch detection head includes the following steps: The input results are first filtered by a masking mechanism to identify regions that may contain the target. Then, they are processed sequentially by 3×3 convolution, BatchNorm function, and ReLU function. The 64-dimensional shared features of the regions are then extracted and split into classification branch, height branch, offset branch, size branch, and angle branch. The height branch, offset branch, size branch, and angle branch are concatenated by the Concat function. The resulting output is obtained by element-wise summation of the output, the classification branch, and the input results.
[0034] like Figure 7 As shown, the processing in the cascaded attention network includes the following steps: The input result is processed by the Region Generation Network (RPN) to obtain the RPN processing result. The RPN processing result is then processed sequentially by the POOL module, the first Encoder module, and the first multi-head attention mechanism module, and then detected by the first 3D detection head. The detection result of the first 3D detection head is then processed sequentially by the POOL module and the second Encoder module. The result obtained from the first 3D detection head is then processed together with the result obtained from the first multi-head attention mechanism module and then processed by the second multi-head attention mechanism module. Finally, the detection result of the second 3D detection head is processed sequentially by the POOL module and the third Encoder module. The result obtained from the second 3D detection head is then processed together with the results obtained from the first and second multi-head attention mechanism modules and then processed by the third multi-head attention mechanism module. Finally, the result obtained from the third 3D detection head is detected to obtain the output result.
[0035] like Figure 8 As shown, the first, second, and third multi-head attention mechanism modules have the same structure, but differ in their input results. The processing steps include the following: In the first multi-head attention mechanism module, the output of the first Encoder module is used as the input. This input is divided into three paths, each of which is processed by a linear layer to obtain three results, defined as a query matrix, a key matrix, and a value matrix. In the second multi-head attention mechanism module, the output of the first multi-head attention mechanism module is copied into two results, which are defined as the query matrix Query and the key matrix Key, respectively. The output of the second Encoder module is defined as the value matrix Value. In the third multi-head attention mechanism module, the output of the first multi-head attention mechanism module is defined as the query matrix Query, the output of the second multi-head attention mechanism module is defined as the key matrix Key, and the output of the third Encoder module is defined as the value matrix Value. These three matrices are combined to form the input features. The input features are copied 6 times, and each set of input features is fed into the scaling dot product attention module for processing, resulting in multiple scaling dot product attention module processing results. These results are concatenated by the Concat function and then processed by the linear layer to obtain the output result. The processing procedure in the scaling dot product attention module is as follows: The query matrix and key matrix are processed by the MatMul function. The resulting matrix multiplication is then processed by the Scale function, Mask (opt.), and SoftMax function. The resulting matrix and value matrix are then processed by the MatMul function to obtain the output. Example 2
[0036] The 3D object detection algorithm studied in complex snowy weather scenarios is mainly applied in the fields of advanced autonomous driving and intelligent transportation environmental perception. It places extremely high demands on the algorithm's anti-interference ability and detection accuracy under adverse weather conditions (such as wind, snow, noise interference, target occlusion, and sparse point clouds). Therefore, the performance of the algorithm of this invention is compared with that of current mainstream 3D object detection algorithms under the same conditions. Mainstream 3D object detection algorithms are primarily based on single radar point clouds and multimodal fusion algorithms based on camera-point cloud. This embodiment selects the current mainstream advanced detection network for comparison.
[0037] The algorithm of this invention adopts the algorithm of Example 1, which is named Voxel-MACN in this example.
[0038] This embodiment compares Voxel-RCNN, RCAVoxel-RCNN, PointPillars, Point-RCNN, PV-RCNN, Part-A2, F-PointNet, AVOD, and Snow-CLOCs with the Voxel-MACN network algorithm of Embodiment 1 of this invention under the same snowy dataset (CADC dataset) and experimental parameters. The specific results are shown in Table 1.
[0039]
[0040] Note: In the Modality column of the table, "L" means using only LiDAR point cloud data, and "L+C" means using multimodal fusion data from LiDAR and camera.
[0041] As shown in Table 1, the method of this invention exhibits superior performance across all evaluation metrics. The model of this invention achieves the highest 3D target detection accuracy and the highest BEV (bird's-eye view) detection accuracy across all three difficulty levels of vehicle targets, significantly outperforming all other single-modal and multi-modal comparison models. Furthermore, the method of this invention also achieves the highest performance on the core challenge of target detection—the Hard level metric (representing extremely distant, sparse, and severely occluded vehicle targets in extreme snowy conditions). Specifically, the Car 3D (Hard) metric is significantly higher than the baseline algorithm (Voxel-RCNN) by 5.14%, and the BEV (Hard) metric is 1.57% higher than the baseline algorithm. It also achieves accuracy leaps of 4.41% and 3.48% respectively under Easy and Mod difficulty levels. This superior accuracy is achieved under extremely complex wind and snow background noise interference. Compared to multimodal algorithms (such as F-PointNet and AVOD) that rely not only on point clouds but also incorporate image data, this invention uses only a single point cloud modality and still achieves comprehensive superiority in accuracy. Experimental results strongly demonstrate that the Voxel-MACN method of this invention, by introducing the SCAC spatial and channel noise reduction module and the DMAHead decoupled detection head, successfully achieves a perfect balance between noise filtering and sparse feature enhancement, and is implemented in a highly efficient multi-level attention cascade (MACN) architecture. This is very suitable for real-world autonomous driving applications with complex environmental perception, limited computing resources, and extremely high safety requirements.
[0042] To intuitively evaluate the performance of the Voxel-MCAN algorithm model of this invention, we randomly selected a portion of point cloud data from the test set for inference and visualization, as shown below. Figure 9 As shown. This data covers various snowy weather scenarios, including distant targets and distant, obscured targets. For example... Figure 9As shown, Voxel-MCAN exhibits excellent performance in snowy conditions, accurately identifying both distant and occluded targets. This demonstrates that the model can effectively analyze occluded vehicles, thus providing more reliable support for decision-making and planning in intelligent driving.
[0043] To verify the detection performance of the Voxel-MCAN algorithm in occluded scenes, this embodiment randomly selects point cloud data from the test set for inference and compares it with the Voxel-RCNN algorithm. Figure 10 The visualization results show that the areas marked with red circles within the blue boxes represent occluded scenes. A comparison reveals that Voxel-RCNN can only identify vehicles when occluded at close range, while the Voxel-MCAN algorithm can effectively identify vehicles occluded at a distance. Key areas of the model inference results visualization are highlighted with red circles.
[0044] In snowy conditions, due to the limitations of LiDAR scanning principles, the point cloud of distant targets is often sparse. This can be addressed by visualizing the analysis and reasoning results. Figure 11 It can be seen that the Voxel-MCAN algorithm of this invention also outperforms the RCAVoxel-RCNN algorithm in long-range target detection performance. For example... Figure 11 As shown above, the blue magnified box contains visible vehicle targets. While not all visible vehicles were identified, a portion of them were successfully detected, whereas RCAVoxel-RCNN failed to identify any at all. For the area below the image, the Voxel-MCAN algorithm was also able to effectively identify distant, occluded target vehicles.
[0045] Therefore, it can be seen that the algorithm of this invention has greatly improved the detection capability of distant targets and distant occlusion situations. Areas with differing detection results are highlighted with red circles in the figure.
[0046] The above-disclosed content is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Therefore, all equivalent technical changes made based on the description and drawings of the present invention are included within the scope of protection of the present invention. Furthermore, the elements therein can be updated as technology develops. The above units are merely examples, and those skilled in the art can adopt corresponding units according to actual needs when implementing this solution.
Claims
1. A 3D target detection method integrating residual submanifold convolution and dual attention, characterized in that, Includes the following steps: A. Construct a 3D object detection network model, which includes: a 3D backbone network, a 2D bird's-eye view feature network, a decoupled multi-branch detection head, and a cascaded attention network. B. Train the 3D object detection network model to obtain a trained 3D object detection network model; C. Preprocess the point cloud data of the lidar to be tested to obtain preprocessed point cloud data; D. Input the preprocessed point cloud data into the trained 3D object detection network model. First, it enters the 3D backbone network for voxel feature extraction, obtaining comprehensive and complementary feature information at different resolution levels. The 2D bird's-eye view feature network compresses the feature information into bird's-eye view BEV features along the height direction, and enhances the bird's-eye view BEV features to obtain enhanced features. After the enhanced feature solution is predicted by the decoupled multi-branch detection head, it is input into the cascaded attention network. The region proposal is gradually optimized and refined using multi-stage target region features, and the final 3D target bounding box is output, which is the final 3D object detection result.
2. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The preprocessing process for the point cloud data of the lidar under test includes the following steps: a. Divide the original irregular 3D point cloud data into three-dimensional voxel meshes of equal size; b. Model the space using voxel structures to obtain a three-dimensional voxel model, which is the preprocessed point cloud data.
3. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The 3D backbone network is a multi-branch feature fusion residual submanifold network (RMFN), and its processing includes the following steps: The input results are sequentially processed by Submanifold convolution (SubM) operations, which expand the number of feature channels to 16 while maintaining the resolution by 1x, resulting in the SubM processing results. The SubM processing results are then processed by residual substream modules three times to obtain the three residual substream module processing results. The processing results of the three residual substream modules and the processing result of SubM are summed element-wise to obtain the fused features; The fused features are processed by the first SRBSubM module. The results of the first SRBSubM module are then processed by the second SRBSubM module four times. The four results are then summed element-wise. Finally, the results are processed by the third SRBSubM module and the first Spareconv module to obtain the output.
4. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 3, characterized in that: The first SRBSubM module, the second SRBSubM module, and the third SRBSubM module have the same structure, only the number of input channels and the number of output channels are set differently. in; The first SRBSubM module has 16 input channels and 32 output channels; the second SRBSubM module has 32 input channels and 64 output channels; the third SRBSubM module has 64 input channels and 128 output channels; the first Spareconv module has 128 input channels and 128 output channels. The processing procedure includes the following steps: The input result is processed sequentially through the second Spareconv module, the BatchNorm function, and the ReLU function to obtain intermediate processing results. The intermediate processing results are then processed sequentially through the first SubMConv module, the BatchNorm function, the ReLU function, the second SubMConv module, and the BatchNorm function. The resulting product is multiplied by the intermediate processing results and then processed by the ReLU function to obtain the output result. The number of input and output channels of the second Spareconv module is the same as that of its corresponding SRBSubM module; the number of input and output channels of both the first and second SubMConv modules is the same as that of their corresponding SRBSubM modules.
5. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The processing steps in the 2D bird's-eye view feature network include the following: The input result is first processed by transport linear projection to change the number of channels from 128 to 512, and then processed by convolutional layer to change the number of channels to 64. Then it is processed by convolutional layer, SCAC module, SCAC module, CBAM module and convolutional layer in sequence to obtain intermediate results. The number of channels 64 is kept unchanged throughout these processes. The intermediate results are divided into two paths. The first path is processed by a convolutional layer to change the number of channels to 128, resulting in the first result. The second path is processed by a convolutional layer to change the number of channels to 128, and then sequentially processed by a convolutional layer, a SCAC module, a CBAM module, a convolutional layer, and another convolutional layer to obtain the second result. The number of channels remains unchanged at 128 throughout these processes. After element-wise summation of the first and second results, an output with 256 channels is obtained.
6. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The processing procedure in the SCAC module includes the following steps: The input result is split into informative and low-information features using the GroupNorm function. The informative and low-information features are then processed by the Sigmoid function to obtain the informative features. and redundancy features ; Then it splits into two branches, the first branch consisting of feature-rich... The first branch result is obtained by summing the redundant feature elements; the second branch is obtained by summing the features rich at the element level. and redundancy features After element-wise summation, the second branch result is obtained; the first and second branch results are then concatenated using the Concat function to obtain the output result. ; Output Channels are divided into characteristic αC and characteristic (1-α)C There are two parts, where α is a learnable parameter that is randomly set during deep learning and then automatically tuned as training progresses; Feature αC After 1×1 convolution processing, it is divided into two paths: one path is processed by GWC convolution, and the other path is processed by PWC convolution. The two convolution results are summed element-wise to obtain the output result Y1, and then processed by average pooling to obtain the first average pooling result. Feature (1-α)C After 1×1 convolution, followed by PWC convolution, the result is similar to the feature (1-α)C. After element-wise summation, the output result Y2 is obtained. Then, after average pooling, the second average pooling result is obtained. After the first and second average pooling results are processed by the SoftMax function, the results are divided into two paths. The first path is multiplied by the output result Y1 to obtain the first path result. The first path is multiplied by the output result Y2 to obtain the second path result. The first path result and the second path result are summed element-wise to obtain the output result.
7. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The processing steps in the CBAM module are as follows: the input results are processed by average pooling and max pooling respectively. The two pooling results are processed by the Share MLP module respectively. The two processed results are then summed element-wise and then processed by the Sigmoid function to obtain the intermediate result. The intermediate results are processed by average pooling and max pooling respectively. The two pooling results are concatenated by the Concat function. The concatenated result is then processed by the Conv layer module and the Sigmoid function. The result is multiplied by the intermediate results to obtain the output result.
8. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The prediction process in the decoupled multi-branch detection head includes the following steps: The input results are first filtered by a masking mechanism to identify regions that may contain the target. Then, they are processed by 3×3 convolution, BatchNorm function, and ReLU function in sequence. Finally, 64-dimensional shared features of the region are extracted and split into classification branch, height branch, offset branch, size branch, and angle branch. The height branch, offset branch, size branch, and angle branch are concatenated using the Concat function. The resulting output is then summed element-wise with the classification branch and the input result.
9. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 1, characterized in that: The processing steps in the cascaded attention network include the following: The input result is processed by the Region Generation Network (RPN) to obtain the RPN processing result. The RPN processing result is then processed sequentially by the POOL module, the first Encoder module, and the first multi-head attention mechanism module, and then detected by the first 3D detection head. The detection result of the first 3D detection head is then processed sequentially by the POOL module and the second Encoder module. The result obtained from the first 3D detection head is then processed together with the result obtained from the first multi-head attention mechanism module and then processed by the second multi-head attention mechanism module. Finally, the detection result of the second 3D detection head is processed sequentially by the POOL module and the third Encoder module. The result obtained from the second 3D detection head is then processed together with the results obtained from the first and second multi-head attention mechanism modules and then processed by the third multi-head attention mechanism module. Finally, the result obtained from the third 3D detection head is detected to obtain the output result.
10. The 3D target detection method fusing residual submanifold convolution and dual attention as described in claim 9, characterized in that: The first, second, and third multi-head attention mechanism modules have the same structure, but differ in their input results. The processing steps include the following: In the first multi-head attention mechanism module, the output of the first Encoder module is used as the input. This input is divided into three paths, each of which is processed by a linear layer to obtain three results, defined as a query matrix, a key matrix, and a value matrix. In the second multi-head attention mechanism module, the output of the first multi-head attention mechanism module is copied into two results, which are defined as the query matrix Query and the key matrix Key, respectively. The output of the second Encoder module is defined as the value matrix Value. In the third multi-head attention mechanism module, the output of the first multi-head attention mechanism module is defined as the query matrix Query, the output of the second multi-head attention mechanism module is defined as the key matrix Key, and the output of the third Encoder module is defined as the value matrix Value. These three matrices are combined to form the input features. The input features are copied multiple times, and each set of input features is fed into the scaling dot product attention module for processing, resulting in multiple scaling dot product attention module processing results. These results are concatenated by the Concat function and then processed by the linear layer to obtain the output result. The processing procedure in the scaling dot product attention module is as follows: The query matrix and key matrix are processed by the MatMul function. The resulting matrix multiplication is then processed by the Scale function, Mask (opt.), and SoftMax function. The resulting matrix and value matrix are then processed by the MatMul function to obtain the output.