Millimeter wave radar BEV target detection method based on sparse attention mechanism

By introducing sparse attention mechanism and PointPillars architecture in the millimeter-wave radar point cloud data processing, the problem of sparseness and noise interference of point cloud data is solved, and a more efficient and robust target detection effect is achieved.

CN120236069AActive Publication Date: 2025-07-01ZHEJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510709608.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The sparsity and noise interference of millimeter-wave radar point cloud data lead to a decrease in target detection accuracy, making it difficult to achieve accurate detection in complex environments.

Method used

The millimeter-wave radar BEV target detection method based on sparse attention mechanism is adopted, and the point cloud features are encoded into BEV feature maps through the PointPillars architecture, and a self-attention mechanism and a variable attention mechanism guided by anchor box are introduced to improve feature extraction and detection accuracy.

Benefits of technology

It significantly improves the efficiency and robustness of feature extraction, reduces the impact of noise on detection results, and improves the accuracy of object detection and the speed of network reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236069A_ABST
    Figure CN120236069A_ABST
Patent Text Reader

Abstract

The invention discloses a millimeter-wave radar BEV target detection method based on a sparse attention mechanism. The method comprises the following steps: acquiring millimeter-wave radar point cloud and preprocessing the millimeter-wave radar point cloud; the millimeter wave radar BEV target detection network construction comprises the following steps: carrying out point cloud feature coding based on a PointPill architecture to obtain a feature map, and carrying out feature enhancement on the feature map based on an anchor frame guided self-attention mechanism; and performing iterative processing on the enhanced features by using a cascade optimization module introducing a variable attention mechanism to obtain a prediction result, performing post-processing on network output to obtain a final result, and completing a BEV target detection task. According to the method, the efficiency and robustness of feature extraction are remarkably improved, meanwhile, the influence of noise on a detection result is reduced, and the target detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection, and particularly to a millimeter-wave radar BEV object detection method based on a sparse attention mechanism. Background Art

[0002] A millimeter-wave radar is a sensor that uses the millimeter-wave frequency band (usually 24 GHz, 77 GHz, or 79 GHz) for detection. By emitting electromagnetic waves and receiving the target reflection signals, it generates point cloud data. The wavelength of millimeter waves ranges from 1 mm to 10 mm, with a relatively high frequency and a short wavelength, which enables the millimeter-wave radar to achieve high-precision distance, speed, and angle measurements. At the same time, millimeter waves have strong penetration ability and can maintain good detection performance under harsh weather conditions such as rain, snow, and fog. Therefore, it has important value in object detection in complex environments.

[0003] The point cloud data generated by the millimeter-wave radar is obtained by processing the reflection signals. Each point usually contains information such as the distance, azimuth angle, elevation angle, speed, and reflection intensity of the target. However, compared with the point cloud generated by lidar, the millimeter-wave radar point cloud has obvious sparsity and low-resolution characteristics. This is because the beam of the millimeter-wave radar is relatively wide, and the energy distribution of its transmitted signal is relatively dispersed, resulting in a low density of the generated point cloud data and making it difficult to accurately describe the geometric shape and detailed features of the target. In addition, the millimeter-wave radar point cloud data is also easily affected by noise, such as multipath effects, clutter in the environment, and interference from other electromagnetic signals. These noises will further reduce the quality of the point cloud data and increase the difficulty of object detection.

[0004] The sparsity and noise interference characteristics of the millimeter-wave radar point cloud data pose unique technical challenges to object detection technology. First, the sparse point cloud data is difficult to provide sufficient information to accurately identify the boundaries and shapes of objects. Especially when the target is far away or the target size is small, the detection accuracy is likely to decrease. Second, noise interference will cause a large number of false alarm points or missed detection points in the point cloud data, affecting the robustness and reliability of object detection. In addition, the millimeter-wave radar point cloud data usually has a high dynamic range, and the separation of the target speed information from the static background is also an important issue.

[0005] The point cloud data generated by the millimeter-wave radar itself has the characteristics of sparsity and low resolution, which makes it more difficult to extract and identify object features. Especially in complex scenes or when the target is far away, the detection accuracy drops significantly. Second, radar signals are easily affected by noise, such as multipath effects, weather conditions (such as rain and snow), and other electromagnetic interferences in the environment. These noises will reduce the quality of the point cloud data and thus affect the robustness of the detection algorithm.

[0006] In terms of computational efficiency, when using a convolutional neural network to process millimeter-wave radar point cloud features, the receptive field is limited, which affects the target detection effect; when using the self-attention mechanism to encode point cloud features, the computational complexity is high.

[0007] In traditional target detection, fixed anchor box parameters are generally used, which are difficult to adapt to the geometric characteristics of different targets, resulting in inaccurate detection results. Summary of the Invention

[0008] The purpose of the present invention is to propose a millimeter-wave radar BEV target detection method based on a sparse attention mechanism in view of the deficiencies of the prior art.

[0009] The purpose of the present invention is achieved through the following technical solutions: A millimeter-wave radar BEV target detection method based on a sparse attention mechanism, the method includes the following steps:

[0010] S1. Obtain millimeter-wave radar point cloud and perform preprocessing;

[0011] S2. Construct a millimeter-wave radar BEV (Bird's-Eye View) target detection network, including: encoding point cloud features based on the PointPillars architecture to obtain a feature map, and enhancing the features of the feature map based on the anchor box-guided self-attention mechanism; using a cascade optimization module to iteratively process the enhanced features to obtain a prediction result;

[0012] The cascade optimization module is cascaded six layers in total, and each layer includes:

[0013] Adopt the self-attention mechanism of target features and the variable attention mechanism for extracting point cloud features based on PointPillars to perform variable sampling on the features; decode the sampling results and correct the anchor boxes;

[0014] S3. Train the target detection network, and use the combination of Focal Loss and SmoothL1 Loss as the loss function for training;

[0015] S4. Use the trained network for target detection, post-process the network output, obtain the final result, and complete the BEV target detection task.

[0016] Furthermore, the obtaining millimeter-wave radar point cloud and performing preprocessing includes:

[0017] First, the millimeter-wave radar emits a frequency-modulated continuous wave and receives the echo signal, and uses the fast Fourier transform to extract the distance, speed, and azimuth angle information of the target;

[0018] Generate point cloud data containing target coordinates and radar reflection intensity through multi-frame data accumulation and point cloud clustering; to adapt to network input, the point cloud is normalized to a unified coordinate system and outliers are removed;

[0019] Subsequently, to facilitate the data processing of the deep learning model, the point cloud is regularized to a unified quantity scale.

[0020] Furthermore, the specific process of encoding the point cloud features based on the PointPillars architecture to obtain the feature map is as follows:

[0021] First, use the PointPillars architecture to divide the point cloud into a regular three-dimensional columnar grid. The spatial range of each columnar body is a fixed size on the plane. For each point in each columnar body, calculate its offset relative to the center of the columnar body, concatenate the original features with the offset to obtain an enhanced feature vector, then encode the enhanced point cloud features through a multi-layer perceptron, map the encoded columnar body features to a two-dimensional BEV grid according to their spatial positions to form an initial BEV feature map, and finally use a two-dimensional convolutional network to aggregate the BEV feature map to obtain the feature map.

[0022] Furthermore, the specific process of enhancing the feature map based on the anchor box-guided self-attention mechanism is as follows:

[0023] Complete the structured organization of the target representation, including the anchor boxes of the target pose and size and the feature vectors encoding the target features; the anchor boxes are initialized in a uniformly distributed manner in the BEV space. Each anchor box is defined by its center coordinates, size, and heading angle, all of which are learnable tensors and are used to continuously optimize during the training process to adapt to the geometric characteristics of different targets; each target is represented by a 256-dimensional feature vector, and these feature vectors are initialized to all zeros and are used to gradually learn the semantic information of the target during the training process. Combine the anchor boxes with the feature vectors so that the model can capture both the geometric attributes and high-level semantic features of the target.

[0024] Furthermore, the self-attention mechanism of the target features enhances the context information of the target features by calculating the global correlation between the feature vectors. For each feature vector, calculate the query, key, and value vectors respectively through a learnable weight matrix, then calculate the attention score matrix through the query vector and the key vector, and obtain the enhanced feature vector by weighted aggregation of the value vectors. The formula for calculating the attention score matrix is:

[0025]

[0026] where, is the dimension of the feature vector, is the query vector corresponding to the i-th feature vector, is the key vector corresponding to the k-th feature vector, and N is the number of feature vectors;

[0027] An enhanced feature vector is obtained by weighted aggregation of the value vectors:

[0028]

[0029] is the value vector corresponding to the j-th feature vector.

[0030] Furthermore, the deformable attention mechanism for extracting point cloud features based on PointPillars includes:

[0031] For each target feature output by the self-attention mechanism , a set of sampling point coordinates is predicted, where K is the number of sampling points. Subsequently, the sampling point features are obtained by bilinear interpolation on the BEV feature map:

[0032]

[0033] And the sampling point features are aggregated by the attention weights:

[0034]

[0035] where is the learnable attention weight. This mechanism enables the target features to adaptively focus on the key regions and improves the detection ability for non-rigid targets.

[0036] Furthermore, the specific steps of feature decoding and anchor box correction for the sampling results include:

[0037] Input the target features output by the deformable attention mechanism . Nonlinear transformation is performed on the target features through a multi-layer fully connected network:

[0038]

[0039] where is the learnable weight matrix, is the bias term;

[0040] Bounding box correction: For each target feature , the offset of the anchor box is predicted:

[0041]

[0042] And the anchor box is corrected by the following formula:

[0043]

[0044] where They are the parameters after anchor box decoding.

[0045] Class prediction: Predict the target class probability through the fully connected layer and the Softmax function:

[0046]

[0047] Among them and are the learnable weight matrix and bias term.

[0048] Furthermore, the combination of Focal Loss and SmoothL1 Loss as the loss function for training specifically includes:

[0049] Among them, is the probability that the anchor box prediction belongs to a certain type of target, is the Focal Loss hyperparameter, is the difference between the predicted target bounding box and the ground truth, and weights are used to weight the two types of losses.

[0050] On the other hand, the present invention also provides a millimeter-wave radar BEV target detection device based on a sparse attention mechanism, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, the above-mentioned millimeter-wave radar BEV target detection method based on a sparse attention mechanism is implemented.

[0051] On the other hand, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the above-mentioned millimeter-wave radar BEV target detection method based on a sparse attention mechanism is implemented.

[0052] Advantages of the present invention:

[0053] The present invention converts sparse point clouds into dense BEV feature maps through the PointPillars framework, significantly improving the efficiency and robustness of feature extraction, and at the same time reducing the impact of noise on the detection results.

[0054] The present invention introduces a deformable attention mechanism to dynamically sample BEV point cloud features and adaptively focus on the key regions of the target, significantly improving the detection accuracy of the target. At the same time, the computational complexity is reduced and the network inference speed is increased.

[0055] The present invention enables the anchor box to be dynamically adjusted through a learnable anchor box initialization and optimization mechanism, better matching the actual shape and pose of the target, thereby improving the detection accuracy. Description of the Drawings

[0056] Figure 1 Flow chart of the millimeter-wave radar point cloud BEV target detection method provided by the embodiment of the present invention;

[0057] Figure 2 Schematic diagram of the millimeter-wave radar point cloud BEV target detection method provided by the embodiment of the present invention;

[0058] Figure 3 Schematic diagram of the millimeter-wave radar point cloud BEV target detection device provided by the embodiment of the present invention. Detailed implementation manners

[0059] The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.

[0060] As Figure 1 and Figure 2 shown, a millimeter-wave radar BEV target detection method based on a sparse attention mechanism provided by this embodiment includes:

[0061] S1. Obtain the millimeter-wave radar point cloud and perform preprocessing;

[0062] The acquisition and preprocessing of the millimeter-wave radar point cloud are key steps in the target detection task. First, the millimeter-wave radar emits a frequency-modulated continuous wave (FMCW) and receives the echo signal, and uses the fast Fourier transform (FFT) to extract the distance, speed, and azimuth angle information of the target. Through the accumulation of multiple frames of data and point cloud clustering, point cloud data containing target coordinates and radar reflection intensity is generated. To adapt to the network input, the point cloud is normalized to a unified coordinate system, and outlier points are removed. Subsequently, to facilitate the data processing of the deep learning model, the present invention regularizes the point cloud to a unified number scale (for example, it is stipulated that each frame of point cloud contains 2,000 points. If it is more than this value, 2,000 points are randomly sampled from all point clouds; if it is less than this value, it is filled with all-zero values), and used as the input of the PointPillars feature extraction module.

[0063] S2. Construct a millimeter-wave radar BEV target detection network, which is composed of a point cloud feature encoding module, a feature enhancement module, a deformable feature sampling module, and a cascaded optimization module, and adopts 6-layer iterative optimization to gradually improve the detection accuracy layer by layer.

[0064] In the point cloud feature encoding stage, the PointPillars architecture is adopted for feature extraction. By converting the millimeter-wave radar point cloud into a regular columnar grid, a multi-layer perceptron (MLP) is used to extract the local geometric features of each columnar body, and then a feature map in the bird's-eye view is constructed through a two-dimensional convolutional network.

[0065] The feature enhancement module innovatively constructs an anchor-guided self-attention mechanism. 900 prior anchor boxes are preset in the BEV (Bird's-Eye View) space (the specific number can be adjusted according to the target density of the application scenario), and each anchor box is parameterized and represented by a 256-dimensional feature vector. The self-attention mechanism is introduced to effectively enhance the ability of the target feature to perceive context information by calculating the global correlation between the anchor box features, especially having a significant effect on the feature completion of partially occluded targets.

[0066] The deformable feature sampling module adopts a two-stage attention fusion strategy. First, 13 dynamic sampling point coordinates are predicted based on the 256-dimensional target features, and these sampling points form a region of interest on the BEV feature map. Subsequently, through the deformable attention mechanism, bilinear interpolation is used to obtain the feature vectors at the corresponding positions, and feature aggregation is performed through attention weights. This design breaks through the limitation of the fixed receptive field of traditional convolutional operations, enabling the network to adaptively adjust the feature sampling region according to the target morphology, and is particularly suitable for processing sparse millimeter-wave radar point cloud features.

[0067] The cascaded optimization architecture adopts 6 layers of iterative processing units, and each layer contains a feature decoder and an anchor box correction module. The feature decoder performs non-linear transformation on the sampled features through a 3-layer fully connected network, and then predicts the probabilities of various categories of the anchor boxes. The anchor box correction module adopts a divide-and-conquer strategy: a 128-dimensional sub-network is used to predict the center point offset, and a 128-dimensional sub-network estimates the size and pose residuals. Through the layer-by-layer refinement correction mechanism, the network finally outputs the accurate coordinates , heading angle , size and category probabilities.

[0068] Specifically, the calculation process of the millimeter-wave radar BEV object detection network is as follows:

[0069] The point cloud feature extraction module of this network based on PointPillars converts the input point cloud data from the sparse three-dimensional space into a dense two-dimensional BEV feature map. The dimension of the input data is (B, 2000, 5), where B (Batch Size) represents the batch processing size, 2000 represents the number of points to which each frame of point cloud is regularized, and the 5-dimensional features include the three-dimensional coordinates (x, y) of the target, the velocity components (vx, vy), and the radar reflection intensity (RCS).

[0070] First, the point cloud is divided into regular three-dimensional columnar grids, and the spatial range of each columnar body is a fixed size on the (x, y) plane (for example, [0.2m × 0.2m]). For each columnar body, the point cloud features are processed through the following steps:

[0071] (1) Feature Enhancement: For each point, calculate its offset relative to the center of the cylinder :

[0072]

[0073] Concatenate the original feature with the offset to obtain an enhanced feature vector:

[0074]

[0075] (2) Feature Encoding: For each cylinder, use a multi-layer perceptron (MLP) to encode the point cloud features:

[0076]

[0077] where N is the number of points in the cylinder, and the MLP outputs the cylinder features with a fixed dimension (set to 256 dimensions in this application).

[0078] (3) Generation of BEV Feature Map.

[0079] Map the encoded cylinder features to a two-dimensional BEV grid according to their spatial positions to form an initial BEV feature map. Subsequently, further aggregate the BEV feature map through a two-dimensional convolutional network to output a dense feature map with dimensions (B, H, W, C), where H and W are the resolutions of the BEV map, and C is the number of feature channels. This feature map will be used as the input for the subsequent detection module.

[0080] Organize the anchor boxes (Anchor) representing the target pose and size and the feature vectors (Feature) encoding the target features. This part of the model mainly completes the structured organization of target representation, including the anchor boxes (Anchor) of the target pose and size and the feature vectors (Feature) encoding the target features. The anchor boxes are initialized in a uniformly distributed manner in the BEV space, and each anchor box is defined by its center coordinates (x, y), size ((w, l), and heading angle are defined. These parameters are set as learnable tensors and are continuously optimized during training through gradient descent to adapt to the geometric characteristics of different targets. At the same time, each target is represented by a 256-dimensional feature vector, which is initialized to all zeros and gradually learns the semantic information of the target during training. By combining the anchor boxes with the feature vectors, the model can simultaneously capture the geometric attributes and high-level semantic features of the target, providing rich context information for subsequent target detection and classification. This design not only enhances the adaptability of the model to complex scenarios but also improves the accuracy and robustness of target detection.

[0081] The self-attention mechanism for target features, the variable attention mechanism for extracting point cloud features based on PointPillars, the feed-forward network, the anchor box correction, and the class prediction module. They will be introduced separately as follows:

[0082] Self-attention mechanism for target features

[0083] The input of the self-attention mechanism is the set of target feature vectors generated in the third part , where each represents the feature vector of a target. The self-attention mechanism enhances the context information of the target features by calculating the global correlation between the feature vectors. The specific process is as follows:

[0084] First, for each feature vector , calculate the query (Query) , key (Key) , and value vector :

[0085]

[0086] where is a learnable weight matrix. Subsequently, calculate the attention score matrix :

[0087]

[0088] where is the dimension of the feature vector, i and j represent the indices of the target feature vector set, and exp(.) is the natural exponential function. Finally, obtain the enhanced feature vector by weighted aggregation of the value vectors:

[0089]

[0090] This mechanism enables each target feature to capture global context information and improves the modeling ability for complex scenes.

[0091] Variable attention mechanism for extracting point cloud features based on PointPillars

[0092] The input of the variable attention mechanism includes the BEV point cloud features and the target features output by the self-attention mechanism . This mechanism enhances the spatial information perception ability of the target features by dynamically sampling the point cloud features. The specific process is as follows:

[0093] For each target feature , predict a set of sampling point coordinates , where K is the number of sampling points. Subsequently, the sampling point features are obtained on the BEV feature map through the bilinear interpolation function Obtain sampling point features:

[0094]

[0095] And aggregate the sampling point features through attention weights:

[0096]

[0097] Where is the learnable attention weight. This mechanism enables the target features to adaptively focus on the key regions and improves the detection ability for non-rigid targets.

[0098] Feed-forward network (FFN)

[0099] The input of the feed-forward network is the target features output by the deformable attention mechanism . Nonlinear transformation of the target features is performed through a multi-layer fully connected network:

[0100]

[0101] Where is the learnable weight matrix, is the bias term, and RELU(.) is the rectified linear unit function; this module further extracts the high-order semantic features of the target and provides support for subsequent bounding box correction and class prediction.

[0102] Bounding box correction and class prediction module

[0103] The input of this module is the target features output by the feed-forward network , and the output is the corrected anchor box parameters, the target class probabilities, and the updated target features. The specific process is as follows:

[0104] Bounding box correction: For each target feature , predict the offset of the anchor box:

[0105]

[0106] And correct the anchor box through the following formula:

[0107]

[0108] Where are the parameters after anchor box decoding, are the corrected anchor box parameters.

[0109] Class prediction: Predict the target class probabilities through a fully connected layer and the Softmax function:

[0110]

[0111] Among them and are learnable weight matrices and bias terms.

[0112] The self-attention mechanism of the target feature, the deformable attention mechanism based on PointPillars for extracting point cloud features, the forward network, the anchor box correction, and the class prediction module are cascaded in 6 layers for iterative processing. The output of the last layer of the model, the target pose and the target class probability, are used as the final result.

[0113] S3. Train the target detection network, and use the combination of Focal Loss and SmoothL1 Loss as the loss function for training; specifically,

[0114] For the class loss function, use Focal Loss to alleviate the class imbalance problem. For the target box regression, use the SmoothL1 Loss function SmoothL1(.), as follows:

[0115]

[0116] Among them, is the Focal Loss value, is the SmoothL1 Loss value, is the total loss; is the probability that the anchor box prediction belongs to a certain type of target, is the Focal Loss hyperparameter, generally taking 0.25 and 2. is the difference between the predicted target bounding box and the ground truth. In this application, weights are used to weight the two types of losses.

[0117] S4. Use the trained network for target detection, post-process the network output, obtain the final result, and complete the BEV target detection task.

[0118] Corresponding to the foregoing embodiment of a millimeter-wave radar BEV target detection method based on a sparse attention mechanism, the present invention also provides an embodiment of a millimeter-wave radar BEV target detection device based on a sparse attention mechanism.

[0119] See Figure 3 , an embodiment of a millimeter-wave radar BEV target detection device provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement a millimeter-wave radar BEV target detection method in the foregoing embodiment.

[0120] An embodiment of a millimeter-wave radar BEV target detection device based on a sparse attention mechanism provided by the present invention can be applied to any device with data processing capabilities. Such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware, or a combination of software and hardware. Taking software implementation as an example, as a logically defined device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. At the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the millimeter-wave radar BEV target detection device provided by the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, generally according to the actual functions of the device with data processing capabilities where the embodiment device is located, other hardware may also be included, which will not be elaborated here.

[0121] The specific implementation process of the functions and roles of each unit in the above device can be found in detail in the implementation process of the corresponding steps in the above method, which will not be elaborated here.

[0122] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0123] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a millimeter-wave radar BEV target detection method based on a sparse attention mechanism in the above embodiment.

[0124] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.

[0125] The present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the method for millimeter-wave radar BEV target detection based on a sparse attention mechanism described above.

[0126] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.

[0127] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. The present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A millimeter-wave radar BEV object detection method based on a sparse attention mechanism, characterized in that The method includes the following steps: S1. Obtain the millimeter-wave radar point cloud and perform preprocessing; S2. Construct a millimeter-wave radar BEV object detection network, including: encoding the point cloud features based on the PointPillars architecture to obtain a feature map, and enhancing the features of the feature map based on the anchor box-guided self-attention mechanism; using a cascaded optimization module to iteratively process the enhanced features to obtain a prediction result; The cascaded optimization module is cascaded six layers in total, and each layer includes: Adopt the self-attention mechanism of the target features and the variable attention mechanism for extracting point cloud features based on PointPillars to perform variability sampling on the features; perform feature decoding and anchor box correction on the sampling results; S3. Train the object detection network, and use the combination of focal loss and smooth L1 loss as the loss function for training; S4. Use the trained network to perform object detection, post-process the network output, obtain the final result, and complete the BEV object detection task.

2. The millimeter-wave radar BEV target detection method based on the sparse attention mechanism according to claim 1, characterized in that The obtaining of the millimeter-wave radar point cloud and performing preprocessing includes: First, the millimeter-wave radar emits a frequency-modulated continuous wave and receives the echo signal, and uses the fast Fourier transform to extract the distance, speed, and azimuth angle information of the target; Generate the point cloud data containing the target coordinates and radar reflection intensity through multi-frame data accumulation and point cloud clustering; to adapt to the network input, the point cloud is normalized to a unified coordinate system, and the outlier points are removed; Subsequently, in order to facilitate the data processing of the deep learning model, the point cloud is regularized to a unified quantity scale.

3. A millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that The encoding of the point cloud features based on the PointPillars architecture to obtain a feature map is specifically as follows: First, use the PointPillars architecture to divide the point cloud into regular three-dimensional columnar grids, the spatial range of each columnar body is a fixed size on the plane, for each point in each columnar body, calculate its offset relative to the center of the columnar body, splice the original features with the offset to obtain an enhanced feature vector, then encode the enhanced point cloud features through a multi-layer perceptron, map the encoded columnar body features to a two-dimensional BEV grid according to their spatial positions to form an initial BEV feature map, and finally use a two-dimensional convolutional network to aggregate the BEV feature map to obtain a feature map.

4. A millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that, The enhancement of the features of the feature map based on the anchor box-guided self-attention mechanism is specifically as follows: Complete the structured organization of the target representation, including the anchor box of the target pose and size and the feature vector encoding the target features; the anchor boxes are initialized in a uniformly distributed manner in the BEV space, each anchor box is defined by its center coordinates, size, and heading angle, all of which are learnable tensors and are used to continuously optimize during the training process to adapt to the geometric characteristics of different targets; each target is represented by a 256-dimensional feature vector, and these feature vectors are initialized to all zeros and are used to gradually learn the semantic information of the target during the training process. Combine the anchor boxes with the feature vectors so that the model can capture both the geometric attributes and high-level semantic features of the target.

5. A millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that, The self-attention mechanism of the target feature enhances the context information of the target feature by calculating the global correlation between feature vectors. For each feature vector, query, key, and value vectors are calculated through a learnable weight matrix, and then the attention score matrix is calculated through the query vector and the key vector, and the enhanced feature vector is obtained by weighted aggregation of the value vectors. The formula for calculating the attention score matrix is: Among them, is the dimension of the feature vector, is the query vector corresponding to the i-th feature vector, is the key vector corresponding to the k-th feature vector, is the key vector corresponding to the j-th feature vector, and N is the number of feature vectors; exp(.) represents the natural exponential function; The enhanced feature vector is obtained by weighted aggregation of the value vectors: is the value vector corresponding to the j-th eigenvector.

6. The millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that The deformable attention mechanism for extracting point cloud features based on PointPillars includes: For each target feature output by the self-attention mechanism , a set of sampling point coordinates is predicted , where K is the number of sampling points. Subsequently, the sampling point features are obtained on the BEV feature map through the bilinear interpolation function : Among them is the BEV point cloud feature And the sampled point features are aggregated through the attention weights: wherein is a learnable attention weight; this mechanism enables the target features to adaptively focus on key regions and improves the detection ability for non-rigid targets.

7. A millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that, The specific steps of feature decoding and anchor box correction for the sampling result include: Target features output by the input-variable attention mechanism , perform a non-linear transformation on the target features through a multi-layer fully connected network: wherein is a learnable weight matrix, is a bias term; RELU(.) is a rectified linear unit function; Bounding box correction: For each target feature , predict the offset of the anchor box: And the anchor box is corrected by the following formula: wherein are the parameters after anchor box decoding; are the corrected anchor box parameters; Class prediction: predicting the target class probability through a fully connected layer and the Softmax function: wherein and are learnable weight matrices and bias terms.

8. A millimeter-wave radar BEV target detection method based on a sparse attention mechanism according to claim 1, characterized in that The combination of the focal loss and the SmoothL1(.) loss as the loss function for training specifically includes: Among them, is the focal loss value, is the smooth L1 loss value, is the total loss; is the probability that the anchor box prediction belongs to a certain type of target, is the focal loss hyperparameter, is the difference between the predicted target bounding box and the ground truth, is the anchor box parameter, and weights are used to weight the two types of losses.

9. A millimeter-wave radar BEV object detection device based on a sparse attention mechanism, comprising a memory and one or more processors, wherein executable code is stored in the memory, and is characterized in that, When the processor executes the executable code, it implements a millimeter-wave radar BEV target detection method based on a sparse attention mechanism as described in any one of claims 1-8.

10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a millimeter-wave radar BEV target detection method based on a sparse attention mechanism as described in any one of claims 1-8.

Citation Information

Patent Citations

  • 4D millimeter wave radar target detection, tracking and speed measurement method based on deep learning

    CN116403180A

  • Fusion 3D target detection method based on 4D millimeter wave radar and image

    CN117274749A

  • 3D target detection method based on image and 4D millimeter wave radar fusion

    CN117542010A

  • 4D millimeter wave radar 3D target detection method and system based on time sequence

    CN118279904A

  • Target detection method and device for multi-modal feature fusion in drive test scene

    CN119723270A