A method for detecting radar RF image targets

The local and global features of radar radio frequency images are extracted by combining convolutional neural networks and Transformer methods, and the redundant targets are removed using a non-maximum suppression algorithm, which solves the problems of low accuracy and repeated targets in the prior art, and achieves higher precision radar radio frequency image target detection.

CN114842196BActive Publication Date: 2025-07-22NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210493562.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2025-07-22
Estimated Expiration
2042-05-07

AI Technical Summary

Technical Problem

The existing radar RF image target detection method only uses pure convolutional neural network model, and cannot effectively extract the global features of radar RF image, resulting in low detection accuracy and lack of effective post-processing methods to remove duplicate targets.

Method used

The local and global features of radar radio frequency images are extracted using a combination of convolutional neural network and Transformer, and redundant targets are removed through a non-maximum suppression algorithm based on heat map prediction. The specific steps include preprocessing, feature enhancement, model construction, training and post-processing.

Benefits of technology

It improves the accuracy of radar radio frequency image target detection, can effectively remove repeated predicted targets, and improves the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842196B_ABST
    Figure CN114842196B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for radar RF image target detection. For radar signals, first, preprocessing is performed to obtain RF images, then feature enhancement is carried out, and then a model combining a convolutional neural network and a Transformer is constructed and trained. Finally, the target detection result is obtained through a heatmap-based non-maximum suppression algorithm. The method of combining a convolutional neural network and a Transformer used in the present invention can extract local and global features of radar RF images and achieve good results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer image processing, relates to radar target detection technology, and specifically relates to a method for detecting radar radio frequency image targets. Background Art

[0002] In the field of computer vision, target detection is a very important task. Through target detection technology, a computer can identify objects to be recognized in an image. Currently, target detection technology has been widely applied in fields such as camera monitoring, autonomous driving, and robot navigation.

[0003] Since camera image data is easy to obtain, has high precision, and is easy to label, the main task of current target detection is camera image data. Researchers have proposed many methods for this, mainly divided into two categories: two-stage detectors and one-stage detectors. Two-stage detectors first select target candidate boxes in the image, and then classify and locate the candidate boxes; one-stage detectors usually directly regard the detection problem as a regression problem and predict image pixels as the categories of targets and bounding boxes.

[0004] Although camera image data has many advantages, it cannot perform target detection well in situations such as strong and weak light environments, rainy and foggy days, occlusion, and blur. More robust sensors and recognition technologies are also required in perception systems such as autonomous driving. Millimeter-wave radar can play a very good role in situations where camera data performs poorly. Therefore, it is very necessary to study the target detection method for millimeter-wave radar.

[0005] The radar radio frequency image obtained by performing fast Fourier transform on millimeter-wave radar data contains rich Doppler and object motion information and has good ability to recognize objects. The target detection method for radar radio frequency images has great application value. However, the target detection method for camera images cannot play a good role in radar radio frequency images. Therefore, it is necessary to propose a target detection method for radar radio frequency images.

[0006] Existing radar radio frequency image target detection methods often only use a pure convolutional neural network model with an encoder-decoder structure to perform target detection on radar radio frequency images and then directly output the results. The disadvantages of this method are first that it can only extract local features of radar radio frequency images and cannot extract global features of radar radio frequency images well; second, the directly output results contain repeatedly predicted targets, resulting in low accuracy of detection results. Therefore, there is still room for improvement in the performance of current radar radio frequency image target detection methods. Summary of the Invention

[0007] The problems to be solved by the present invention are as follows: The existing radar RF image target detection methods only use the model of pure convolutional neural network, which is not sufficient to obtain good accuracy in the radar RF image target detection task because the pure convolutional neural network cannot well capture the global features of the radar RF image; the existing radar RF image target detection methods lack a post-processing process, and a post-processing method that can effectively remove duplicate targets is needed.

[0008] The technical solution of the present invention is: A radar RF image target detection method uses a neural network to perform target detection on a radar RF image, extracts local and global features of the radar RF image by combining a convolutional neural network and a Transformer, and performs non-maximum suppression based on heat map prediction on the results to obtain the target detection result, including the following steps:

[0009] 1) Preprocess the frequency signal received by the radar to obtain a range-angle radar RF image;

[0010] 2) Perform feature enhancement processing on the radar RF image;

[0011] 3) Construct a model that combines a convolutional neural network and a Transformer for radar RF image target detection, including an encoder module, a Transformer module, and a decoder module:

[0012] 3.1) The encoder module consists of 3 3D convolutional layers of 9×5×5 and 3 multi-scale convolutional modules;

[0013] 3.2) The Transformer module has 6 encoding layers, and each encoding layer includes two sub-layers: a multi-head self-attention mechanism and a multi-layer perceptron. Each multi-head attention mechanism layer includes three vectors of dimension D: Q, K, and V. By calculating the dot product of Q and K and dividing by the scale coefficient to obtain the weight information corresponding to Q and K, and using softmax to normalize the weight function and perform weighted summation on V to obtain the attention value. Among them, the attention algorithm adopts to implement;

[0014] 3.3) The decoder module consists of 3 transposed convolutional layers of 3×6×6 and 1 convolutional layer of 9×5×5, and includes three skip connection structures;

[0015] 4) Set the initial training parameters, and the initial training parameters include the learning rate, the number of iterations, the peak threshold, and the target similarity threshold;

[0016] 5) Train the model combining the convolutional neural network and Transformer, use the trained detector for object detection, and employ a non-maximum suppression algorithm based on heatmap prediction to remove duplicate predictions of objects;

[0017] 6) Calculate whether the accuracy and recall rate of object detection on the validation set and test set meet the detection requirements. If not, set new initial parameters to retrain the model combining the convolutional neural network and Transformer until the detection requirements are met.

[0018] Further, step 1) specifically includes:

[0019] 1.1) Perform a distance fast Fourier transform on the radar signal;

[0020] 1.2) Estimate the distance of the radar signal processed in 1.1);

[0021] 1.3) Use a low-pass filter to remove high-frequency noise from the result processed in 1.2);

[0022] 1.4) Perform an angle fast Fourier transform on the signal processed in 1.3);

[0023] 1.5) Select the parts of the radar radio frequency image generated in 1.4) where the chirp frequencies in the millimeter-wave radar signal are 0, 64, 128, and 192 to form a frame of radar radio frequency image data with 4 chirps.

[0024] Further, in step 2), feature enhancement processing is performed on the radar radio frequency image, which is specifically implemented as follows: The convolutional part consists of a distance-angle convolutional layer and a temporal convolutional layer, and a temporal max pooling layer is used to simplify the radar radio frequency images of multiple chirps.

[0025] Further, the initial training parameters set in step 4) specifically include: Set 60 epochs; Set the batchsize to 32; Use the Adam optimizer, where the initial learning rate is 0.001, beta1 is 0.9, and beta2 is 0.999. The train-step used is 1, and the train-stride is 4.

[0026] Further, in step 5) when training the model combining the convolutional neural network and Transformer, the loss function for object regression is:

[0027]

[0028] Where, is the final loss, D represents the confidence map of the true annotation, represents the pixel index, cls represents the class label, and (i, j) represents the pixel index.

[0029] Furthermore, step 5) uses a non-maximum suppression algorithm based on heatmap prediction to remove duplicates of redundant targets. The calculation method is as follows: Input the model object detection result map screened by the confidence threshold. For the object detection results of the current frame, record the coordinates and confidence of the target points, and place the points in set P. Select the peak point p with the highest confidence in set P, remove it from set P, and add it to set P * . Calculate the point p * and the similarity S with the remaining points p i . Compare it with the set similarity threshold. If it is higher than the threshold, delete the point p i from set P. Loop to select the highest point from P and repeat the above process until P is empty, and retain the target points in P * . Among them, the similarity S calculation method of two target points is as follows:

[0030]

[0031] S is the similarity of two target points, L is the actual distance between the two points, and κ cls has a value for each category, indicating the scale size of the category.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] (1) The present invention uses a method combining a convolutional neural network and a Transformer to better extract the local and global features of radar RF images, which can improve the accuracy of radar RF image object detection;

[0034] (2) The present invention proposes a multi-scale convolution module, which uses multiple branches to extract the input feature map and uses residual connections to retain the features of the input part, and can better extract the multi-scale information of radar RF images;

[0035] (3) The present invention proposes a non-maximum suppression method based on heatmap prediction, which can better remove redundant prediction targets of the detection model and make the detection results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a flowchart of the radar RF image object detection method based on Transformer proposed by the present invention.

[0037] Figure 2 is a specific schematic diagram of the radar RF image object detection method based on Transformer proposed by the present invention.

[0038] Figure 3 This is the visualization result of the radar RF image after the range - angle fast Fourier transform of the radar data of the present invention.

[0039] Figure 4 This is a schematic diagram of the feature enhancement module proposed by the present invention.

[0040] Figure 5 This is the structure diagram of the detection model proposed by the present invention.

[0041] Figure 6 This is the multi - scale convolution module in the detection model proposed by the present invention.

[0042] Figure 7 This is a schematic diagram of the detection result of the radar RF image target detection method based on Transformer of the present invention and the related camera scene. Detailed implementation manners

[0043] The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings.

[0044] As Figure 1 and Figure 2 shown, the radar RF image target detection method based on Transformer proposed by the present invention includes the following steps:

[0045] Step S1: Pre - process the signal received by the radar to obtain the radar RF image.

[0046] In the embodiment of the present invention, step S1 specifically includes the following steps:

[0047] Step S1.1: Perform range fast Fourier transform on the radar signal;

[0048] Step S1.2: Perform range estimation based on the result of step S1.1;

[0049] Step S1.3: Use a low - pass filter to remove high - frequency noise based on the result of step S1.2;

[0050] Step S1.4: Perform angle fast Fourier transform based on the result of step S1.3;

[0051] Step S1.5: Select the parts with chirp frequencies of 0, 64, 128, and 192 in the millimeter - wave radar signal from the radar RF image generated in step S1.4 to form a frame of radar RF image data with 4 chirps. As Figure 3 shown, it is the visualization effect of the radar RF image.

[0052] Step S2: Perform feature enhancement processing on the radar RF image. In the embodiment of the present invention, it is processed using a feature enhancement module, such as Figure 4 shown. The feature enhancement module consists of a distance-angle convolutional layer and a temporal convolutional layer to form the convolutional part, and simplifies the radar RF images of multiple chirps through a temporal max pooling layer.

[0053] Step S3: Construct a model that combines a convolutional neural network and a Transformer for radar RF image target detection. In the embodiment of the present invention, constructing a model that combines a convolutional neural network and a Transformer specifically includes:

[0054] Step S3.1: The encoder module consists of 3 3D convolutional layers of 9×5×5 and 3 multi-scale convolutional modules;

[0055] Step S3.2: The Transformer module has a total of 6 encoding layers. Each encoding layer includes two sub-layers: a multi-head self-attention mechanism and a multi-layer perceptron. Each multi-head attention mechanism layer includes three vectors of dimension D: Q, K, and V. By calculating the dot product of Q and K and dividing by the scale factor to obtain the weight information corresponding to Q and K, using softmax to normalize the weight function and perform weighted summation on V to obtain the attention value. Among them, the attention algorithm is adopted to implement;

[0056] Step S3.3: The decoder module consists of 3 transposed convolutional layers of 3×6×6 and 1 convolutional layer of 9×5×5, and contains three skip connection structures.

[0057] In the example of the present invention, the constructed model that combines a convolutional neural network and a Transformer is as Figure 5 shown, which combines the convolutional neural network and the Transformer module, and can better extract the local and global features of the radar RF image. Among them, the multi-scale convolutional module is as Figure 6 shown, using multiple branches and residual connections, and can well extract the multi-scale information of the features.

[0058] Step S4: Set the initial training parameters of the model as follows:

[0059] Set 60 epochs; set batchsize to 32; use the Adam optimizer, with an initial learning rate of 0.001, beta1 of 0.9, and beta2 of 0.999; set train-step to 1 and train-stride to 4.

[0060] Step S5: Train the model combining the convolutional neural network and Transformer, use the trained detector for object detection, and adopt the non-maximum suppression algorithm based on heatmap prediction to remove duplicates of repeatedly predicted objects. The object regression loss function used for training the model is:

[0061]

[0062] where is the final loss, D represents the confidence map of the true annotation, represents the pixel index, cls represents the class label, and (i, j) represents the pixel index.

[0063] In step S5, the non-maximum suppression algorithm based on heatmap prediction is adopted to remove duplicates of redundant objects, and the calculation method is as follows:

[0064] Input the object detection result map of the model screened by the confidence threshold. For the object detection result of the current frame, record the coordinates and confidence of the object points, and place the points in the set P. Select the peak point p with the highest confidence in the set P, remove it from the set P, and add it to the set P * . Calculate the similarity S between this point p * and the remaining points p i . Compare it with the set similarity threshold. If it is higher than the threshold, delete the point p i from the set P, and loop to select the highest point from P and repeat the above process until P is empty, and keep the object points in P * .

[0065] The calculation method of the similarity S of the object points is as follows:

[0066]

[0067] where S is the similarity between two object points, L is the actual distance between the two points, and κ cls has a value for each category, mainly referring to the scale size of this category, which can be specified empirically.

[0068] Step S6: Calculate whether the accuracy and recall rate of object detection on the validation set and the test set meet the detection requirements. If not, set new initialization parameters and retrain the model combining the convolutional neural network and Transformer until the detection requirements are met. The accuracy is defined as The recall rate is defined as where N TP is the number of true objects that are predicted as true objects, N FP is the number of false objects that are predicted as true objects, and N FN is the number of true objects that are predicted as false objects.

[0069] As Figure 7 shown, the target detection results of the radar RF image formed after using the non-maximum suppression algorithm based on heatmap prediction are as follows from top to bottom: the scene RGB image, the visualization of the radar RF image, the ground truth, and the prediction result. The detection is evaluated by accuracy and recall as follows: on the cruw dataset, the accuracy reaches 77.8% and the recall reaches 87.5%.

Claims

1. A radar radio frequency image target detection method, characterized in that Target detection is performed on radar RF images using a neural network. The local and global features of the radar RF images are extracted by combining a convolutional neural network and a Transformer, and a non-maximum suppression method based on heatmap prediction is used for the results to obtain the target detection results, including the following steps: 1) Preprocess the frequency signal received by the radar to obtain a range-angle radar RF image; 2) Perform feature enhancement processing on the radar RF image; 3) Construct a model that combines a convolutional neural network and a Transformer for target detection of radar RF images, including an encoder module, a Transformer module, and a decoder module: 3.1) The encoder module consists of 3 3D convolutional layers of 9×5×5 and 3 multi-scale convolutional modules; 3.2) The Transformer module has a total of 6 encoding layers. Each encoding layer includes two sub-layers: a multi-head self-attention mechanism and a multi-layer perceptron. Each multi-head attention mechanism layer includes three vectors of dimension D: Q, K, and V. By calculating the dot product of Q and K and dividing by the scale factor the weight information corresponding to Q and K is obtained. The softmax function is used to normalize the weight function and perform weighted summation on V to obtain the attention value. Among them, the attention algorithm is implemented using ; 3.3) The decoder module consists of 3 transposed convolutional layers of 3×6×6 and 1 convolutional layer of 9×5×5, which contains three skip connection structures; 4) Set the initial training parameters, where the initial training parameters include the learning rate, the number of iterations, the peak threshold, and the target similarity threshold; 5) Train the model that combines the convolutional neural network and the Transformer, use the trained detector for target detection, and use a non-maximum suppression algorithm based on heatmap prediction to remove duplicates of repeatedly predicted targets; 6) Calculate whether the accuracy and recall rate of target detection meet the detection requirements on the validation set and the test set. If not, set new initialization parameters and retrain the model that combines the convolutional neural network and the Transformer until the detection requirements are met.

2. The method for detecting a radar RF image target according to claim 1, wherein, Step 1) specifically includes: 1.1) Perform range fast Fourier transform on the radar signal; 1.2) Perform range estimation on the radar signal processed in 1.1); 1.3) Use a low-pass filter to remove high-frequency noise from the result processed in 1.2); 1.4) Perform angle fast Fourier transform on the signal processed in 1.3); 1.5) Select the parts with chirp frequencies of 0, 64, 128, and 192 in the millimeter-wave radar signal processed in 1.4) to form a frame of radar RF image data with 4 chirps.

3. A method for detecting radar RF image targets according to claim 1, characterized in that, Step 2) Perform feature enhancement processing on the radar RF image, and the specific implementation is: the convolutional part consists of a layer of range-angle convolutional layer and a layer of temporal convolutional layer, and a layer of temporal max pooling layer is used to simplify the radar RF image of multiple chirps.

4. A radar RF image target detection method according to claim 1, characterized in that, The initial training parameters set in step 4) specifically include: set 60 epochs; the batchsize is set to 32; use the Adam optimizer, where the initial learning rate is 0.001, beta1 is 0.9, beta2 is 0.999, and the train-step used is 1, train-stride is 4.

5. A radar RF image target detection method according to claim 1, characterized in that, In step 5) when training the model that combines the convolutional neural network and the Transformer, the loss function of target regression is: Among them, l is the final loss, D represents the confidence map of the true annotation, represents the pixel index, cls represents the class label, and (i, j) represents the pixel index.

6. A radar RF image target detection method according to claim 1, characterized in that In step 5), a non-maximum suppression algorithm based on heatmap prediction is used to remove redundant targets, and the calculation method is as follows: Input the model object detection result map screened by the confidence threshold. For the object detection result of the current frame, record the coordinates and confidence of the target points, and place the points in the set P. Select the peak point p with the highest confidence in the set P, remove it from the set P, and add it to the set P * . Calculate the point p * and the similarity S with the remaining points p i . Compare it with the set similarity threshold. If it is higher than the threshold, delete the point p from the set P i . Repeatedly select the highest point from P and repeat the above process until P is empty, and retain the target points in P * . Among them, the calculation method of the similarity S between two target points is as follows: S is the similarity between two target points, L is the actual distance between the two points, κ cls There is a numerical value for each category, indicating the scale size of that category.

7. A radar RF image target detection method according to claim 1, characterized in that In step 6), the accuracy rate is defined as The recall rate is defined as where N TP is the number of true targets predicted as true targets, and N FP is the number of false targets predicted as true targets, and N FN is the number of true targets predicted as false targets.

Citation Information

Patent Citations

  • Detection method and device of image

    CN107369154A

  • Medium-repetition-frequency radar target detection method based on attention mechanism

    CN113805151A