Class increment target detection method and device based on radar point cloud data
By converting radar point cloud data into pseudo-images and training teacher models, and combining radar scattered cross-section information for knowledge distillation, the problem of forgetting historical target detection capabilities in radar point cloud data is solved, and the target detection performance is significantly improved.
Patent Information
- Application Number
- CN202510146319.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing incremental target detection methods fail to make full use of radar-specific information, resulting in serious forgetting of historical target detection capabilities in radar point cloud data, especially in radar scattering cross-sections.
By converting radar point cloud data into pseudo-images, and using pseudo-images to train the teacher model, identifying the teacher's attention feature map, combining new class target annotation data to train the student model, and output prediction boxes and classification results.
It significantly enhances the detection ability of historical targets with smaller radar scattering cross-sections, ensuring that while learning new targets, the student model can maintain the detection efficiency of historical targets more stably, thereby improving the overall target detection performance.
Smart Images

Figure CN120198707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar and target detection, and particularly to a class-incremental target detection method and device based on radar point cloud data. Background Art
[0002] In the field of modern technology, the application scope of radar target detection technology is becoming increasingly wide, and it plays an important role in intelligent security, aerospace, autonomous driving, etc. In the actual application process, the categories of targets to be detected are often uncontrollable, and new categories that have never appeared before may occur. At this time, if the existing target detection model is trained and expanded using new category target data, a catastrophic forgetting problem will occur in which the detection performance of old category targets deteriorates sharply. Existing class-incremental target detection methods are all for optical image data, and information such as the radar cross section (RCS) unique to radar data is not fully utilized. How to combine the unique information of radar to improve the effect of incremental learning is a key issue in the field of radar target detection.
[0003] Class-incremental target detection is a technology that enables the target detection model to have the ability to detect new category targets while alleviating the catastrophic forgetting of old category targets. It is also the basis for related application fields of radar target detection such as the construction of intelligent decision-making systems, the optimization of multi-sensor collaborative perception, and the accurate analysis of environmental situations. Class-incremental target detection has important technical value. Compared with optical images, millimeter-wave radar point cloud data is sparse and irregular, and is easily affected by noise, resulting in the student model being difficult to inherit the knowledge from the teacher model well during the incremental learning process, and the detection performance is limited. Therefore, how to fully exploit the unique information of radar, overcome the catastrophic forgetting problem in the incremental learning process, and design a class-incremental target detection method suitable for radar point cloud data has important research significance.
[0004] Most of the existing incremental detection methods use knowledge distillation to distill the knowledge of detecting old category targets from the teacher network. The basic idea is as follows: First, a mature teacher model is trained using existing old category data, which contains knowledge such as in-depth feature understanding and classification logic of old category targets. Then, a student model is constructed; during the training process of the student model, on the one hand, it is allowed to contact new category target data and perform normal learning based on the new category data; on the other hand, a knowledge distillation mechanism is introduced to enable the student model to learn the old category knowledge mastered by the teacher model. Finally, the student model gradually has the effective detection ability for new and old categories. However, in the process of class-incremental target detection, the existing incremental detection methods only use the outputs at all levels of the network itself as the basis for knowledge distillation. In the case of radar point cloud data, most of the existing methods fail to fully consider the characteristics of millimeter-wave radar itself, and the detection ability for historical category targets with small radar cross sections is forgotten more seriously. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a class-incremental target detection method and device based on radar point cloud data, so as to solve the problem that most of the existing methods fail to fully consider the characteristics of millimeter-wave radar itself, and the detection ability for historical class targets with small radar cross-sections is severely forgotten.
[0006] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of the present invention provides a class-incremental target detection method based on radar point cloud data, including:
[0008] Convert the radar point cloud data into a pseudo-image, which is applicable to the region-based fast convolutional neural network model, and the pseudo-image includes channel data corresponding to the radar cross-section;
[0009] Use the pseudo-image to train the teacher model to obtain a trained teacher model, and the teacher model is a region-based fast convolutional neural network model;
[0010] Determine the teacher attention feature map according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each class of targets in the historical class target annotation data;
[0011] Use the new class target annotation data, the trained teacher model and the teacher attention feature map to train the student model to obtain a trained student model;
[0012] Input the new pseudo-image into the trained student model so that the trained student model outputs a prediction box and a classification result, and the new pseudo-image includes the latest class target annotation data, historical class target annotation data and new class target annotation data.
[0013] The second aspect of the present invention provides a class-incremental target detection device based on radar point cloud data, including:
[0014] A conversion module for converting radar point cloud data into a pseudo-image, which is applicable to the region-based fast convolutional neural network model, and the pseudo-image includes channel data corresponding to the radar cross-section;
[0015] A determination module for determining the teacher attention feature map according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each class of targets in the historical class target annotation data;
[0016] A final output module is used to input a new pseudo-image into the trained student model, so that the trained student model outputs prediction boxes and classification results. The new pseudo-image contains data with the latest class target annotations, data with historical class target annotations, and data with new class target annotations. The trained student model is a model obtained by training the student model using data with new class target annotations, the trained teacher model, and the teacher attention feature map. The trained teacher model is a model obtained by training the teacher model using pseudo-images. The teacher model is a region-based fast convolutional neural network model.
[0017] Compared with the prior art, the class-incremental object detection method and device based on radar point cloud data provided by the present invention convert radar point cloud data into pseudo-images; use the pseudo-images to train a teacher model to obtain a trained teacher model; determine a teacher attention feature map according to the channel data corresponding to the radar cross-section and multiple radar cross-sections of each class of objects in the data with historical class target annotations; use the data with new class target annotations, the trained teacher model, and the teacher attention feature map to train the student model to obtain a trained student model; input the new pseudo-image into the trained student model so that the trained student model outputs prediction boxes and classification results. In this way, the radar cross-section can be fused to carry out knowledge distillation. Compared with the prior art method that simply relies on the network output to implement knowledge distillation, it can more effectively fit the characteristics of the radar itself; significantly enhances the detection ability of historical class objects with small radar cross-sections during the class-incremental object detection process, ensuring that the trained student model can more stably maintain the detection efficiency of historical class objects while learning new class objects, thereby improving the overall object detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0019] Figure 1 Schematically shows a flowchart of a class-incremental object detection method based on radar point cloud data;
[0020] Figure 2 Schematically shows an algorithm flowchart of a class-incremental object detection method based on radar point cloud data;
[0021] Figure 3 Schematically shows a flowchart of determining a teacher attention feature map;
[0022] Figure 4Schematically shows a schematic diagram of the class incremental target detection method based on radar point cloud data of the present invention and the existing methods in terms of class target detection capabilities;
[0023] Figure 5 Schematically shows the structural diagram of the class incremental target detection device based on radar point cloud data. Detailed implementation manners
[0024] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully communicated to those skilled in the art.
[0025] It should be noted that: unless otherwise specified, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.
[0026] The method in the embodiments of the present invention will be described in detail below.
[0027] Figure 1 Schematically shows the flowchart of the class incremental target detection method based on radar point cloud data in the embodiments of the present invention. See Figure 1 As shown, the class incremental target detection method based on radar point cloud data may include:
[0028] S101. Convert the radar point cloud data into a pseudo-image.
[0029] Among them, the pseudo-image is applicable to the region-based fast convolutional neural network model (Faster Region-based Convolutional Neural Networks, Faster R-CNN), and the pseudo-image includes channel data corresponding to the radar cross section.
[0030] Specifically, converting the radar point cloud data into a pseudo-image includes:
[0031] Step A1: Divide the radar point cloud data on the horizontal plane according to the observation range to obtain grids, and project the coordinates of each point in the radar point cloud data to the coordinates of the corresponding grid to obtain a grid map.
[0032] To eliminate the error caused by the fact that the values of (x, y) of each point in the radar point cloud data often do not correspond one-to-one to the coordinates of the grid, the present invention quantifies the coordinates (x, y) of each point in the radar point cloud data according to the cell size of the grid and then projects each point in the radar point cloud data.
[0033] Step A2: Use multiple physical quantities of the radar point cloud data as channels of the grid map, and fill the channels into the corresponding grids to obtain the filled grid map.
[0034] Among them, the multiple physical quantities include the coordinates of each point, the distance of each point from the radar, the radial velocity of each point after ego-vehicle motion compensation, the azimuth angle of each point, and the radar cross section.
[0035] After projection, use the physical quantities of the coordinates (x, y) of each point, the distance of each point from the radar, the radial velocity of each point after ego-vehicle motion compensation, the azimuth angle of each point, and the radar cross section in the radar point cloud data as channels of the grid map, and fill the channels into the corresponding grids to obtain the filled grid map. If multiple points are projected into the same grid at the same time, the physical quantities of these points are averaged and then filled. In order to eliminate the deviation between the point cloud coordinate position and the grid position caused by the quantization of the (x, y) coordinates in Step A1, the true (x, y) coordinates of each point are also filled into the grid as channels.
[0036] Step A3: Normalize the multiple physical quantities in the filled grid map to obtain a pseudo-image.
[0037] Normalize the multiple physical quantities in the filled grid map to the range of 0 to 1. The expression of the pseudo-image is:
[0038]
[0039] Among them, X(·) is the pseudo-image, X P is the radar point cloud data, Norm is the normalization process, x is the abscissa of each point in the radar point cloud data, y is the ordinate of each point in the radar point cloud data, r is the distance of each point from the radar, a is the azimuth angle of each point, v is the radial velocity of each point after ego-vehicle motion compensation, σ is the radar cross section value, is to project the coordinates of each point in the radar point cloud data to the coordinates of the corresponding grid, that is, after quantifying the cell size of the grid, project each point in the radar point cloud data.
[0040] S102: Use the pseudo-image to train the teacher model to obtain the trained teacher model.
[0041] Among them, the teacher model is a region-based fast convolutional neural network model.
[0042] Specifically, Figure 2 Schematically shows the algorithm flowchart of the class-incremental object detection method based on radar point cloud data. Refer to Figure 2 as shown, using the pseudo-image to train the teacher model to obtain the trained teacher model, including:
[0043] Step B1: Input the pseudo-image into the feature extraction network so that the feature extraction network outputs a deep feature map.
[0044] Among them, the feature extraction network is a 16-layer Visual Geometry Group network (VGG16).
[0045] Step B2: Set anchor boxes for each pixel of the deep feature map, perform feature mapping on the deep feature map using one convolutional layer, and process the anchor boxes using two convolutional layers to obtain the positions of the anchor boxes and the class scores of the anchor boxes.
[0046] Anchor boxes with sizes of 2.2, 3, 4.5, and 7 pixels are predefined for each pixel of the deep feature map, and the aspect ratios of the anchor boxes set at each pixel are 1, 0.5, and 2. Therefore, a total of 12 anchor boxes with different sizes and aspect ratios are defined for each pixel of the deep feature map.
[0047] The position of the anchor box includes the center point coordinates, the height, and the width of the anchor box. The classes of the anchor box include two categories: background and foreground. If there is a target in the anchor box, it is foreground; if there is no target in the anchor box, it is background.
[0048] Step B3: Perform non-maximum suppression on the positions of the anchor boxes and the class scores of the anchor boxes to obtain predefined candidate boxes.
[0049] Perform non-maximum suppression (NMS) on the positions of the anchor boxes and the class scores of the anchor boxes, remove the boxes with redundant spatial positions among all the anchor boxes, and obtain predefined candidate boxes. The regions surrounded by the predefined candidate boxes are highly likely to be candidate regions containing the target.
[0050] Step B4: According to a grid of a predefined size, evenly divide the candidate regions of the predefined candidate boxes to obtain predefined candidate regions corresponding to each grid of the predefined size, and perform a max-pooling operation on the predefined candidate regions to obtain the pooled candidate regions.
[0051] The predefined size can be 7×7 or 6×6. There can be multiple predefined sizes, and no specific limitation is made on the predefined size here.
[0052] Exemplarily, use a 7×7 grid to evenly divide the candidate regions of each predefined candidate box, perform a max-pooling operation within each grid, and finally the sizes of the pooled candidate regions obtained are all 7×7.
[0053] Step B5: Input the pooled candidate regions into the fully connected layer so that the fully connected layer outputs the candidate positions corresponding to the pooled candidate regions and the candidate target classes to which they belong.
[0054] Step B6: Perform non-maximum suppression on the candidate positions and their corresponding candidate target categories to obtain the final candidate bounding boxes and the final classification results.
[0055] Through non-maximum suppression, the preset candidate bounding boxes with redundant spatial positions can be removed again to obtain the final candidate bounding boxes and the final classification results.
[0056] Step B7: Determine the classification loss function and the regression loss function based on the preset candidate bounding boxes, the final candidate bounding boxes, the final classification results, the true candidate bounding boxes, and the true classification results, and perform backpropagation on the classification loss function and the regression loss function to obtain the trained teacher model.
[0057] S103: Determine the teacher attention feature map according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each type of target in the historical class target annotation data.
[0058] Calculating the teacher attention feature map can be used for subsequent knowledge distillation.
[0059] Specifically, Figure 3 Schematically shows the flowchart for determining the teacher attention feature map. See Figure 3 As shown, determining the teacher attention feature map according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each type of target in the historical class target annotation data includes:
[0060] Step C1: Sequentially perform average pooling and downsampling on the channel data corresponding to the radar cross-section to obtain the downsampled data, and copy the downsampled data to each channel of the grid map to obtain the radar cross-section map.
[0061] Among them, the size of the radar cross-section map is the same as that of the deep feature map, and the radar cross-section map includes the radar cross-section values in each grid.
[0062] The size H×W of the channel data corresponding to the radar cross-section after average pooling is downsampled to the size H / 8×W / 8, that is, the size of the downsampled data is H / 8×W / 8. H is the height of the channel data corresponding to the radar cross-section after average pooling, and W is the width of the channel data corresponding to the radar cross-section after average pooling.
[0063] Step C2: Calculate the average value of the radar cross-sections of each type of target corresponding to according to the multiple radar cross-sections of each type of target in the historical class target annotation data, and determine the difference of the multiple radar cross-sections according to the radar cross-section values in each grid and the average value of the radar cross-sections of each type of target.
[0064] Specifically, the expression for the difference between multiple radar cross - sections is as follows:
[0065]
[0066] Among them, is the difference between the radar cross - section value in the grid with coordinates (i, j) and the average value of the radar cross - sections of the k - th type of target in the historical class target annotation data. RCSij is the radar cross - section value in the grid with coordinates (i, j). is the difference between the average value of the radar cross - sections of the k - th type of target in the historical class target annotation data. i is the abscissa of the grid, and j is the ordinate of the grid.
[0067] Step C3: Determine the attention feature weight map according to the differences between multiple radar cross - sections and hyperparameters.
[0068] Specifically, the attention feature weight map includes the weight value corresponding to the grid with coordinates (i, j). The expression for the weight value corresponding to the grid with coordinates (i, j) is:
[0069]
[0070] Among them, weight_mapij is the weight value corresponding to the grid with coordinates (i, j), and β is the hyperparameter. is the difference between the radar cross - section value in the grid with coordinates (i, j) and the average value of the radar cross - sections of the 1st type of target in the historical class target annotation data. is the difference between the radar cross - section value in the grid with coordinates (i, j) and the average value of the radar cross - sections of the 2nd type of target in the historical class target annotation data. is the difference between the radar cross - section value in the grid with coordinates (i, j) and the average value of the radar cross - sections of the k - th type of target in the historical class target annotation data. i is the abscissa of the grid, and j is the ordinate of the grid.
[0071] Combine the weight values corresponding to the grids of all coordinates to obtain the attention feature weight map.
[0072] Step C4: Perform pixel - by - pixel multiplication and weighting on the attention feature weight map and the deep feature map to obtain the teacher attention feature map.
[0073] Generate the attention feature weight map based on the radar cross - section, use the attention feature weight map to weight the deep feature map, and use the obtained teacher attention feature map for subsequent knowledge distillation, so as to improve the detection ability for historical categories, especially targets with small radar cross - sections.
[0074] S104. Train the student model using the data labeled with new class targets, the trained teacher model, and the teacher attention feature map to obtain the trained student model.
[0075] The trained student model has the ability to detect new class targets while retaining the teacher model's ability to detect historical class targets.
[0076] Specifically, Figure 2 Schematically shows the algorithm flowchart of the class incremental target detection method based on radar point cloud data. See Figure 2 As shown, training the student model using the data labeled with new class targets, the trained teacher model, and the teacher attention feature map to obtain the trained student model includes:
[0077] Step D1: Input the data labeled with new class targets into the feature extraction network so that the feature extraction network outputs a new deep feature map.
[0078] Step D2: According to the new deep feature map and the grid of the preset size, use the operations of non-maximum suppression, average division, max pooling, and fully connected layer to determine the final classification loss function and the final regression loss function.
[0079] Specifically, step D2 includes:
[0080] Step D21: Set a new anchor box for each pixel point of the new deep feature map, perform feature mapping on the new deep feature map using one convolutional layer, and process the new anchor box using two convolutional layers to obtain the new position of the new anchor box and the new class score of the new anchor box.
[0081] Step D22: Perform the operation of non-maximum suppression on the new position of the new anchor box and the new class score of the new anchor box to obtain a new preset candidate box.
[0082] Step D23: According to the grid of the preset size, evenly divide the candidate areas of the new preset candidate box to obtain new preset candidate areas corresponding to each grid of the preset size, and perform max pooling operation on the new preset candidate areas to obtain the pooled new candidate areas.
[0083] Step D24: Input the pooled new candidate areas into the fully connected layer so that the fully connected layer outputs the new candidate positions corresponding to the pooled new candidate areas and the new candidate target classes to which they belong.
[0084] Step D25: Perform the operation of non-maximum suppression on the new candidate positions and the new candidate target classes to which they belong to obtain new candidate boxes and new classification results.
[0085] Step D26: Determine the final classification loss function and the final regression loss function based on the new preset candidate boxes, new candidate boxes, new classification results, true candidate boxes, and true classification results.
[0086] Step D3: Set the final anchor boxes for each pixel of the teacher attention feature map and the new deep feature map, use one layer of convolution to perform feature mapping on the teacher attention feature map and the new deep feature map respectively, and use two convolutional layers to process the final anchor boxes to obtain the positions of the final anchor boxes and the class scores of the final anchor boxes.
[0087] Step D4: Perform non-maximum suppression on the positions of the final anchor boxes and the class scores of the final anchor boxes to obtain intermediate candidate boxes and intermediate class scores.
[0088] Among them, the intermediate class scores include the first foreground score of the trained teacher model and the second foreground score of the student model, and the intermediate candidate boxes include the first intermediate candidate boxes of the trained teacher model and the second intermediate candidate boxes of the student model.
[0089] Step D5: Determine the total loss function for candidate region distillation in the intermediate candidate boxes based on the first foreground score, the second foreground score, the first intermediate candidate boxes, and the second intermediate candidate boxes.
[0090] Steps D3 to D5 are for candidate region distillation to preserve the ability to detect historical class objects at the feature level.
[0091] The total loss function for candidate region distillation in the intermediate candidate boxes is the sum of the candidate region classification distillation loss function and the candidate region regression distillation loss function in the intermediate candidate boxes.
[0092] Let the first foreground score of the trained teacher model be \(S_t\), the second foreground score of the student model be \(S_s\), the information of the first intermediate candidate box of the trained teacher model be \(B_t=(x_t,y_t,w_t,h_t)\), and the information of the second intermediate candidate box of the student model be \(B_s=(x_s,y_s,w_s,h_s)\), where \(B_t\) is the information of the first intermediate candidate box of the trained teacher model, \(x_t\) is the abscissa of the first intermediate candidate box of the trained teacher model, \(y_t\) is the ordinate of the first intermediate candidate box of the trained teacher model, \(w_t\) is the width of the first intermediate candidate box of the trained teacher model, \(h_t\) is the height of the first intermediate candidate box of the trained teacher model, \(B_s\) is the information of the second intermediate candidate box of the student model, \(x_s\) is the abscissa of the second intermediate candidate box of the student model, \(y_s\) is the ordinate of the second intermediate candidate box of the student model, \(w_s\) is the width of the second intermediate candidate box of the student model, and \(h_s\) is the height of the second intermediate candidate box of the student model. Let the candidate region set of the first intermediate candidate box of the trained teacher model and the candidate region set of the second intermediate candidate box of the student model, that is, the candidate region set of the intermediate candidate boxes of the teacher model and the student model, be \(C\), and the number of elements in the set \(C\) be \(n\). The candidate region classification distillation loss \(L\) in the intermediate candidate box CRCD is calculated as follows:
[0093]
[0094] where \(L\) CRCD is the candidate region classification distillation loss in the intermediate candidate box, \(C\) is the candidate region set of the intermediate candidate boxes of the teacher model and the student model, \(i'\) is an element in the set \(C\), is the first foreground score of the trained teacher model in the candidate region of the \(i'\)-th intermediate candidate box, is the second foreground score of the student model in the candidate region of the \(i'\)-th intermediate candidate box, and \(n\) is the number of elements in the set \(C\).
[0095] The candidate region regression distillation loss function \(L\) in the intermediate candidate box CRRD is calculated as follows:
[0096]
[0097] where \(L\) CRRD is the candidate region regression distillation loss function in the intermediate candidate box, \(C\) is the candidate region set of the intermediate candidate boxes of the teacher model and the student model, \(i'\) is an element in the set \(C\), is the abscissa of the \(i'\)-th intermediate candidate box of the trained teacher model, is the abscissa of the \(i'\)-th intermediate candidate box of the student model, is the ordinate of the \(i'\)-th intermediate candidate box of the trained teacher model, is the vertical coordinate of the \(i'\)-th intermediate candidate box of the student model, is the width of the \(i'\)-th intermediate candidate box of the trained teacher model, is the width of the \(i'\)-th intermediate candidate box of the student model, is the height of the \(i'\)-th intermediate candidate box of the trained teacher model, is the height of the \(i'\)-th intermediate candidate box of the student model, is the first foreground score of the candidate region of the \(i'\)-th intermediate candidate box of the trained teacher model, is the second foreground score of the candidate region of the \(i'\)-th intermediate candidate box of the student model, and \(n\) is the number of elements in set \(C\).
[0098] The calculation formula of the total loss function for candidate region distillation in the intermediate candidate box is:
[0099] \(L\) CRTD \(=\) \(L\) CRCD \(+\) \(L\) CRRD ;
[0100] where, \(L\) CRTD is the total loss function for candidate region distillation in the intermediate candidate box, \(L\) CRCD is the candidate region classification distillation loss in the intermediate candidate box, \(L\) CRRD is the candidate region regression distillation loss function in the intermediate candidate box.
[0101] Step D6: According to the candidate regions in the intermediate candidate boxes and the grids of the preset size, use the operations of average division, max pooling, fully connected layer, and non-maximum suppression to determine the final candidate boxes and the final classification results.
[0102] Step D6 is to perform top-level distillation and preserve the ability to detect historical class targets at the classification level.
[0103] Among them, the final candidate boxes include the first final candidate box and the second final candidate box.
[0104] Specifically, step D6 includes:
[0105] Step D61: According to the grids of the preset size, perform the operation of average division on the candidate regions in the intermediate candidate boxes to obtain the intermediate candidate regions corresponding to each grid of the preset size, and perform the max pooling operation on the intermediate candidate regions to obtain the pooled intermediate candidate regions.
[0106] Step D62: Input the pooled intermediate candidate regions into the fully connected layer so that the fully connected layer outputs the intermediate candidate positions and the corresponding intermediate candidate target categories of the pooled intermediate candidate regions.
[0107] Step D63: Perform non-maximum suppression on the intermediate candidate positions and their corresponding intermediate candidate target categories to obtain the final candidate boxes and the final classification results.
[0108] Step D7: Determine the total loss function of top-level distillation based on the final classification results, the number of classification heads, the final classification results, the number of predicted targets, the first final candidate boxes, and the second final candidate boxes.
[0109] The total loss function of top-level distillation is the sum of the top-level distillation classification loss function and the top-level distillation regression loss function.
[0110] Let the final classification results be C′ be the number of classification heads, c′ be the c′-th classification head, p be the p-th predicted target. For the trained teacher model and student model, calculate the corresponding normalization factors according to the following formulas respectively:
[0111]
[0112] where Dp is the normalization factor of the trained teacher model and student model, is the final classification result of the p-th predicted target under the c′-th classification head.
[0113] After obtaining the normalization factors of the trained teacher model and student model, adjust the output O of the trained teacher model and student model according to the following formula for N types of historical target classes:
[0114]
[0115] where Op is the output of the p-th predicted target of the trained teacher model and student model after adjustment under the historical target class, is the final classification result of the p-th predicted target under the historical target classes from the 1st to the Nth, and Dp is the normalization factor of the trained teacher model and student model.
[0116] Adjust the output O of the trained teacher model and student model according to the following formula for background target classes and new target classes:
[0117]
[0118] where, is the output of the p-th predicted target of the trained teacher model and student model after adjustment under the background target classes and new target classes, C′ is the number of classification heads, N is the number of historical target class categories, is the final classification result of the p-th predicted target under the m-th background target class and new target class, and Dp is the normalization factor of the trained teacher model and student model.
[0119] The adjusted outputs of the trained teacher model and the student model are obtained according to the above formula, and the top-level distillation classification loss function is calculated according to the following formula:
[0120]
[0121] where L TDCL is the top-level distillation classification loss function, X′ is the number of prediction targets, p is the p-th prediction target, N is the number of historical class target categories, c′ is the c′-th classification head, is the output of the p-th prediction target of the trained teacher model after adjustment under the historical class target, is the output of the p-th prediction target of the student model after adjustment under the historical class target, is the output of the p-th prediction target of the trained teacher model after adjustment under the background class target and the new class target, is the output of the p-th prediction target of the student model after adjustment under the background class target and the new class target.
[0122] Let the final candidate box of the trained teacher model be Rt, the final candidate box of the student model for the historical class target be, N be the number of historical class target categories, and X′ be the number of prediction targets. The top-level distillation regression loss function is calculated according to the following formula:
[0123]
[0124] where L TDRL is the top-level distillation regression loss function, p is the p-th prediction target, C′ is the number of classification heads, is the final candidate box of the trained teacher model for the p-th prediction target of the q-th class, is the final candidate box of the student model for the p-th prediction target of the q-th class.
[0125] The expression of the total loss function of top-level distillation is:
[0126] L TDTL = L TDCL + L TDRL ;
[0127] where L TDTL is the total loss function of top-level distillation, L TDCL is the top-level distillation classification loss function, and L TDRL is the top-level distillation regression loss function.
[0128] Step D8: Determine the total loss function of the student model according to the final classification loss function, the final regression loss function, the total loss function of candidate region distillation, and the total loss function of top-level distillation.
[0129] The expression of the total loss function of the student model is as follows:
[0130] Ltotal = L F + L CRTD + L TDTL ;
[0131] Among them, Ltotal is the total loss function of the student model, L F is the final classification loss function and the final regression loss function, L CRTD is the total loss function of candidate region distillation in the intermediate candidate boxes, L TDTL is the total loss function of top-level distillation.
[0132] Determining the final classification loss function, the final regression loss function, the total loss function of candidate region distillation in the intermediate candidate boxes, and the total loss function of top-level distillation are executed simultaneously.
[0133] Step D9: Perform backpropagation on the total loss function of the student model to obtain the trained student model.
[0134] S105. Input the new pseudo-image into the trained student model so that the trained student model outputs a prediction box and a classification result.
[0135] Among them, the new pseudo-image includes data with the latest class object annotations, data with historical class object annotations, and data with new class object annotations.
[0136] Since the present invention makes full use of the radar cross-section information in the radar point cloud data for class incremental detection, it retains the detection ability for historical class objects, especially objects with a small radar cross-section, stronger than the existing methods, and improves the target detection performance.
[0137] The above effects of the present invention can be further illustrated by the following experiments.
[0138] Experimental conditions: This invention processes relevant tasks and trains models using the PyTorch deep learning framework developed by the Facebook AI Research team in the Python programming language on a central processing unit of Intel(R) Xeon(R) Gold 5117M CPU@2.00GHz and a Linux operating system. The data used for training and testing all come from the publicly available Radar Scenes dataset. To achieve the object detection task, two parts of preprocessing were performed on the Radar Scenes dataset. First, in the original Radar Scenes dataset, a group of pulse durations was 60 ms. In this invention, four groups of radar point cloud data with a total of 240 ms were used to generate a burst of radar point cloud data. Second, the Radar Scenes dataset contains 11 target categories, which were merged according to the official documentation of the Radar Scenes dataset to obtain five major categories: cars, large vehicles, two-wheeled vehicles, pedestrians, and crowds.
[0139] The methods compared in the experiment are as follows: One is the incremental detection method based on top-level distillation (Incremental Learning of Object Detectors without catastrophic forgetting, ILOD), denoted as ILOD in the experiment, and the other is the incremental detection method based on multi-level knowledge distillation (Incremental Learning for Object Detectors based on Faster rcnn, Faster ILOD), denoted as Faster ILOD in the experiment.
[0140] Experimental content: According to the specific implementation manner of the class incremental target detection method based on radar point cloud data of the present invention, first train a teacher model that can detect historical class targets, and then use the same data containing only new class annotations to train a student model to measure the detection ability of different class targets. In terms of setting the data division method, it is mainly divided into a single-stage incremental experiment mode and a multi-stage incremental experiment mode. The single-stage incremental experiment mode focuses on integrating and dividing the data at one time, and combines a specific amount of historical class data and new class data in the same stage; the multi-stage incremental experiment mode adopts a phased and progressive data division strategy. First, determine a certain amount of historical class data as the basis, and then introduce new class data in subsequent different stages in turn. The experimental measurement index is the F1 score (F1Score). The F1 score is the harmonic mean that comprehensively considers the precision and recall, and is used to comprehensively evaluate the model performance. Among them, the precision represents the proportion of samples that are truly positive among the samples predicted as positive by the model; the recall represents the proportion of all samples that are actually positive and are accurately predicted as positive by the model. The calculation formula of the F1 score is:
[0141]
[0142] Among them, F1 is the F1 score, Precision is the precision, and Recall is the recall.
[0143] Compare the F1 score index calculated by the class incremental target detection method based on radar point cloud data of the present invention with the indexes of the ILOD method and the Faster ILOD method. The single-stage incremental experiment results are shown in Table 1 and the multi-stage incremental experiment results are shown in Table 2.
[0144] Table 1 Single-stage incremental experiment results
[0145] Method Car Pedestrian Crowd Two-wheeled vehicle Large vehicle Overall The present invention 0.7445 0.513 0.6039 0.6152 0.6213 0.6483 ILOD 0.6804 0.4444 0.3644 0.3574 0.5121 0.4987 Faster ILOD 0.7192 0.4924 0.3132 0.3294 0.4941 0.5058
[0146] Table 2 Multi-stage incremental experiment results
[0147] Method Car Pedestrian Crowd Two-wheeled vehicle Large vehicle Overall The present invention 0.5514 0.4592 0.4658 0.4630 0.4841 0.4963 ILOD 0.3850 0.3178 0.3795 0.4603 0.4925 0.4015 Faster ILOD 0.3669 0.4299 0.4215 0.4380 0.4477 0.4138
[0148] As shown in Table 1, among them, cars, pedestrians, and crowds are set as historical class targets, while two-wheeled vehicles and large vehicles are new class targets. Table 2 also uses cars, pedestrians, and crowds as historical class targets, and two-wheeled vehicles and large vehicles as new class targets. However, in Table 2, an incremental method in batches is adopted, that is, cars first enter the incremental learning process, and then large vehicles are added later.
[0149] Through in-depth analysis of the experimental data, it can be seen that whether it is the single-stage incremental mode or the multi-stage incremental mode, the present invention, with its unique design, fully exploits and utilizes the radar cross-section information in the radar point cloud data. Compared with the other two methods, the present invention shows excellent performance in the detection ability of historical targets, can continuously and stably detect historical targets accurately, especially effectively improves the detection ability for targets with small radar cross-sections such as pedestrians, crowds, and two-wheeled vehicles, effectively avoids the decline of the detection ability of historical targets, and reflects the superiority of the method for class-incremental target detection based on radar point cloud data of the present invention. Figure 4 Schematically shows a schematic diagram of the method for class-incremental target detection based on radar point cloud data of the present invention and the existing methods in terms of the class target detection ability. Refer to Figure 4 As shown, it visually demonstrates the advantages of the method for class-incremental target detection based on radar point cloud data of the present invention compared with the other two methods in terms of detection performance. Among them, the red box is the real target, and the green box is the predicted box. It can be seen that the method of the present invention correctly classifies and regresses all targets, while the other two methods have misdetections and missed detections of targets to varying degrees.
[0150] Based on the above Figure 1 implementation method, it can be seen that in the embodiment of the present invention, the radar point cloud data is converted into a pseudo-image; a teacher model is trained using the pseudo-image to obtain a trained teacher model; according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each type of target in the data labeled with historical targets, a teacher attention feature map is determined; a student model is trained using the data labeled with new targets, the trained teacher model, and the teacher attention feature map to obtain a trained student model; the new pseudo-image is input into the trained student model so that the trained student model outputs a predicted box and a classification result. In this way, the radar cross-section can be fused to carry out knowledge distillation. Compared with the existing method of simply relying on the network output to implement knowledge distillation, it can more effectively fit the characteristics of the radar itself; significantly enhances the detection ability for historical targets with small radar cross-sections in the process of class-incremental target detection, ensures that the trained student model can more stably maintain the detection efficiency for historical targets while learning new targets, and thus improves the overall target detection performance.
[0151] Based on the same inventive concept, as an implementation of the above method for class-incremental target detection based on radar point cloud data, the embodiment of the present invention also provides a device for class-incremental target detection based on radar point cloud data. Figure 5 For the structure diagram of the device in the embodiment of the present invention, refer to Figure 5 As shown, the device for class-incremental target detection based on radar point cloud data may include:
[0152] The conversion module 501 is used to convert radar point cloud data into a pseudo-image, which is applicable to the region-based fast convolutional neural network model. The pseudo-image includes channel data corresponding to the radar cross-section.
[0153] The determination module 502 is used to determine the teacher attention feature map according to the channel data corresponding to the radar cross-section and the multiple radar cross-sections of each type of target in the historical class target annotation data.
[0154] The final output module 503 is used to input the new pseudo-image into the trained student model, so that the trained student model outputs the prediction box and the classification result. The new pseudo-image includes the data of the latest class target annotation, the data of the historical class target annotation, and the data of the new class target annotation. The trained student model is a model obtained by training the student model using the data of the new class target annotation, the trained teacher model, and the teacher attention feature map. The trained teacher model is a model obtained by training the teacher model using the pseudo-image. The teacher model is a region-based fast convolutional neural network model.
[0155] Specifically, the conversion module 501 is used to divide the radar point cloud data on the horizontal plane according to the observation range to obtain grids, project the coordinates of each point in the radar point cloud data to the coordinates of the corresponding grid to obtain a grid map; use multiple physical quantities of the radar point cloud data as the channels of the grid map, and fill the channels into the corresponding grids to obtain a filled grid map. The multiple physical quantities include the coordinates of each point, the distance of each point from the radar, the radial velocity of each point after the ego-vehicle motion compensation, the azimuth angle of each point, and the radar cross-section; perform normalization processing on the multiple physical quantities in the filled grid map to obtain a pseudo-image.
[0156] Specifically, the determination module 502 is used to perform average pooling and downsampling on the channel data corresponding to the radar cross-section in sequence to obtain the downsampled data, and copy the downsampled data to each channel of the grid map to obtain a radar cross-section map. The size of the radar cross-section map is the same as that of the deep feature map. The radar cross-section map includes the radar cross-section values in each grid; calculate the average value of the radar cross-sections of each type of target according to the multiple radar cross-sections of each type of target in the historical class target annotation data, and determine the difference between the multiple radar cross-sections according to the radar cross-section values in each grid and the average value of the radar cross-sections of each type of target; determine the attention feature weight map according to the difference between the multiple radar cross-sections and the hyperparameters; perform pixel-by-pixel multiplication weighting on the attention feature weight map and the deep feature map to obtain the teacher attention feature map.
[0157] It should be noted here that the description of the above embodiments of the class-incremental target detection device based on radar point cloud data is similar to the description of the above embodiments of the class-incremental target detection method based on radar point cloud data, and has beneficial effects similar to those of the embodiments of the class-incremental target detection method based on radar point cloud data. For the technical details not disclosed in the embodiments of the class-incremental target detection device based on radar point cloud data of the embodiments of the present invention, please refer to the description of the embodiments of the class-incremental target detection method based on radar point cloud data of the present invention for understanding.
[0158] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A quasi-incremental target detection method based on radar point cloud data, characterized in that: include: Converting radar point cloud data into a pseudo image, wherein the pseudo image is suitable for a region-based fast convolutional neural network model, and the pseudo image includes channel data corresponding to the radar scattering cross section; Using the pseudo image to train a teacher model to obtain a trained teacher model, wherein the teacher model is the region-based fast convolutional neural network model; Determine a teacher attention feature map according to the channel data corresponding to the radar cross section and multiple radar cross sections of each type of target in the data labeled by historical type targets; Training the student model using the labeled data of the new category target, the trained teacher model and the teacher attention feature map to obtain a trained student model; A new pseudo image is input into the trained student model so that the trained student model outputs a prediction box and a classification result, wherein the new pseudo image includes data labeled with the latest class target, data labeled with the historical class target, and data labeled with the new class target.
2. The quasi-incremental target detection method based on radar point cloud data according to claim 1, characterized in that: The step of converting radar point cloud data into a pseudo image comprises: According to the observation range, the radar point cloud data is divided on a horizontal plane to obtain a grid, and the coordinates of each point in the radar point cloud data are projected to the coordinates of the corresponding grid to obtain a grid map; Using multiple physical quantities of the radar point cloud data as channels of the grid map, and filling the channels into the corresponding grids to obtain a filled grid map, wherein the multiple physical quantities include the coordinates of each point, the distance of each point from the radar, the radial velocity of each point after self-vehicle motion compensation, the azimuth of each point, and the radar scattering cross section; The multiple physical quantities in the filled grid image are normalized to obtain the pseudo image.
3. The quasi-incremental target detection method based on radar point cloud data according to claim 2, characterized in that: The method of training the teacher model using the pseudo image to obtain a trained teacher model includes: Inputting the pseudo image into a feature extraction network so that the feature extraction network outputs a deep feature map, wherein the feature extraction network is a 16-layer network of a visual geometry group; An anchor frame is set for each pixel point of the deep feature map, a feature map is performed on the deep feature map using a layer of convolution, and the anchor frame is processed using two convolution layers to obtain the position of the anchor frame and the category score of the anchor frame; Performing non-maximum suppression on the position of the anchor frame and the category score of the anchor frame to obtain a preset candidate frame; According to grids of preset sizes, the candidate areas of the preset candidate frame are evenly divided to obtain preset candidate areas corresponding to grids of preset sizes, and a maximum pooling operation is performed on the preset candidate areas to obtain pooled candidate areas; Inputting the pooled candidate region into a fully connected layer, so that the fully connected layer outputs the candidate position corresponding to the pooled candidate region and the candidate target category to which it belongs; Performing non-maximum suppression on the candidate positions and the candidate target categories to which they belong, to obtain a final candidate box and a final classification result; According to the preset candidate box, the final candidate box, the final classification result, the real candidate box and the real classification result, a classification loss function and a regression loss function are determined, and the classification loss function and the regression loss function are back-propagated to obtain the trained teacher model.
4. The quasi-incremental target detection method based on radar point cloud data according to claim 3 is characterized in that: Determining the teacher's attention feature map according to the channel data corresponding to the radar cross section and multiple radar cross sections of each type of target in the data labeled by historical type targets includes: Performing average pooling and downsampling on the channel data corresponding to the radar cross section in sequence to obtain downsampled data, and copying the downsampled data to each channel of the grid map to obtain a radar cross section map, wherein the size of the radar cross section map is the same as that of the deep feature map, and the radar cross section map includes radar cross section values in each grid; Calculating an average value of the corresponding radar cross sections of each type of target according to multiple radar cross sections of each type of target in the data marked by the historical type of targets, and determining a difference between multiple radar cross sections according to the radar cross section values in each grid and the average value of the radar cross sections of each type of target; Determining an attention feature weight map according to the differences and hyperparameters of the multiple radar cross sections; The attention feature weight map and the deep feature map are weighted and multiplied pixel by pixel to obtain the teacher attention feature map.
5. The quasi-incremental target detection method based on radar point cloud data according to claim 4, characterized in that: The expression of the difference of the multiple radar cross sections is: in, is the difference between the radar cross section value in the grid with coordinates (i, j) and the average value of the radar cross section of the k-th target in the data marked by the historical class targets, RCSij is the radar cross section value in the grid with coordinates (i, j), is the difference between the average values of the radar cross sections of the kth target in the data marked for the historical targets, i is the abscissa of the grid, and j is the ordinate of the grid.
6. The quasi-incremental target detection method based on radar point cloud data according to claim 4, characterized in that: The attention feature weight map includes the weight value corresponding to the grid with coordinates (i, j), and the expression of the weight value corresponding to the grid with coordinates (i, j) is: Among them, weight_mapij is the weight value corresponding to the grid with coordinates (i, j), β is the hyperparameter, is the difference between the radar cross section value in the grid with coordinates (i, j) and the average value of the radar cross section of the first category target in the data marked by the historical category targets. is the difference between the radar cross section value in the grid with coordinates (i, j) and the average value of the radar cross section of the second type of target in the data marked by the historical type target. It is the difference between the radar cross section value in the grid with coordinates (i, j) and the average value of the radar cross section of the kth target in the data marked by the historical targets, i is the horizontal coordinate of the grid, and j is the vertical coordinate of the grid.
7. The quasi-incremental target detection method based on radar point cloud data according to claim 4, characterized in that: The method of training the student model using the data labeled with the new class target, the trained teacher model and the teacher attention feature map to obtain the trained student model includes: Inputting the data labeled by the new class object into the feature extraction network so that the feature extraction network outputs a new deep feature map; According to the new deep feature map and the grid of the preset size, a final classification loss function and a final regression loss function are determined by using a non-maximum suppression operation, an average partitioning operation, the maximum pooling operation and the fully connected layer; A final anchor frame is set for each pixel point of the teacher's attention feature map and the new deep feature map, and feature mapping is performed on the teacher's attention feature map and the new deep feature map respectively by using the one layer of convolution, and the final anchor frame is processed by using the two convolution layers to obtain the position of the final anchor frame and the category score of the final anchor frame; Performing non-maximum suppression on the position of the final anchor box and the category score of the final anchor box to obtain an intermediate candidate box and an intermediate category score, wherein the intermediate category score includes a first foreground score of the trained teacher model and a second foreground score of the student model, and the intermediate candidate box includes a first intermediate candidate box of the trained teacher model and a second intermediate candidate box of the student model; Determine a total loss function for distillation of a candidate region in the intermediate candidate frame according to the first foreground score, the second foreground score, the first intermediate candidate frame, and the second intermediate candidate frame; According to the candidate area in the intermediate candidate frame and the grid of the preset size, the average partitioning operation, the maximum pooling operation, the fully connected layer and the non-maximum suppression operation are used to determine the final candidate frame and the final classification result, wherein the final candidate frame includes a first final candidate frame and a second final candidate frame; Determine a total loss function for top-level distillation according to the final classification result, the number of classification heads, the final classification result, the number of predicted targets, the first final candidate box, and the second final candidate box; Determine the total loss function of the student model according to the final classification loss function and the final regression loss function, the total loss function of the candidate region distillation and the total loss function of the top layer distillation; The total loss function of the student model is back-propagated to obtain the trained student model.
8. The quasi-incremental target detection method based on radar point cloud data according to claim 7, characterized in that: The method of determining the final classification loss function and the final regression loss function according to the new deep feature map and the grid of the preset size by using the non-maximum suppression operation, the average partitioning operation, the maximum pooling operation and the fully connected layer includes: Setting a new anchor frame for each pixel point of the new deep feature map, performing feature mapping on the new deep feature map using the one layer of convolution, and processing the new anchor frame using the two convolution layers to obtain a new position of the new anchor frame and a new category score of the new anchor frame; Performing a non-maximum suppression operation on the new position of the new anchor frame and the new category score of the new anchor frame to obtain a new preset candidate frame; According to the grids of the preset size, the candidate areas of the new preset candidate frame are evenly divided to obtain new preset candidate areas corresponding to grids of each preset size, and a maximum pooling operation is performed on the new preset candidate areas to obtain new pooled candidate areas; Inputting the pooled new candidate region into the fully connected layer, so that the fully connected layer outputs a new candidate position corresponding to the pooled new candidate region and a new candidate target category to which it belongs; Performing a non-maximum suppression operation on the new candidate position and the new candidate target category to which it belongs, to obtain a new candidate box and a new classification result; The final classification loss function and the final regression loss function are determined according to the new preset candidate box, the new candidate box, the new classification result, the real candidate box and the real classification result.
9. The quasi-incremental target detection method based on radar point cloud data according to claim 7, characterized in that: The determining of the final candidate box and the final classification result according to the candidate area in the intermediate candidate box and the grid of the preset size by using the average division operation, the maximum pooling operation, the fully connected layer and the non-maximum suppression operation includes: According to the grids of the preset size, the candidate regions in the intermediate candidate frame are averagely divided to obtain intermediate candidate regions corresponding to grids of each preset size, and the intermediate candidate regions are maximum pooled to obtain pooled intermediate candidate regions; Inputting the pooled intermediate candidate region into the fully connected layer, so that the fully connected layer outputs the intermediate candidate position corresponding to the pooled intermediate candidate region and the intermediate candidate target category to which it belongs; The non-maximum suppression operation is performed on the intermediate candidate position and the intermediate candidate target category to which it belongs, to obtain the final candidate box and the final classification result.
10. A quasi-incremental target detection device based on radar point cloud data, characterized in that: include: A conversion module, used to convert radar point cloud data into a pseudo image, wherein the pseudo image is suitable for a region-based fast convolutional neural network model, and the pseudo image includes channel data corresponding to the radar scattering cross section; A determination module, used to determine the teacher's attention feature map according to the channel data corresponding to the radar cross section and multiple radar cross sections of each type of target in the data labeled by historical class targets; The final output module is used to input a new pseudo image into the trained student model so that the trained student model outputs a prediction box and a classification result. The new pseudo image contains data labeled with the latest class target, data labeled with the historical class target and data labeled with the new class target. The trained student model is a model obtained by training the student model using the data labeled with the new class target, the trained teacher model and the teacher attention feature map. The trained teacher model is a model obtained by training the teacher model using the pseudo image. The teacher model is the region-based fast convolutional neural network model.