Method for accelerating 3D object detection through pruning and distillation
By combining pruning and distillation techniques in 3D object detection tasks, and using genetic algorithms to perform automatic channel search and knowledge distillation, the problem of slow inference speed of 3D object detection tasks in the existing technology is solved, and efficient model inference and accuracy maintenance are achieved.
Patent Information
- Application Number
- CN202510315488.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art is difficult to improve the inference speed while maintaining high accuracy in 3D object detection tasks, especially in autonomous driving applications, where the limitation of computing resources leads to high deployment costs and limited technology popularization.
Using a combination of pruning and distillation technology, the teacher model is pruned by constructing an automatic channel search strategy based on genetic algorithms, reducing the complexity and calculation amount of the model, and passing the teacher model knowledge to the student model through knowledge distillation, restoring the accuracy loss caused by pruning.
It realizes the significant reduction in model parameters and calculation amount while maintaining high detection accuracy and robustness, and improves inference speed, which is suitable for embedded devices and scenarios with limited computing resources.
Smart Images

Figure CN120147749A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method for accelerating 3D object detection by pruning and distillation. Background Art
[0002] In the field of autonomous driving, the 3D object detection task is crucial for vehicle perception systems, especially for object recognition and obstacle avoidance tasks in complex environments. The multi-modal fusion perception method based on cameras and lidar (LiDAR), which utilizes the visual information provided by cameras and the spatial distance information provided by lidar, can effectively improve the perception accuracy and robustness. However, although the multi-modal fusion method can provide high accuracy, its inference speed is usually slow, especially in autonomous driving applications that require real-time response, resulting in great challenges in deployment on embedded devices. This is mainly due to the limitation of computing resources, which makes the deployment cost relatively high and also restricts the popularization and application of this technology.
[0003] In the field of image classification, common model compression methods, such as pruning, quantization, and distillation, have been widely used to improve the inference speed, reduce the storage occupancy of the model, and attempt to find a balance between accuracy and speed. The pruning method reduces the redundant weights in the model and retains the most important parts for the final decision, thereby reducing the computational amount and improving the inference speed. In the 3D object detection task, although pruning can theoretically improve the inference speed, due to the high accuracy requirements of this task, the pruning method often brings a large loss of accuracy, especially when combining multi-modal data, and key feature information may be lost after pruning.
[0004] At the same time, as another compression technology, model distillation can effectively improve the performance of small models by transferring the knowledge of large teacher models to small student models, while reducing the computational complexity while maintaining the inference accuracy. However, in the field of 3D object detection, most efforts are focused on achieving higher detection accuracy, and model compression, especially the combination of pruning and distillation, is not common, which also limits the practical deployment and application of most models. Therefore, the present invention proposes a method for accelerating 3D object detection by pruning and distillation to solve the problems existing in the prior art. Summary of the Invention
[0005] Aiming at the above problems, the purpose of the present invention is to propose a method for accelerating 3D object detection by pruning and distillation. This method for accelerating 3D object detection by pruning and distillation can, through pruning and distillation techniques, while reducing the model complexity and computational amount, still maintain high detection accuracy and robustness, and can solve the problems existing in the prior art.
[0006] To achieve the objectives of the present invention, the present invention is implemented through the following technical solutions: A method for accelerating 3D object detection using pruning and distillation, comprising the following steps:
[0007] Step 1: Construct a well-trained multi-modal expert model
[0008] Obtain a trained multi-modal 3D object detection model based on the fusion of images and lidar as the teacher model, where the teacher model includes an image feature extraction network, a lidar point cloud feature extraction network, a multi-modal BEV feature fusion module, a shared BEV feature extraction module, and a 3D object detection head. Then, construct a dataset and perform data processing to obtain a processed dataset. Input it into the teacher model for feature extraction to obtain a feature map.
[0009] Step 2: Construct a module-aware genetic channel search pruning strategy
[0010] Construct a module-aware genetic channel search pruning strategy for each module of the teacher model hierarchically. According to the genetic algorithm, encode the channels of each module separately.
[0011] Step 3: Construct a multi-modal 3D student model according to the pruning strategy
[0012] Construct a 3D object detection model that fuses images and lidar point clouds as the student model. Select the optimal individual through Step 2. According to the selected optimal individual, that is, the optimal channel selection of the module, assign the weights of the optimal channels to the channels of the corresponding modules of the student model. For the other unselected channels, their weights are set to 0, which are the channels to be pruned. Take out the weights of the selected channels of each module in turn and assign them to the student model. Iterate the process in Step 2 to calculate the fitness value, and finally form a student model with the optimal channel number and structure.
[0013] Step 4: Construct a joint loss distillation strategy of the teacher model for the student model and perform training
[0014] Construct a joint loss distillation strategy of the teacher model for the student model, and obtain that the loss of knowledge distillation of the teacher model for the student model is the BEV feature loss L bev and the prediction result loss L head , and the overall loss L = L bev +L head . Thus, different knowledge is transmitted to the student model through multi-class joint losses, and then the student model is supervised and trained.
[0015] Step 5: Perform 3D object detection based on the student model after knowledge distillation
[0016] Based on pruning and distillation strategies, and a trained student model, perform 3D object detection tasks, where the method of data input and image processing and enhancement is the same as in Step 1.
[0017] A further improvement lies in that in Step 1, data processing includes image enhancement processing and point cloud data enhancement processing, and both the image enhancement processing and the point cloud data enhancement processing include cropping and rotation processing.
[0018] A further improvement lies in that in Step 1, the image feature extraction network includes an image feature encoding module and a class LSS perspective space conversion module.
[0019] A further improvement lies in that in Step 1, the lidar point cloud feature extraction network includes a point cloud voxel feature encoder and a BEV feature extraction module.
[0020] A further improvement lies in that in Step 2, the specific encoding method is as follows:
[0021] S1: Perform binary genotype encoding on the convolutional channels of specific modules to generate a population P N , where an individual in the population is
[0022] S2: Define the number of genetic iterations as iters times, and then for the initial population P N , calculate the fitness value V fit through the fitness function f fit . According to the fitness value, calculate the roulette probability, and use the roulette probability selection method to select K population individuals;
[0023] S3: Perform crossover and mutation operations of the genetic algorithm on the K population individuals with the best fitness values. Generate M 1 new individuals through crossover inheritance, generate M 2 new individuals through mutation inheritance, and use the inherited K + M 1 +M 2 individuals to update the population;
[0024] S4: Denote the population as Perform the next genetic iteration on the population until the fitness of an individual in the population reaches the expected value or the maximum number of iterations is reached.
[0025] A further improvement lies in that the fitness function f fit is defined as:
[0026]
[0027] Wherein, ACT is the sigmoid activation function with an output range of 0 to 1, and cls_heatmap ori and cls_heatmap cur are the class heatmaps before and after pruning respectively, and N cls is the number of classes in the object detection task.
[0028] The further improvement lies in: in the third step, the specific steps for constructing the BEV feature loss are as follows:
[0029] Sample the BEV features at the corresponding positions in the teacher model and the student model, which include the image BEV feature F cbev , the lidar point cloud BEV feature F lbev , and the multi-modal fusion BEV feature F fbev ; Use the ground truth box gt_bbox3d to generate masks of the same size as the above three BEV features Multiply the masks with each BEV feature point by point to form foreground-aware BEV features, thereby calculating the F cbev , F lbev , and F fbev feature losses of the teacher and student models, which are L cbev , L lbev , and L fbev respectively; The image BEV feature loss L cbev , the lidar BEV feature L lbev , and the multi-modal fusion BEV feature L fbev are added together to form the total BEV feature loss L bev .
[0030] The further improvement lies in: in the third step, the specific method for constructing the prediction result loss is as follows:
[0031] Extract the prediction results P T , P S of the teacher model and the student model S Calculate the loss L T by comparing the student model prediction P pred with the teacher model prediction P S ; Calculate the loss L gt by comparing the student model prediction P pred with the ground truth GT; The prediction loss L gt from the teacher model and the loss L head with the ground truth are added together to form the total prediction loss L
[0032] The beneficial effects of the present invention are:
[0033] (1) The present invention designs a pruning and distillation framework for 3D object detection tasks, effectively combining two model compression techniques to achieve efficient model inference. Compared with the traditional method of using pruning or distillation alone, the present invention can take both accuracy and inference speed into account.
[0034] (2) By constructing an automatic channel search strategy based on genetic algorithms, the present invention can prune the convolutional channels in the teacher model module by module. The genetic algorithm can adaptively select the optimal pruning strategy, automatically discover the channel combination most suitable for a specific task, maximize the accuracy of the pruned model, and effectively reduce the computational amount.
[0035] (3) The present invention uses knowledge distillation to transfer the rich knowledge of the teacher model to the student model. By generating a foreground mask, it focuses on learning local important features, helps the student model recover the losses caused by pruning, and at the same time learns deeper features.
[0036] (4) The present invention improves the model structure and performance. On the premise of ensuring high detection accuracy, it significantly reduces the number of model parameters and the computational amount. By reducing redundant parameters and the computational amount, the inference speed of the model is greatly improved, which makes this method particularly suitable for application in embedded devices and other scenarios with limited computing resources, thus promoting the popularization of 3D object detection technology in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a schematic diagram of the step flow of the present invention.
[0038] Figure 2 is a schematic diagram of the model structure and distillation of the present invention.
[0039] Figure 3 is a schematic diagram of the teacher model structure of the present invention.
[0040] Figure 4 is a schematic diagram of the pruning strategy process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] To deepen the understanding of the present invention, the following will further elaborate on the present invention in combination with embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0042] As can be seen from the background art, in the field of 3D object detection, most pursuits are for higher detection accuracy, and limited considerations are given to the actual deployment and application of most models. In particular, model compression combining pruning and distillation is not common.
[0043] Most of the current popular 3D object detection model compression technologies are single-modal knowledge distillation. Most of them use a powerful pure lidar or multi-modal teacher model to distill a pure vision student model with poor performance. Although many works have proved the effectiveness of the distillation method, the actual acceleration effect is still not good and still cannot meet the real-time operation requirements of embedded devices.
[0044] To solve the above problems, the present invention provides a method for accelerating 3D object detection using pruning and distillation, which is as follows:
[0045] According to Figure 1 As shown, this embodiment proposes a method for accelerating 3D object detection using pruning and distillation, including the following steps:
[0046] Step 1: Construct a well-trained multi-modal expert model
[0047] Obtain a trained multi-modal 3D object detection model based on the fusion of images and lidar as the teacher model, and then obtain the student model. In this embodiment, the student model is isomorphic to the teacher model, only the number of convolutional channels, the number of parameters, etc. are different. Then construct a data set and perform data processing. In this embodiment, the large public autonomous driving data set NuScenes is used, and 6 panoramic RGB images with a size of 1600×900×3 and 1 360-degree lidar point cloud map are loaded from it. The data processing includes image enhancement processing and point cloud data enhancement processing. Both the image enhancement processing and the point cloud data enhancement processing include cropping and rotation processing. Specifically, the cropping of image enhancement scales the RGB image with a size of 1600×900×3 to 704×256×3, and the cropping of point cloud data enhancement only extracts the point cloud data with a range of -54.0, -54.0, -5.0, 54.0, 54.0, 3.0, where the ranges in the X-axis and Y-axis directions are both -54.0, 54.0, and the range in the Z-axis direction is -5.0, 3.0. After that, the processed data is input into the teacher model to start feature extraction.
[0048] For the teacher model as Figure 2 and Figure 3 shown, it includes an image feature extraction network, a lidar point cloud feature extraction network, a multi-modal BEV feature fusion module, a shared BEV feature extraction module, and a 3D object detection head. Among them, the image feature extraction network includes an image feature encoding module and a class LSS perspective space transformation module (vtransform), and the lidar point cloud feature extraction network includes a point cloud voxel feature encoder and a BEV feature extraction module.
[0049] It should be emphasized that the teacher model in the embodiments of the present invention is a well-trained multi-modal 3D object detection model based on the fusion of images and lidar; the overall structure and feature extraction network of the teacher model are similar to the BEVFusion model; the teacher model extracts image features through an image feature encoding module, then converts the 2D image features into 3D geometric space features through a class LSS perspective space conversion module, and then obtains the image BEV features through BEV pooling; the teacher model obtains voxel features through a point cloud voxel feature encoder, and the BEV feature extraction module converts the voxel features into point cloud BEV features; the multi-modal BEV feature fusion module of the teacher model concatenates the image BEV features and the point cloud BEV features along the dimension, and fuses them through convolution operations to form multi-modal BEV features; the BEV features are sent to a shared BEV feature extraction module to extract deep BEV features, and these features are sent to a 3D object detection head; the output of the 3D object detection head includes a BEV feature classification heatmap, an object prediction classification result, and an object prediction regression box.
[0050] Step 2: Construct a genetic channel search pruning strategy based on module perception
[0051] Construct a genetic channel search pruning strategy based on module perception for each module of the teacher model, and encode the channels of each module separately according to the genetic algorithm.
[0052] According to the genetic algorithm, it is first necessary to encode the channels of each module separately. Obviously, encoding the channels of all modules simultaneously is unreasonable, which will lead to a huge search space, difficult search, and waste of resources such as GPU video memory. Therefore, it is particularly important to encode the channels module by module and layer by layer. According to the traditional iterative method of the genetic algorithm, the convolutional channels of various modules in the teacher model are respectively encoded with binary genotypes. The specific encoding method is as follows:
[0053] S1: For a specific module 1≤i≤N model , for the convolutional channels of are respectively encoded with binary genotypes to generate a population P N , where an individual in the population is Among them, is defined as:
[0054]
[0055] is a 0-1 binary encoding, representing the encoding of the j-th channel of the i-th individual in the population. represents that this channel will be pruned. represents that this channel will be retained.
[0056] Population P N is defined as:
[0057]
[0058] S2: Define the number of genetic iterations as iters times, and then for the initial population P N , through the fitness function f fit calculate the fitness value V fit , according to the fitness value, calculate the roulette probability, and use the roulette probability selection method to select K population individuals, where the fitness function f fit is defined as:
[0059]
[0060] In the formula, ACT is the activation function sigmoid with an output range of 0 to 1, cls_heatmap ori and cls_heatmap cur are the class heatmaps before and after pruning respectively, and N cls is the number of classes in the object detection task;
[0061] The roulette probability calculation formula is defined as:
[0062]
[0063] The genetic crossover operation formula is defined as:
[0064]
[0065] where mask cross is a random vector mask composed of 0 - 1, used to randomly select the positions of cross - inheritance of population individuals and to generate new individuals
[0066] The genetic mutation operation formula is defined as:
[0067]
[0068] where RandomMut() is a function to randomly flip the encoding bits, and m_p is the mutation flip probability;
[0069] S3: Perform the crossover and mutation operations of the genetic algorithm on the K population individuals with the best fitness values, generate M 1 new individuals through cross - inheritance, generate M 2 new individuals through mutation inheritance, where K + M 1 +M2 = N gen , using the inherited K + M 1 + M 2 updated populations, and denote the population as
[0070] S4: Denote the population as Perform the next genetic iteration on the population until the fitness of an individual in the population reaches the expected value or the maximum number of iterations is reached.
[0071] Step 3: Construct a multi-modal 3D student model according to the pruning strategy
[0072] Construct a 3D object detection model that fuses images and lidar point clouds as the student model. Select the optimal individual through Step 2. According to the selected optimal individual, that is, the optimal channel selection of the module, for a binary channel mask, that is, the population individual in Step 2 's formula. In the training stage, assign the weights of the channels of to the corresponding channels of the student model's modules. For the other unselected channels, their weights are set to 0 to test the fitness value; sequentially take out the weights of the selected channels of each module, assign them to the student model, iterate the process in Step 2, calculate the fitness value, and finally form a student model with the optimal channel number and structure;
[0073] Step 4: Construct a teacher model for the joint loss distillation strategy of the student model and perform training
[0074] Construct a teacher model for the joint loss distillation strategy of the student model, and obtain the loss of knowledge distillation of the teacher model to the student model as the BEV feature loss L bev and the prediction result loss L head , and the overall loss L = L bev + L head , thus transmitting different knowledge to the student model through the multi-class joint loss;
[0075] The specific steps to construct the BEV feature loss are as follows:
[0076] Sample the BEV features at the corresponding positions in the teacher model and the student model, which include the image BEV feature F chev , the lidar point cloud BEV feature F lbev , and the multi-modal fusion BEV feature F fbev ; use the ground truth box gt_bbox3d to generate a mask of the same size as the above three BEV features Multiply the mask with each BEV feature point by point to form a foreground-aware BEV feature, and thus calculate the F cbev , F lbev, F fbev Feature losses, namely L cbev , L lbev , L fbev ; The image BEV feature loss L cbev , the lidar BEV feature L lbev and the multi-modal fusion BEV feature L fbev are added together to form the total BEV feature loss L bev , that is, L bev = L cbev + L lbev + L fbev ;
[0077] Determine the BEV feature loss L bev as:
[0078]
[0079] Among them,
[0080] C, H, and W are the number of channels, height, and width of the feature map respectively;
[0081] is a foreground 0-1 mask matrix formed by mapping the ground truth box gt_bbox3d, used to filter out foreground feature targets;
[0082]
[0083] and are the image BEV features of the teacher and student models respectively;
[0084] and are the point cloud branch BEV features of the teacher and student models respectively;
[0085] and are the multi-modal fusion BEV features of the teacher and student models respectively; is the foreground loss L of the teacher image BEV feature map and the student model image BEV feature map cbev ;
[0086] is the foreground loss L of the teacher point cloud BEV feature map and the student model point cloud BEV feature map lbev ;
[0087] is the foreground loss L Foreground loss L of the student model's multi-modal BEV feature map ; fbev ;
[0088] Specifically, this loss calculation method is similar to the MSE loss function;
[0089] After that, construct the prediction result loss, and extract the prediction results P of the teacher model and the student model T , P S For the prediction P of the student model S and the prediction P of the teacher model T calculate the loss L pred ; For the prediction P of the student model S and the loss L with the ground truth GT gt ; The prediction loss L from the teacher model pred and the loss L with the ground truth gt are added together to form the total prediction loss L head , that is, L head =βL gt +(1 - β)L pred ;
[0090] Determine the total prediction loss L head as:
[0091]
[0092] where
[0093] β is the hyperparameter for balancing the detection loss, to balance the self-prediction loss L of the student model gt and the distillation loss L pred ;
[0094] x ∈ {cls, bbox, heatmap} are the prediction results of the student model, the prediction results of the teacher model, and the ground truth annotation information respectively;
[0095] FL, L1, GFL are the Focal Loss function FocalLoss for correcting sample imbalance, the L1 loss function L1Loss for penalized regression, and the Gaussian Focal Loss function GaussianFocalLoss based on the center range respectively;
[0096]
[0097] is the loss of the prediction result of the student model and the ground truth annotation information ;
[0098]
[0099] For the prediction results of the student model and the prediction results of the teacher model of the loss.
[0100] Step Five: Perform 3D object detection based on the student model after knowledge distillation
[0101] Based on the pruning and distillation strategies and the trained student model, perform the 3D object detection task, where the method of data input and image processing and enhancement is the same as in Step One.
[0102] Reduce the number of parameters and computational complexity of the multimodal model through the pruning strategy of module-aware genetic channel search, thereby forming a more compact student model. Combining joint loss knowledge distillation, transfer the knowledge of the teacher model to the student model, which not only improves the robustness of the student model, but also alleviates the accuracy loss caused by pruning and improves the detection accuracy of the student model.
[0103] The above-mentioned training processing module and process are carried out on the large-scale public autonomous driving dataset NuScenes, and the performance of the pruned student model is evaluated using the validation set of NuScenes, including indicators such as 3D detection accuracy, number of parameters, computational complexity, and inference time.
[0104] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the framework and scope of application of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for accelerating 3D object detection using pruning and distillation, characterized in that: The following steps are involved: Step 1: Build a well-trained multimodal expert model A mature multimodal 3D target detection model based on image and lidar fusion is obtained as a teacher model, where the teacher model includes an image feature extraction network, a lidar point cloud feature extraction network, a multimodal BEV feature fusion module, a shared BEV feature extraction module and a 3D target detection head. Then, a data set is constructed and processed to obtain a processed data set, which is input into the teacher model for feature extraction to obtain a feature map. Step 2: Construct a module-aware genetic channel search and pruning strategy A module-aware genetic channel search and pruning strategy is constructed hierarchically for each module of the teacher model, and the channels of each module are encoded separately according to the genetic algorithm; Step 3: Construct a multimodal 3D student model based on the pruning strategy Construct a 3D target detection model that integrates image and lidar point cloud as the student model. Select the best individual through step 2. According to the selected best individual, that is, the best channel of the module, assign the weight of the best channel to the channel of the corresponding module of the student model. The weights of other unselected channels are set to 0, which are the channels to be pruned. The weights of the selected channels of each module are taken out in turn and assigned to the student model. Iterate the process in step 2, calculate the fitness value, and finally form a student model with the best number of channels and structure. Step 4: Construct a joint loss distillation strategy for the teacher model and train it Construct a joint loss distillation strategy for the teacher model to the student model, and obtain the loss of knowledge distillation of the teacher model to the student model as BEV feature loss L bev And the prediction result loss L head , the total loss L = L bev +L head ,Thus, different knowledge is transferred to the student model through multi-class joint loss, and then the student model is supervised and trained; Step 5: 3D object detection based on the student model after knowledge distillation Based on the pruning and distillation strategies and the trained student model, the 3D object detection task is performed, where the data input and image processing and enhancement methods are the same as step one.
2. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In the step 1, the data processing includes image enhancement processing and point cloud data enhancement processing, and the image enhancement processing and point cloud data enhancement processing both include cropping and rotation processing.
3. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In the step 1, the image feature extraction network includes an image feature encoding module and an LSS-like viewing space conversion module.
4. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In the step 1, the laser radar point cloud feature extraction network includes a point cloud voxel feature encoder and a BEV feature extraction module.
5. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In the step 2, the specific encoding method is: S1: For specific modules The convolution channels of the genotypes are encoded in binary format to generate the population P N , where an individual in the population is S2: Define the number of genetic iterations as iters, and then N , through the fitness function f fit Calculate the fitness value V fit , according to the fitness value, calculate the roulette probability, and use the roulette probability selection method to select K population individuals; S3: Perform crossover and mutation operations of the genetic algorithm on the K population individuals with the best fitness values, generate M1 new individuals through crossover inheritance, generate M2 new individuals through mutation inheritance, and use the inherited K+M1+M2 individuals to update the population; S4: Let the population be The next genetic iteration is performed on the population until the fitness of the individuals in the population reaches the expected value or the maximum number of iterations is reached.
6. The method for accelerating 3D object detection using pruning and distillation according to claim 5, characterized in that: The fitness function f fit Defined as: Where ACT is the activation function sigmoid with an output range of 0 to 1, cls_heatmap ori and cls_heatmap cur They are the category heat maps before and after pruning, N cls is the number of categories for the target detection task.
7. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In step 3, the specific steps of constructing the BEV feature loss are: Sample the BEV features of the corresponding positions in the teacher model and the student model, which include the image BEV features F cbev 、LiDAR point cloud BEV feature F lbev , multi-modal fusion BEV feature F fbev ; Use the real box gt_bbox3d to generate a mask of the same size as the above three BEV features Mask Multiply each BEV feature point to form the foreground-aware BEV feature, thereby calculating the F of the teacher and student models cbev 、F lbev 、F fbev The feature loss is L cbev , L lbev , L fbev ; The image BEV feature loss L cbev 、LiDAR BEV Features L lbev and multimodal fusion BEV feature L fbev Add together to form the total BEV feature loss L bev .
8. The method for accelerating 3D object detection using pruning and distillation according to claim 1, characterized in that: In step 3, the specific method of constructing the prediction result loss is: Extract the prediction results P of the teacher model and the student model T ,P S The student's model predicts P S And the teacher model predicts P T Calculate the loss L pred ; The student's model predicts P S and the true result GT loss L gt ; The prediction loss L from the teacher model pred The sum and the loss L of the true result gt Add together to form the total prediction loss L head .
Citation Information
Patent Citations
BEV semantic segmentation model training method, system and equipment based on knowledge distillation and medium
CN115690416A
Model training method and device, electronic equipment and storage medium
CN116563807A
3D target detection system and method based on 4D millimeter wave radar and camera fusion
CN117452396A
Multi-modal biological feature recognition method based on third-order knowledge distillation
CN117831138A
Cited By
Depth forgery detection method and system for priori knowledge guidance and double-domain representation optimization
CN120748054A
Scenarized abnormal behavior classification method and system, equipment and storage medium
CN121412844A