Target detection knowledge distillation method based on feature fusion and global response

By adopting a knowledge distillation method based on feature fusion and global response in the object detection technology, combining feature distillation and response distillation, the problem of small object detection in complex scenarios is solved, and efficient small object detection performance is achieved.

CN120047735APending Publication Date: 2025-05-27BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117114.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When existing object detection technology deals with remote sensing images with complex scene background, large change in target scales and large numbers of small targets, it is difficult to effectively detect small targets. Due to hardware resource limitations, lightweight models have insufficient detection performance.

Method used

A method of distillation of object detection knowledge based on feature fusion and global response is proposed. By combining feature distillation and response distillation, the student model is guided by using the feature information of the teacher model to extract feature information that is conducive to the student model, and improve the accuracy and efficiency of small object detection.

Benefits of technology

It significantly improves the accuracy and efficiency of small object detection, is suitable for improving detection performance without changing the model structure, and effectively saves model computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047735A_ABST
    Figure CN120047735A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection knowledge distillation method based on feature fusion and global response. Two methods of feature distillation and student distillation are fused to realize feature information transfer from a teacher model to a student model. In the feature distillation part, firstly decoupling a foreground region and a background region, and then introducing a product-moment correlation coefficient to measure a linear relationship among different feature levels of the foreground region; in a distillation response part, a distillation condition screening mechanism is provided for probability distribution of model prediction information, and then a non-target class matching module and a frame probability module are respectively used in a classification branch and a regression branch to realize full transfer of teacher model prediction information, so that the detection performance of a lightweight model is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a target detection knowledge distillation method based on feature fusion and global response. Background Art

[0002] Due to its high precision and flexibility, target detection technology has become an important means of obtaining ground information and is widely used in fire monitoring, infectious disease prevention and control, precision agriculture and other fields.

[0003] However, for remote sensing images with complex scene backgrounds, large object scale variations, and mostly small objects, small objects usually occupy fewer pixels, have limited information, and are easily disturbed by background noise. In addition, there is an extremely unbalanced quantitative relationship between foreground objects and background objects, with the number of background objects far exceeding that of foreground objects. There is a large amount of noise information in the features, which increases the difficulty of detection.

[0004] In addition, due to the limitation of edge device hardware resources, most platforms can only deploy lightweight models. This further increases the difficulty of target detection tasks in detection scenarios with complex backgrounds and a large number of small targets. At present, although some studies have focused on improving the structure of target detection networks to enhance the ability to extract small target features, these methods usually require structural adjustments to existing models, which in turn affects the computational efficiency of the models and limits their deployment and use in practical applications. Therefore, how to improve the performance of target detection algorithms while maintaining computational efficiency has become an urgent problem to be solved.

[0005] Methods for lightweighting deep learning models can further reduce the number of model parameters and complexity while maintaining model accuracy, making them easier to deploy. These methods mainly include knowledge distillation, pruning, and quantization. However, pruning may lead to irregular model structures, and unstructured pruning is difficult to accelerate reasoning; while quantization can significantly reduce storage requirements, it may lead to loss of accuracy. Knowledge distillation avoids these shortcomings while maintaining performance, providing a more flexible and effective method for lightweighting models.

[0006] Knowledge distillation is a model compression method that transfers the prediction information of a complex teacher network to a lightweight student network. This method can significantly improve the performance of the student network without incurring additional computational overhead. At present, the existing knowledge distillation frameworks can be roughly divided into two categories according to the different locations of the distilled information transfer. The first category is response-based knowledge distillation, which was first applied to image classification tasks. Specifically, the student model learns the logarithmic distribution of the prediction layer of the teacher model and processes the probability distribution of the output of the teacher and student models through the Softmax function with temperature T. However, response-based knowledge distillation usually relies more on the output of the last layer of the neural network (such as soft labels), so the feature information of the intermediate layer of the teacher network model is difficult to be effectively utilized. The soft labels are derived from the classification probability distribution of the network, and their performance is limited by the learning effect of prior knowledge.

[0007] The second category is feature-based knowledge distillation. Since the target detection network usually contains more complex module relationships, this task is more difficult than the image classification task, which makes it more difficult to capture inter-class knowledge, which greatly increases the difficulty of knowledge transfer. Romero et al. first proposed to use the features of the middle layer of the teacher model to improve the prior knowledge training of the student model. The specific method is to improve the performance of the student network by making the student model imitate the output of the feature activation of the middle layer of the teacher model. Cao et al. proposed to use the Pearson correlation coefficient to simulate the corresponding related features, so as to focus on the associated information from the teacher model. Lee et al. proposed a knowledge distillation algorithm based on singular value decomposition to extract the potential knowledge in the feature map and pass it to the student model as the guiding information of the correlation between feature maps.

[0008] A knowledge distillation method for object detection based on feature fusion and global response is designed, which takes into account the extreme imbalance of the number of foreground and background in image features, and the extraction ability of teacher model features. Extracting feature information that is beneficial to the student model from the complex image features is the key to improving the generalization of student model distillation. Summary of the invention

[0009] In view of the detection difficulties such as redundant image background information, unbalanced target categories, and large changes in target scale, the main purpose of the present invention is to propose a target detection knowledge distillation method based on feature fusion and global response, combined with the feature-based and response-based knowledge distillation methods, to make full use of the feature information in the teacher model to guide the student model, to detect small targets in detection scenes with complex backgrounds and a large number of small targets, and to improve the accuracy and efficiency of small target detection.

[0010] To achieve the above object, the present invention provides a target detection knowledge distillation method based on feature fusion and global response, comprising the following steps:

[0011] Step 1: In the backbone network of the teacher and student models, extract the feature level information of different scales of the teacher model for the target to be detected;

[0012] Step 2: In the different feature level information obtained in step 1, further decouple the foreground area features and use the product-moment correlation coefficient to calculate the feature loss of the teacher model and the student model;

[0013] Step 3: Use the screening mechanism to select some of the teacher model’s results for response distillation guidance;

[0014] Step 4: In the network predictor part of the teacher and student models, for the classification branch of the object detection task, the NTCM module is used to normalize the non-target category probabilities of the teacher model and the student model;

[0015] Step 5: In the network predictor part of the teacher and student models, for the regression branch of the target detection task, the similarity between the position probabilities of the two bounding boxes is calculated by the positioning distillation method. The target positioning box position (x, y, w, h), i.e., the representation of the center point coordinates, width and height, is converted to (t, b, l, r), i.e., the distance from the sampling point to the top, bottom, left and right. The output position probability value of the teacher and student network detection head part is and Perform distillation;

[0016] Step 6: Calculate the feature distillation loss in step 2 and the response distillation loss in steps 4 and 5, and add them together to get the total loss of this method;

[0017] Step 7: Get the trained student model.

[0018] Furthermore, in step 1, feature information of l scales is selected from the existing feature level. Specifically, three feature maps of different scales are selected, where l∈{320×320, 80×80, 20×20}, and the corresponding intermediate feature maps of different resolutions are extracted and recorded as the teacher's features. and student characteristics

[0019] Furthermore, in step 2, the foreground features of the teacher model and the foreground features of the student model are decoupled by segmenting the mask, thereby generating a binary foreground mask M fg ∈{0,1} H×W To extract foreground features, use the product-moment correlation coefficient Model the linear relationship between foreground features after decoupling.

[0020] Furthermore, the foreground features of the teacher model are: The student model foreground features are:

[0021] Furthermore, the product-moment correlation coefficient is used to align the foreground feature relationship between the teacher model and the student model:

[0022] Get the loss of each level feature distillation

[0023] Furthermore, in step 3, the prediction results of the teacher model and the student model are calculated, and the IoU score (I t and I s ), which represents the intersection-over-union ratio of the predicted bounding box and the true box, and is used to measure the positioning accuracy. t >I s When the distillation step 4 is carried out.

[0024] Further, in step 4, T t Represents the output category probability of the teacher network, S t Represents the probability of the student network category, and the loss formula for calculating the target category is L target =-T t log(S t ). i Represents the probability of non-target categories output by the teacher model, S i represents the probability of non-target categories in the student model, T j and S j Represents the sum of target and non-target category probabilities output by the teacher model and the student model respectively;

[0025] Normalize the non-target category probabilities of the teacher model and the student model, and the formula is Among them, after the normalization operation, the loss formula of the non-target category is L non-target =-∑ i≠t N(T i / τ)log(N(S i / τ)), τ is the temperature parameter, which increases the smoothness of the distribution;

[0026] Combining the target category loss and the non-target category loss, we get a new classification loss formula L NTCM =ε 1 L target +ε 2 L non-target ε 1 , ε 2 Represent the weights of the distillation loss for the target category and non-target category, respectively.

[0027] Further, in step 5, the distillation loss formula of the response distillation BPM module is obtained as follows: in are four sides, τ is the temperature value

[0028] Further, in step 7, the overall response distillation loss is L r =δ 1 L NTCM +δ 2 L BPM , where δ 1 , δ 2 is a hyperparameter used to balance the loss.

[0029] Further, in step 6, the total loss of the entire distillation framework is L = L ori +λL f +βL r Among them, λ and β are hyperparameters used to balance the loss, L f It refers to the feature distillation loss related to the details of the distilled small target features, L r is the response distillation loss for global information, L ori is the original loss of the student model.

[0030] Compared with the existing methods, the method of the present invention has the following advantages:

[0031] 1. The present invention discloses a target detection knowledge distillation method based on feature fusion and global response, which significantly improves the detection effect of small targets. It not only effectively saves the space required for model calculation, but also improves the calculation efficiency.

[0032] 2. It has a non-target class matching module and a border probability module, which makes up for the lack of feature knowledge distillation in capturing correlation information between pixels, captures global information in the image, and has high detection accuracy.

[0033] 3. Suitable for small target detection, which can improve detection performance without changing the model structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A framework diagram of a target detection knowledge distillation method based on feature fusion and global response disclosed in the present invention;

[0035] Figure 2 is a flow chart of the screening mechanism of step 3;

[0036] Figure 3 This is a comparison of the training results of the VisDrone dataset using this distillation framework. DETAILED DESCRIPTION

[0037] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0039] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0040] The following combination Figure 1-Figure 3 The specific embodiments of the present invention are described in detail. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0041] like Figure 1 , Figure 2 , Figure 3 As shown, the present invention proposes a target detection knowledge distillation method based on feature fusion and global response, which integrates the two methods of feature distillation and student distillation to realize the feature information transfer from the teacher model to the student model. In the feature distillation part, the foreground area and the background area are first decoupled, and then the product-moment correlation coefficient is introduced to measure the linear relationship between different feature levels in the foreground area; in the response distillation part, a distillation condition screening mechanism is proposed for the probability distribution of the model prediction information, and then the non-target class matching module and the border probability module are used in the classification branch and the regression branch respectively to realize the full transfer of the teacher model prediction information, thereby comprehensively improving the detection performance of the lightweight model.

[0042] The invention discloses a method for object detection knowledge distillation based on feature fusion and global response, using two RTX2080Ti, based on Python and Pytorch deep learning API. In this embodiment, the RetinaNet-101 network is selected as the teacher model, and the RetinaNet-50 network is selected as the student model.

[0043] See also Figure 1 The present invention provides a target detection knowledge distillation method based on feature fusion and global response, comprising the following steps:

[0044] Step 1: Extract the feature level information of different scales of the teacher model.

[0045] After the backbone network of the teacher model and the student model, feature information of l scales is selected from the existing feature level. Specifically, three different scale feature maps are selected, where l∈{320×320, 80×80, 20×20}, and the corresponding intermediate feature maps of different resolutions are extracted and recorded as the teacher's features. and student characteristics

[0046] Step 2: Decouple the foreground region features and use the product-moment correlation coefficient to calculate the feature loss of the teacher model and the student model.

[0047] In order to focus on the salient area, the foreground features of the teacher model and the foreground features of the student model are decoupled by segmentation mask, and then a binary foreground mask M is generated. fg ∈{0,1} h×W To extract foreground features, use the product-moment correlation coefficient The linear relationship between foreground features after modeling decoupling is weakened, thereby weakening the impact of feature amplitude differences on the distillation process and solving the problems of insufficient small target features and inconsistent multi-scale features.

[0048] Teacher Model Foreground Features Student Model Prospect Characteristics The product-moment correlation coefficient is used to align the foreground feature relationships between the teacher model and the student model.

[0049] in,

[0050] Get the loss of each level feature distillation

[0051] This loss function focuses on the feature relationship information of the teacher and relaxes the constraint on the size of the student features.

[0052] The overall loss is the weighted sum of each scale: Among them, λ lis a hyperparameter used to balance the contributions of different scales. This method ensures that the student model can effectively align its multi-scale feature representation with the feature representation of the teacher model, especially for the foreground information, thereby improving the overall distillation performance.

[0053] Step 3: Use the screening mechanism to select some of the teacher model’s results for response distillation guidance.

[0054] The teacher model is not always better than the student model. Directly using the output of the teacher model may mislead the student model. Therefore, before performing response-based knowledge distillation, a screening mechanism is proposed to determine whether the teacher has sufficient guidance ability before distillation, and only perform response knowledge distillation when the conditions are met. Focus on effective prediction information transfer, improve the efficiency and reliability of knowledge distillation, and strengthen the student model's ability to distinguish complex targets and distractors.

[0055] After the network predictors of the teacher model and the student model, the prediction results of the teacher model and the student model are obtained, including the category to which the predicted target box belongs and the regression parameters of the target box. The area of ​​each target box is divided by the area of ​​the corresponding real box in the GT image to calculate the IoU score (I t and I s ), which represents the intersection-over-union ratio of the predicted bounding box and the true box, and is used to measure the positioning accuracy. t >I s When the distillation step 4 is performed, the specific screening mechanism is as follows Figure 2 shown.

[0056] Step 4: Design a non-target class matching module (NTCM module) and a bounding box probability module (BPM module) to capture the relationship between pixels in the feature map and help the student model obtain global context information; that is, in the feature distillation module, the detailed information of the teacher model and the student model for small targets is successfully extracted. In order to make up for the lack of the above-mentioned type of knowledge distillation in capturing the correlation information between pixels, a response-based distillation module is used to capture the global information in the image and use it to guide the training of the student model.

[0057] For the classification branch of the target detection task, the traditional teacher model and student model have inconsistent totals of non-target category probability distributions, which makes it difficult to align the distributions between non-target categories, thus limiting the distillation effect. In the network predictor part of the teacher model and the student model, the NTCM module is used to normalize the non-target category probabilities of the teacher model and the student model to ensure that they have the same total and better aligned distribution.

[0058] T t Represents the output category probability of the teacher network, S t Represents the probability of the student network category, and the loss formula for calculating the target category is Ltarget =-T t log(S t ). i Represents the probability of non-target categories output by the teacher model, S i represents the probability of non-target categories in the student model, T j and S j Represents the sum of target and non-target category probabilities output by the teacher model and the student model respectively.

[0059] Normalize the non-target category probabilities of the teacher model and the student model, and the formula is Among them, after the normalization operation, the loss formula of the non-target category is L non-target =-∑ i≠t N(T i / τ)log(N(S i / τ)), τ is the temperature parameter, which increases the smoothness of the distribution.

[0060] Combining the target category loss and the non-target category loss, we get a new classification loss formula L NTCM =ε 1 L target +ε 2 L non-target ε 1 , ε 2 Represent the weights of the distillation loss for the target category and non-target category, respectively.

[0061] Step 5: For the regression branch of the target detection task, the similarity between the position probabilities of the two bounding boxes is calculated by positioning distillation. The target positioning box position (x, y, w, h), i.e., the representation of the center point coordinates, width and height, is converted to (t, b, l, r), i.e., the distance from the sampling point to the top, bottom, left and right. The output position probability values ​​of the teacher and student network detection heads are calculated. and Distillation is performed.

[0062] The distillation loss formula of the response distillation BPM module is further obtained as follows: in are four sides, τ is the temperature value, and in this embodiment, τ=1.5.

[0063] Step 6: Calculate the feature distillation loss and response distillation loss, and add them together to get the total loss of this method.

[0064] After steps 4 and 5, the overall response distillation loss of the knowledge distillation method of the present invention is L r =δ 1 L NTCM +δ 2L BPM Among them, δ 1 , δ 2 is a hyperparameter used to balance the loss.

[0065] The total loss of the entire distillation framework is L = L ori +λL f +βL r Among them, λ and β are hyperparameters used to balance the loss, L f It refers to the feature distillation loss related to the details of the distilled small target features, L r is the response distillation loss for global information, L ori is the original loss of the student model.

[0066] Step 7: Get the trained student model.

[0067] Embodiment 2

[0068] Example 2 is a specific application of the target detection knowledge distillation method based on feature fusion and global response proposed in Example 1. The specific process is as follows: data set and experimental setting, selection of Adam optimizer, optimization of network training parameters by setting the above-mentioned loss calculation formula hyperparameters, selection of teacher model and student network, extraction of foreground areas in multiple levels of feature maps, design of non-target class matching module (NTCM module) and border probability module (BPM module) after screening mechanism, and finally obtaining the student model through training, and applying the student model to various small target detection tasks.

[0069] 1. Dataset and experimental settings.

[0070] The present invention uses the VisDrone dataset to illustrate the embodiment, uses the RetinaNet-101 network as the teacher model, and the RetinaNet-50 network as the student model. The setting used in the present invention is λ=5×10 -2 , β=2.5×10 -2 All models were optimized using the Adam optimizer for a total of 60 epochs, with momentum set to 0.9 and weight decay set to 0.0001.

[0071] 2. Extract foreground features of the teacher model and the student model.

[0072] like Figure 1 As shown in the figure, we focus on the target area of ​​the image, decouple the foreground features of the teacher model and the foreground features of the student model through segmentation masks, and use the product-moment correlation coefficient to model the linear relationship between the foreground features after decoupling, thereby weakening the impact of feature amplitude differences on the distillation process.

[0073] 3. Through the screening module, the part of the teacher model's prediction results that are better than the student model is passed to the student model.

[0074] The prediction results of the teacher model are not all better than those of the student model, such as Figure 2 As shown, the intersection and union ratios of the prediction results of the teacher model and the student model with the real box are calculated respectively. When the intersection and union ratio of the teacher model is greater than that of the student model, the next step of calculation is performed.

[0075] 4. Guide the classification branch of the student model through the NTCM module.

[0076] The non-target category probabilities of the teacher model and the student model are normalized to ensure that they have the same sum, better aligning the distribution of target and non-target probability values ​​in the teacher and student models.

[0077] 5. Guide the regression branch of the student model through the BPM module.

[0078] The target positioning box position representation method is converted into a representation method of the distance from the sampling point to the top, bottom, left, and right, and the similarity between the position probabilities of the two bounding boxes of the teacher model and the student model is calculated.

[0079] 6. Obtain the trained student model.

[0080] Some experimental results are as follows Figure 3 As shown, it can be seen that the model after distillation can detect more small targets than before distillation.

[0081] 7. Performance comparison.

[0082] The present invention compares the detection accuracy of the method of the present invention and the student model on the VisDrone unmanned aerial vehicle dataset. The results show that the detection accuracy of the method of the present invention is significantly improved, the AP value is increased by 2.8 compared with that before using the distillation method, and the detection performance of small targets is also improved.

[0083] The present invention uses the Pytorch deep learning framework and adopts the feature and response distillation method to distill the teacher model, which greatly improves the final accuracy of the student detection model, reduces the number of model parameters and improves the computational efficiency of the model. Therefore, it is more suitable for edge devices with more limited computing power such as drone equipment.

Claims

1. A knowledge distillation method for object detection based on feature fusion and global response, characterized in that: The following steps are involved: Step 1: In the backbone network of the teacher and student models, extract the feature level information of different scales of the teacher model for the target to be detected; Step 2: In the different feature level information obtained in step 1, further decouple the foreground area features and use the product-moment correlation coefficient to calculate the feature loss of the teacher model and the student model; Step 3: Use the screening mechanism to select some of the teacher model’s results for response distillation guidance; Step 4: In the network predictor part of the teacher and student models, for the classification branch of the object detection task, the NTCM module is used to normalize the non-target category probabilities of the teacher model and the student model; Step 5: In the network predictor part of the teacher and student models, for the regression branch of the target detection task, the similarity between the position probabilities of the two bounding boxes is calculated by the positioning distillation method. The target positioning box position (x, y, w, h), i.e., the representation of the center point coordinates, width and height, is converted to (t, b, l, r), i.e., the distance from the sampling point to the top, bottom, left and right. The output position probability value of the teacher and student network detection head part is and Perform distillation; Step 6: Calculate the feature distillation loss in step 2 and the response distillation loss in steps 4 and 5, and add them together to get the total loss of this method; Step 7: Get the trained student model.

2. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 1, feature information of l scales is selected from the existing feature level. Specifically, three feature maps of different scales are selected, where l∈{320×320, 80×80, 20×20}, and the corresponding intermediate feature maps of different resolutions are extracted and recorded as the teacher's features. and student characteristics 3. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 2, the foreground features of the teacher model and the foreground features of the student model are decoupled by segmenting the mask, thereby generating a binary foreground mask M fg ∈{0,1} H×W To extract foreground features, use the product-moment correlation coefficient Model the linear relationship between foreground features after decoupling.

4. The object detection knowledge distillation method based on feature fusion and global response according to claim 3 is characterized in that: The teacher model foreground features are: The student model foreground features are:

5. The object detection knowledge distillation method based on feature fusion and global response according to claim 4 is characterized in that: Use the product-moment correlation coefficient to align the foreground feature relationship between the teacher model and the student model: Get the loss of each level feature distillation 6. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 3, the prediction results of the teacher model and the student model are calculated, and the IoU score (I t and I s ), which represents the intersection-over-union ratio of the predicted bounding box and the real box, and is used to measure the positioning accuracy. t >I s When the distillation step 4 is carried out.

7. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 4, T t Represents the output category probability of the teacher network, S t Represents the probability of the student network category, and the loss formula for calculating the target category is L target =-T t log(S t );T i Represents the probability of non-target categories output by the teacher model, S i represents the probability of non-target categories in the student model, T j and S j Represents the sum of target and non-target category probabilities output by the teacher model and the student model respectively; Normalize the non-target category probabilities of the teacher model and the student model, and the formula is Among them, after the normalization operation, the loss formula of the non-target category is L non-target =-∑ i≠t N(T i / τ)log(N(S i / τ)), τ is the temperature parameter, which increases the smoothness of the distribution; Combining the target category loss and the non-target category loss, we get a new classification loss formula L NTCM =ε1L target +ε2L non-target ; ε1 and ε2 represent the weights of the distillation loss of the target category and the non-target category, respectively.

8. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 5, the distillation loss formula of the response distillation BPM module is obtained as follows: in are four sides, and τ is the temperature value.

9. The object detection knowledge distillation method based on feature fusion and global response according to claim 1, characterized in that: In step 7, the overall response distillation loss is L r =δ1L NTCM+ δ2L BPM , where δ1 and δ2 are hyperparameters used to balance the loss.

10. The object detection knowledge distillation method based on feature fusion and global response according to claim 9, characterized in that: In step 6, the total loss of the entire distillation framework is L = L ori +λL f +βL r ; Among them, λ, β are hyperparameters used to balance the loss, L f It refers to the feature distillation loss related to the details of the distilled small target features, L r is the response distillation loss for global information, L ori is the original loss of the student model.