Target detection model compression method and system based on channel attention
By introducing the distillation loss function and 8-bit static quantization compression that are noted by channel detection model, the problem of imbalance between positioning and classification supervision and neglecting scale differences in the prior art is solved, improving the detection accuracy of the target detection model and reducing resource requirements.
Patent Information
- Application Number
- CN202510099976.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
AI Technical Summary
The existing knowledge distillation method is difficult to balance the degree of supervision of positioning and classification in the target detection task, and ignores regional differences at different scales, resulting in a decrease in the detection accuracy of the student model.
A method of compression of object detection model based on channel attention is proposed. By constructing the distillation loss function of channel attention, combining the object detection task loss function, student model parameters are optimized, so that it can better capture target information at different scales, and 8-bit static quantization compression is performed to improve computational efficiency.
Improves the performance of student models in detection accuracy, ensures deployment on resource-constrained devices, and meets deployment conditions on devices with computing resource-constrained devices.
Smart Images

Figure CN120106148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision technology, and in particular to a channel attention-based target detection model compression method and system. Background Art
[0002] Object detection is one of the core tasks in the field of computer vision, which aims to identify objects in an image and provide each object with its category label and precise bounding box location. With the rapid development of deep learning technology, neural network-based object detection models have achieved leading performance indicators in object detection tasks.
[0003] Neural network target detection models with high detection accuracy are often large in scale, complex in calculation, and have high storage requirements, making them difficult to deploy on resource-constrained devices. In order to effectively deploy the model to a specific device, commonly used optimization techniques include knowledge distillation, quantization compression, and model pruning. Knowledge distillation captures the characteristic information and patterns output by a complex "teacher" model and transfers its knowledge to a smaller "student" model; quantization compression uses lower-precision data types during model calculations, which not only improves the computational efficiency of arithmetic operations, but also reduces storage requirements during calculations; model pruning evaluates the significance of the impact of each weight or neuron on the final result and removes relatively unimportant weights or neurons to achieve model size compression.
[0004] Since object detection is a comprehensive task involving bounding box positioning and category classification, most of the existing knowledge distillation methods are designed based on classification models, ignoring the regression task of bounding box positioning. Directly using traditional knowledge distillation methods often leads to an imbalance in the degree of positioning and classification supervision, which in turn leads to a decrease in overall detection accuracy. In addition, object detection models usually contain multiple different task modules, such as candidate region generation, feature extraction, classification and regression. How to allocate distillation weights among these different tasks is a major problem. Finally, since there are objects of different sizes, shapes, and positions in the object detection task, existing methods often ignore regional differences of different scales, making it difficult for student models to focus on different objects of different scales. Summary of the invention
[0005] The purpose of the present invention is to propose a channel attention-based target detection model compression method and system to improve the detection performance of the student model obtained through knowledge distillation.
[0006] The technical solution to achieve the purpose of the present invention is: a channel attention-based target detection model compression method, the steps are as follows:
[0007] Step 1: Collect the target images to be detected, and expand the original data set into the target detection data set through color transformation and geometric transformation;
[0008] Step 2: Build the YoloX series detection model as the teacher model and student model. The teacher model and the student model have the same structure but different number of channels. The backbone network is darknet, which is used for multi-scale feature extraction; the neck network is a feature pyramid network, which is used for multi-scale feature fusion; the head network is divided into two branches: regression and classification, which are used for bounding box regression and target category classification respectively;
[0009] Step 3: Construct the target detection task loss function as the teacher model loss function and train the teacher model;
[0010] Step 4: Load the backbone network of the teacher model as a multi-scale feature extractor, input the images in the target detection dataset into the feature extractor, and construct a multi-scale feature dataset;
[0011] Step 5: Construct a channel-attention distillation loss function, combine it with the target detection task loss function, construct a student model loss function, and jointly optimize the student model on the target detection dataset and the multi-scale feature dataset;
[0012] Step 6: Perform 8-bit static quantization compression on the student model, which is then used for the classification of the image to be detected. The image to be detected is input into the compressed student model to complete the actual image target detection task.
[0013] Further, step 1: collect the target images to be detected, and expand the original data set into the target detection data set through color transformation and geometric transformation. The specific method is as follows:
[0014] Download the images related to the target to be detected and mark the targets appearing in the images;
[0015] Considering the number of target categories to be detected and the distribution of the number of images in each category, color transformation and geometric transformation are used. Color transformation includes random combinations of brightness, contrast, saturation and hue, and geometric transformation includes random combinations of rotation, translation, scaling and other transformations. The original dataset is expanded to a target detection dataset containing a total of 2,000 images with a balanced category distribution.
[0016] Furthermore, the target-related image to be detected is an image containing a vehicle, and the targets appearing in the image are labeled and classified into five types of targets, namely, sedans, SUVs, trucks, vans, and buses.
[0017] Further, step 2: construct the YoloX series detection model as the teacher model and the student model. The teacher model and the student model have the same structure but different number of channels. The backbone network is darknet, which is used for multi-scale feature extraction of the input image; the neck network is a feature pyramid network, which is used for multi-scale feature fusion; the head network is divided into two branches: regression and classification, which are used for bounding box regression and target category classification respectively, where:
[0018] The basic number of channels for the teacher model is 128, and the basic number of channels for the student model is 96.
[0019] Further, step 3: construct the target detection task loss function as the teacher model loss function, train and build the teacher model, the specific method is:
[0020] Construct the target detection task loss function L det , including the positioning loss L loc , confidence loss L conf With the classification loss L cls , which are used to measure the bounding box positioning performance, the binary classification performance of whether there is a target, and the target multi-category classification performance of the target detection model, respectively.
[0021] Positioning loss L loc The intersection-over-union ratio is used to measure the difference between the predicted bounding box and the true bounding box:
[0022]
[0023] Where N is the total number of samples; is an indicator value, which is 1 when the target is contained in sample i, otherwise it is 0; is the predicted bounding box b of sample i i With the ground-truth bounding box The intersection-and-union ratio between them;
[0024] Confidence loss L conf To measure whether the predicted bounding box contains the target and its accuracy, use the binary cross entropy loss:
[0025]
[0026] in, It is also an indicator value, which is 1 when sample i does not contain the target, otherwise it is 0; BCE(c i ,t) is the prediction confidence c of sample i i The binary cross entropy loss between t and the target value t;
[0027] Classification loss L cls To measure the difference between the predicted category and the true category, cross entropy loss is used;
[0028] Object detection loss L det The weighted sum of the three losses
[0029] L det =L loc +λ conf L conf +λ cls L cls ;
[0030] A large learning rate is set so that the teacher model stops early after reaching a high accuracy after several rounds of iterations. The condition for judging early stopping is that the accuracy of the model on the validation set does not improve for five consecutive rounds.
[0031] Further, step 5: construct the channel attention distillation loss function, combine the target detection task loss function, construct the student model loss function, and jointly optimize the student model on the target detection dataset and the multi-scale feature dataset. The specific method is:
[0032] Construct a channel-attention distillation loss function to align the multi-scale feature maps of the student model backbone network with the output of the teacher model. Since the number of channels of the feature maps output by the teacher model and the student model is different, a 1x1 convolution kernel is first used to reduce the dimension of the feature map extracted by the teacher model, and then the weight value of each channel is calculated through the global maximum pooling and activation function. The mean square loss between the output feature map of the student model and the output feature map of the teacher model after dimension reduction is calculated channel by channel, and the loss function terms of different channels are weighted to obtain the distillation loss function.
[0033] If the backbone networks of the teacher model and the student model and their parameters are recorded as (g t ,W t ) and (g s ,W s ), extract features at different scales, use superscript k to represent different scales, and record the different scale features extracted by the teacher model as:
[0034]
[0035] Then the distillation loss function is:
[0036]
[0037] Combined with the target detection task loss function, the student model loss function is constructed as:
[0038] L total =L det +λ dis L dis
[0039] A large learning rate is set so that the student model stops early after reaching a high accuracy after several rounds of iterations. The condition for judging early stopping is that the accuracy of the model on the validation set does not improve for 5 consecutive rounds.
[0040] Further, step 6: the student model is subjected to 8-bit static quantization compression, which is subsequently used for the classification of the image to be detected. The image to be detected is input into the compressed student model to complete the actual image target detection task. The specific method is as follows:
[0041] Stratified sampling is performed from the target detection dataset to obtain a subset with uniform distribution of image sample categories as a calibration dataset, and the images in the calibration dataset are input into the student model in batches for forward inference;
[0042] Count the dynamic distribution range of the weights and activation values of the student model on the calibration data set, and map the dynamic distribution to the representation range of 8-bit integers to determine the quantization coefficient;
[0043] The student model is quantized into an 8-bit quantized model for actual image target detection tasks.
[0044] A channel-attention-based target detection model compression system implements the channel-attention-based target detection model compression method, realizes channel-attention-based target detection model compression and target detection, and is divided into six modules to respectively execute steps 1 to 6.
[0045] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the channel-attention-based target detection model compression method is implemented to achieve channel-attention-based target detection model compression and target detection.
[0046] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the channel-attention-based target detection model compression method to achieve channel-attention-based target detection model compression and target detection.
[0047] Compared with the prior art, the present invention has the following significant advantages: through multi-scale feature extraction of the teacher model, the student model can better capture target information of different scales; in addition, training strategies of different levels of detail are set for the teacher model and the student model to ensure efficient acquisition of a lightweight target detection model with higher accuracy; through further quantization compression, the computing efficiency is improved and the storage requirements are reduced to meet the deployment conditions on devices with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Flowchart of the channel attention based object detection model compression method. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] The present invention discloses a channel attention-based target detection model compression method, comprising:
[0051] Step 1: Collect the target images to be detected, and expand the original data set into a target detection data set of appropriate size through color transformation and geometric transformation.
[0052] Download images related to the target to be detected from the Internet and other channels, and annotate the targets appearing in the images. Then, considering the number of target categories to be detected and the distribution of the number of images in each category, the original data set is expanded to a target detection data set containing 2,000 images with balanced category distribution by using color transformation (including random combinations of brightness, contrast, saturation and hue) and geometric transformation (including random combinations of rotation, translation, scaling and other transformations).
[0053] Step 2: Build YoloX-m and YoloX-s as the teacher model and student model respectively.
[0054] Both the teacher model and the student model are yolox series target detection models. The backbone network of the yolox series detection model is darknet, which is used for multi-scale feature extraction; the neck network is the feature pyramid network (FeaturePyramidNetwork), which performs multi-scale feature fusion; the head network is divided into two branches: regression and classification, which are used for bounding box regression and target category classification respectively.
[0055] Cross-stage local (CSPLayer) is a basic unit for cross-layer connection. It is used as the basic unit of the backbone network Darknet and the neck network FPN in the yolox series network. The backbone network Darknet and the neck network feature pyramid network are constructed using the cross-stage local network as the basic unit, and two convolutional network branches are constructed as the bounding box regression and category classification prediction head networks respectively. The student model and the teacher model are both built using the above-mentioned backbone, neck, and head networks. The difference between the two is the number of channels. The basic number of channels of the teacher model is 128, and the number of channels of each layer thereafter is a multiple of 128, such as 256, 512; and the basic number of channels of the student model is 96. In this way, a teacher model and a student model with the same structure but different number of channels are constructed.
[0056] The target detection network and its parameters are denoted as f and W respectively, then the forward calculation process of the network can be expressed as [b,c,p]=f(x;W), where x represents the image data, b represents the predicted bounding box, c represents the confidence that the target exists in each predicted bounding box, and p represents the probability that the object in the predicted bounding box belongs to each category.
[0057] Step 3: Construct the target detection task loss function based on the teacher model, set a larger learning rate, and end the training process early when the accuracy of the validation set does not improve for 5 consecutive rounds.
[0058] Construct the target detection task loss function L det , including the positioning loss L loc , confidence loss L conf With the classification loss L cls , which are used to measure the bounding box positioning performance, the binary classification performance of whether there is an object, and the multi-category classification performance of the target detection model. Among them:
[0059] The positioning loss often uses the intersection over union (IoU) to measure the difference between the predicted bounding box and the true bounding box:
[0060]
[0061] Where N is the total number of samples; is an indicator value, which is 1 when the target is contained in sample i, otherwise it is 0; is the predicted bounding box b of sample i i With the ground-truth bounding box The intersection ratio between them.
[0062] Confidence loss measures whether the predicted bounding box contains the target and its accuracy, usually using binary cross entropy loss (BCE, Binary Cross-Entropy Loss):
[0063]
[0064] in, It is also an indicator value, which is 1 when sample i does not contain the target, otherwise it is 0; BCE(c i ,t) is the prediction confidence c of sample i i and the target value t.
[0065] The classification loss measures the difference between the predicted category and the true category. Here, the cross-entropy loss is used.
[0066] The target detection loss is the weighted sum of the above three losses
[0067] L det =L loc +λ conf L conf +λ cls L cls .
[0068] The stochastic gradient descent optimizer and the hot restart cosine annealing learning rate planner are used to update the parameters. The initial learning rate is set to 0.01, so that the teacher model stops early after reaching a high accuracy after several rounds of iterations. The condition for judging early stopping is that the accuracy of the model on the validation set does not improve for five consecutive rounds.
[0069] Step 4: Load the teacher model and delete the neck network and head network parts, keep the backbone network as the multi-scale feature extractor, input the images in the target detection dataset into the feature extractor, and construct a multi-scale feature dataset.
[0070] Load the teacher model trained in step 3 and detect the head network, retain only the backbone network part as the feature extractor, and fix the weight of the feature extractor; use the feature extractor to extract the multi-scale feature map of the image in the target detection dataset and store it as a multi-scale feature dataset.
[0071] Step 5: Construct the target detection task loss function based on the student model, and also construct the channel attention distillation loss function, and jointly optimize the student model parameters on the target detection dataset and the multi-scale feature dataset.
[0072] First, the target detection task loss function is constructed based on the student model; then, a channel-attentive distillation loss function is constructed to align the multi-scale feature maps of the student model backbone network to the output of the teacher model. Since the number of channels of the feature maps output by the teacher model and the student model is different, a 1x1 convolution kernel is first used to reduce the dimension of the feature map extracted by the teacher model, and then the weight value of each channel is calculated through the global maximum pooling and activation function. The mean square loss between the output feature map of the student model and the output feature map of the teacher model after dimension reduction is calculated channel by channel, and the loss function terms of different channels are weighted by the above weight values to obtain the distillation loss function. The target detection task loss function and the distillation loss function are weighted by the two weight values, and the sum of the two weights is 1; and the weight of the distillation loss term decays exponentially with the increase in the number of training rounds, so that the model eventually tends to update the parameters in the direction of the gradient of the target detection task loss function.
[0073] Conventional knowledge distillation is aimed at category classification tasks. The supervision methods used are hard label supervision and soft label supervision. Hard label supervision is to train the student model through the standard cross entropy loss function:
[0074]
[0075] Among them, y i is the category prediction output by the teacher network, represented by a one-hot code; p i represents the predicted probability of the student model for category i.
[0076] The soft label refers to the probability distribution of the output of the teacher model adjusted by the temperature parameter T. This distribution contains more knowledge about the similarity between categories:
[0077]
[0078] Among them, z i With s i are the log odds of the teacher model and the student model’s predictions on category i, respectively.
[0079] In order to ensure that the student model has the feature extraction ability of the original target detection model and has a certain flexibility in bounding box regression and target classification tasks, this method constructs a distillation loss based on the backbone network output of the teacher model and the student model. If the backbone network and its parameters of the teacher model and the student model are respectively denoted as (g t ,W t ) and (g s ,W s ), extract features at three scales. The superscript k represents different scales, and the different scale features extracted by the teacher model are recorded as:
[0080]
[0081] Then the distillation loss function is:
[0082]
[0083] The final loss function is:
[0084] L total =L det +λ dis L dis
[0085] The stochastic gradient descent optimizer is used to update the parameters of the student model, the learning rate is set to 0.001, and a fixed number of iterations are performed until the loss function converges.
[0086] Step 6: Perform 8-bit static quantization compression on the student model.
[0087] Layered sampling is performed from the target detection dataset to obtain a subset with uniform distribution of image sample categories as the calibration dataset. The images in the calibration dataset are input into the student model in batches for forward inference, and the values of weights and activation values in the model are counted; the dynamic distribution range of the weights and activation values of the student model on the calibration dataset is counted, and the quantization coefficient is determined by mapping the range to the representation range of an 8-bit integer. The calculation result x for the multi-scale feature extraction layer of the student model out , its quantization coefficient is S, then the quantized calculation result is:
[0088]
[0089] Step 7: Input the image to be tested into the compressed student model to complete the actual image target detection task.
[0090] The present invention also proposes a channel-attention-based target detection model compression system, implements the channel-attention-based target detection model compression method, and realizes channel-attention-based target detection model compression and target detection.
[0091] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the channel-attention-based target detection model compression method is implemented to achieve channel-attention-based target detection model compression and target detection.
[0092] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the channel-attention-based target detection model compression method to achieve channel-attention-based target detection model compression and target detection.
[0093] Example
[0094] In order to verify the effectiveness of the solution of the present invention, the following simulation experiment is carried out.
[0095] In the first step, 523 vehicle target images to be detected are downloaded from the Internet, including five types of targets: sedans, SUVs, trucks, vans, and buses; the image data are labeled; and the original data set is expanded into a target detection data set with balanced category distribution through color transformation and geometric transformation, which contains a total of 2,000 images.
[0096] The second step is to build YoloX-m and YoloX-s models as the teacher model and student model respectively.
[0097] The third step is to construct the target detection task loss function to train the teacher model; use the stochastic gradient descent optimizer, set the initial learning rate to 0.03, and adjust the learning rate according to the cosine annealing rule with hot restart. The training process ends when the validation set accuracy does not improve for 5 consecutive rounds.
[0098] The fourth step is to load the teacher model, delete the detection head network, and retain the backbone network as the feature extractor; traverse the images in the target detection dataset, input the feature extractor for forward calculation, store the feature maps output by different layers of the network, and construct a multi-scale feature map dataset.
[0099] The fifth step is to construct the target detection task loss function and the channel attention distillation loss function, set the target detection loss function weight and the distillation loss function weight to 0.0 and 1.0 respectively, let the distillation loss function weight decay exponentially to 0.1 during training, and keep the sum of the target detection loss function weight and the distillation loss function weight to 1.0; the optimizer settings are the same as the teacher network training settings, and the initial learning rate is set to 0.01; train fully until the total loss function converges.
[0100] The sixth step is to stratify and sample a subset from the target detection dataset as the calibration dataset to ensure that the image sample categories are evenly distributed. The images in the subset are used as the input of the student model for forward calculation, and the weights, biases, and activation values of each layer of the model are statistically analyzed to determine the mapping relationship from floating-point numbers to 8-bit integers.
[0101] In the seventh step, the image to be tested is input into the compressed student model to complete the actual image target detection task.
[0102] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0103] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A channel attention-based object detection model compression method, characterized in that: Here are the steps: Step 1: Collect the target images to be detected, and expand the original data set into the target detection data set through color transformation and geometric transformation; Step 2: Build the YoloX series detection model as the teacher model and student model. The teacher model and the student model have the same structure but different number of channels. The backbone network is darknet, which is used for multi-scale feature extraction; the neck network is a feature pyramid network, which is used for multi-scale feature fusion; the head network is divided into two branches: regression and classification, which are used for bounding box regression and target category classification respectively; Step 3: Construct the target detection task loss function as the teacher model loss function and train the teacher model; Step 4: Load the backbone network of the teacher model as a multi-scale feature extractor, input the images in the target detection dataset into the feature extractor, and construct a multi-scale feature dataset; Step 5: Construct a channel-attention distillation loss function, combine it with the target detection task loss function, construct a student model loss function, and jointly optimize the student model on the target detection dataset and the multi-scale feature dataset; Step 6: Perform 8-bit static quantization compression on the student model, which is then used for the classification of the image to be detected. The image to be detected is input into the compressed student model to complete the actual image target detection task.
2. The object detection model compression method based on channel attention according to claim 1 is characterized in that: Step 1: Collect the target images to be detected, and expand the original data set into the target detection data set through color transformation and geometric transformation. The specific method is as follows: Download the images related to the target to be detected and mark the targets appearing in the images; Considering the number of target categories to be detected and the distribution of the number of images in each category, color transformation and geometric transformation are used. Color transformation includes random combinations of brightness, contrast, saturation and hue, and geometric transformation includes random combinations of rotation, translation, scaling and other transformations. The original dataset is expanded to a target detection dataset containing a total of 2,000 images with a balanced category distribution.
3. The channel-attention-based target detection model compression method according to claim 2, characterized in that: The relevant images of the target to be detected are images containing vehicles. The targets appearing in the images are marked and classified into five categories: sedans, SUVs, trucks, vans, and buses.
4. The channel-attention-based target detection model compression method according to claim 1, characterized in that: Step 2: Build the YoloX series detection model as the teacher model and the student model. The teacher model and the student model have the same structure but different number of channels. The backbone network is darknet, which is used for multi-scale feature extraction of the input image; the neck network is a feature pyramid network, which is used for multi-scale feature fusion; the head network is divided into two branches: regression and classification, which are used for bounding box regression and target category classification respectively. The basic number of channels for the teacher model is 128, and the basic number of channels for the student model is 96.
5. The channel-attention-based target detection model compression method according to claim 1, characterized in that: Step 3: Construct the target detection task loss function as the teacher model loss function and train the teacher model. The specific method is: Construct the target detection task loss function L det , including the positioning loss L loc , confidence loss L conf With the classification loss L cls , which are used to measure the bounding box positioning performance, the binary classification performance of whether there is a target, and the target multi-category classification performance of the target detection model, respectively. Positioning loss L loc The intersection-over-union ratio is used to measure the difference between the predicted bounding box and the true bounding box: Where N is the total number of samples; is an indicator value, which is 1 when the target is contained in sample i, otherwise it is 0; is the predicted bounding box b of sample i i With the ground-truth bounding box The intersection-and-union ratio between them; Confidence loss L conf To measure whether the predicted bounding box contains the target and its accuracy, use the binary cross entropy loss: in, It is also an indicator value, which is 1 when sample i does not contain the target, otherwise it is 0; BCE(c i , t) is the prediction confidence c of sample i i The binary cross entropy loss between t and the target value t; Classification loss L cls To measure the difference between the predicted category and the true category, cross entropy loss is used; Object detection loss L det The weighted sum of the three losses L det =L loc +λ conf L conf +λ cls L cls ; A large learning rate is set so that the teacher model stops early after reaching a high accuracy after several rounds of iterations. The condition for judging early stopping is that the accuracy of the model on the validation set does not improve for five consecutive rounds.
6. The channel-attention-based target detection model compression method according to claim 1, characterized in that: Step 5: Construct a channel-attention distillation loss function, combine it with the target detection task loss function, construct a student model loss function, and jointly optimize the student model on the target detection dataset and the multi-scale feature dataset. The specific method is: Construct a channel-attention distillation loss function to align the multi-scale feature maps of the student model backbone network with the output of the teacher model. Since the number of channels of the feature maps output by the teacher model and the student model is different, a 1x1 convolution kernel is first used to reduce the dimension of the feature map extracted by the teacher model, and then the weight value of each channel is calculated through the global maximum pooling and activation function. The mean square loss between the output feature map of the student model and the output feature map of the teacher model after dimension reduction is calculated channel by channel, and the loss function terms of different channels are weighted to obtain the distillation loss function. If the backbone networks of the teacher model and the student model and their parameters are recorded as (g t , W t ) and (g s , W s ), extract features at different scales, use superscript k to represent different scales, and record the different scale features extracted by the teacher model as: Then the distillation loss function is: Combined with the target detection task loss function, the student model loss function is constructed as: L total =L det +λ dis L dis A large learning rate is set so that the student model stops early after reaching a high accuracy after several rounds of iterations. The condition for judging early stopping is that the accuracy of the model on the validation set does not improve for 5 consecutive rounds.
7. The object detection model compression method based on channel attention according to claim 1 is characterized in that: Step 6: Perform 8-bit static quantization compression on the student model, which is then used for the classification of the image to be detected. The image to be detected is input into the compressed student model to complete the actual image target detection task. The specific method is as follows: Stratified sampling is performed from the target detection dataset to obtain a subset with uniform distribution of image sample categories as a calibration dataset, and the images in the calibration dataset are input into the student model in batches for forward inference; Count the dynamic distribution range of the weights and activation values of the student model on the calibration data set, and map the dynamic distribution to the representation range of 8-bit integers to determine the quantization coefficient; The student model is quantized into an 8-bit quantized model for actual image target detection tasks.
8. A channel attention-based object detection model compression system, characterized in that: Implement the channel attention-based target detection model compression method described in any one of claims 1-7 to achieve channel attention-based target detection model compression and target detection, and perform steps 1 to 6 in six modules respectively.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the channel-attention-based target detection model compression method according to any one of claims 1 to 7 is implemented to realize channel-attention-based target detection model compression and target detection.
10. A computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the channel-attention-based target detection model compression method according to any one of claims 1 to 7 is implemented to achieve channel-attention-based target detection model compression and target detection.