Lightweight target detection method based on YOLO
By introducing lightweight FasterNet module, Dim-SimAM module and joint pruning-optimization module into the YOLOv8 object detection model, the problem of high demand for computing resources of existing object detection models is solved, real-time object recognition on low-computing equipment is achieved, and model accuracy and inference speed are improved.
Patent Information
- Application Number
- CN202510654862.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing target detection models have high demand for computing resources, making it difficult to achieve real-time target recognition on low-computing equipment such as drones, and the lightweight model accuracy is too low and difficult to recover.
Using the lightweight object detection method based on YOLOv8, the introduction of the lightweight FasterNet module, the Dim-SimAM module and the joint pruning-optimization module reduces the amount of model parameters and computational complexity, while improving the model accuracy through pruning and knowledge distillation.
While ensuring accuracy, it significantly reduces the calculation amount and computational complexity of the model, realizes real-time inference operations on low-cost edge devices, and improves the portability and inference speed of the model.
Smart Images

Figure CN120182585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and provides a lightweight object detection method based on YOLO. Background Art
[0002] Drone applications need to quickly locate targets in complex backgrounds (such as farmland and urban building complexes) for real-time tasks such as identifying crop diseases and pests and inspecting the flow of people. However, in order to achieve high accuracy and powerful object recognition capabilities, existing object detection models have very high requirements for computing resources. The training and inference of complex models usually require the support of a GPU with powerful computing capabilities. The performance of the processors that drones can carry is very limited, and the high computational load leads to a sharp increase in power consumption, severely shortening the flight time of drones. In summary, there is a great need for algorithms with low power consumption and low computational volume.
[0003] Object detection algorithms integrate object localization and classification. Detectors are divided into two categories: two-stage algorithms based on candidate regions and one-stage algorithms based on regression. The two-stage algorithm is divided into two steps: object localization and classification. First, region proposals are generated and then classified. For example, Faster R-CNN (Fast Region-based Convolutional Network), which has high detection accuracy but slow inference speed. The one-stage algorithm directly uses a CNN (Convolutional Neural Network) to extract features and simultaneously predicts classification and localization. It is simple and efficient and suitable for real-time scenarios, such as OverFeat and YOLO. Among them, the YOLO series of one-stage object detection algorithms have more balanced advantages and have been well verified in industrial applications. The YOLOv8 framework has achieved cutting-edge performance in terms of both accuracy and speed. Existing object detection network lightweighting solutions have some insurmountable disadvantages: it is impossible to balance detection accuracy and the number of model parameters, the accuracy of the pruned model is too low, and it is difficult to restore the accuracy of the pruned model. Summary of the Invention
[0004] The present invention aims to at least solve one of the technical problems existing in the related art. Therefore, the present invention provides a lightweight object detection method based on YOLO, which breaks through the limitations of computing power and can perform real-time inference operations on low-cost edge devices while ensuring accuracy.
[0005] The present invention provides a lightweight object detection method based on YOLO, including: S1: Obtain and preprocess an image to obtain a preprocessed image; S2: Establish a lightweight object detection model based on the YOLOv8 framework. The lightweight object detection model includes a backbone network, a feature extraction network, and a joint pruning-optimization module. The backbone network includes a lightweight FasterNet module, and the feature extraction network includes a Dim-SimAM module. S3: Set training parameters and use the preprocessed images to train the lightweight object detection model to obtain a lightweight object detection training model. S4: Input the required detection images into the lightweight object detection training model to obtain detection results.
[0006] According to a lightweight object detection method based on YOLO provided by the present invention, the lightweight FasterNet module is used to replace the feature extraction module in the backbone network.
[0007] According to a lightweight object detection method based on YOLO provided by the present invention, the running steps of the lightweight FasterNet module are as follows: S11: Divide the FasterNet input data into channel group data, perform PConv convolution on the channel group data by channel to obtain the first module output data. S12: Perform depthwise separable convolution on the first module output data to obtain the second module output data. S13: Perform GLU stacking on the second module output data, and use the ReLU function to activate to obtain the third module output data. S14: Perform convolution on the third module output data to obtain the fourth module output data. S15: Multiply the FasterNet input data and the fourth module output data element by element to obtain the FasterNet output data.
[0008] According to a lightweight object detection method based on YOLO provided by the present invention, the Dim-SimAM module replaces the RepNCSPELAN4 module in the feature extraction network.
[0009] According to a lightweight object detection method based on YOLO provided by the present invention, the Dim-SimAM module includes: S21: Solve the intermediate feature map and the minimized energy function of the input feature map according to the SimAM algorithm. S22: Solve the dynamic parameters according to the input feature map and the minimized energy function. S23: Update according to the dynamic parameters and the intermediate feature map to obtain the updated SimAM algorithm. S24: Calculate the output feature map based on the updated SimAM algorithm for the input feature map.
[0010] According to a lightweight object detection method based on YOLO provided by the present invention, the dynamic parameter adjusts the benchmark of the energy function by adding the variance of the intermediate feature map.
[0011] According to a lightweight object detection method based on YOLO provided by the present invention, when the input feature map is a multi-layer feature map, the downsampling and upsampling operations fuse different feature maps, perform multi-scale downsampling on the input feature map, independently apply SimAM at each scale to generate attention weights, align the resolution by bilinear interpolation upsampling, and weighted-fuse the input feature map to obtain the output feature map.
[0012] According to a lightweight object detection method based on YOLO provided by the present invention, the joint pruning-optimization module includes the following steps: S31: Calculate the sparsity loss of the lightweight object detection model; S32: Calculate the distillation loss of the lightweight object detection model; S33: Calculate the task loss of the lightweight object detection model, where the task loss is the sum of the box regression loss and the object confidence loss of the YOLO model; S34: Calculate the total loss, where the total loss is the sum of the task loss, sparsity loss, and distillation loss. When the total loss converges, the training of the lightweight object detection model is completed.
[0013] According to a lightweight object detection method based on YOLO provided by the present invention, the steps for calculating the sparsity loss include: S311: Sparsify the BN layer coefficients of the lightweight object detection model using regularization to obtain the pruned lightweight object detection model, and calculate the layer sparsity parameters; S312: Calculate the sparsity loss according to the layer sparsity parameters.
[0014] According to a lightweight object detection method based on YOLO provided by the present invention, the steps for calculating the distillation loss include: S321: Perform knowledge distillation on the lightweight object detection model to obtain a teacher model and a student model; S322: Calculate the dynamic response of the teacher model, and adjust the pruning threshold according to the dynamic response; S323: Calculate the distillation loss according to the pruning threshold, teacher model, and student model.
[0015] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: A lightweight object detection method based on YOLO provided by the present invention realizes the following advantages by introducing a lightweight FasterNet module, a Dim-SimAM module, and a joint pruning-optimization module: (1) The present invention solves the problems of large model size and limited portability. By using a partial convolution algorithm to reduce redundant calculations and memory access, the parameter calculation amount and computational complexity of the model can be significantly reduced.
[0016] (2) Without increasing the original network parameters, three-dimensional attention weights are inferred for the feature map, and global adaptive weighting is performed on the input feature map, enhancing the feature extraction ability of the model and effectively solving the problem of accuracy loss caused by lightweighting.
[0017] (3) The combination of model pruning and knowledge distillation improves the inference speed and reduces the number of network parameters while ensuring the accuracy.
[0018] Additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of a lightweight object detection method based on YOLO provided by the present invention.
[0021] Figure 2 It is a structural diagram of the lightweight FasterNet module.
[0022] Figure 3 It is a parameter information diagram during the model training process.
[0023] Figure 4 It is a performance comparison diagram between the present invention and different algorithms. Detailed Embodiments
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0025] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0026] Embodiment The following combines Figures 1 to 4 to describe the present invention.
[0027] As Figure 1 shown, Figure 1 is a flowchart of a lightweight object detection method based on YOLO provided by the present invention, including the following steps: S1: Obtain and preprocess an image to obtain a preprocessed image; S2: Establish a lightweight object detection model based on the YOLOv8 framework. The lightweight object detection model includes a backbone network, a feature extraction network, and a joint pruning-optimization module; the backbone network includes a lightweight FasterNet module, and the feature extraction network includes a Dim-SimAM module; S3: Set training parameters and use the preprocessed image to train the lightweight object detection model to obtain a lightweight object detection training model; S4: Input the detection image to be detected into the lightweight object detection training model to obtain a detection result.
[0028] Among them, the lightweight FasterNet module is used to replace the feature extraction module in the backbone network. The Dim-SimAM module replaces the RepNCSPELAN4 module of the feature extraction network.
[0029] Specifically, as Figure 2The running steps of the lightweight FasterNet module are as follows: S11: Divide the FasterNet input data into channel group data, and perform PConv convolution on the channel group data by channel to obtain the first module output data; S12: Perform depthwise separable convolution on the first module output data to obtain the second module output data; S13: Perform GLU stacking on the second module output data, and activate it using the ReLU function to obtain the third module output data; S14: Perform convolution on the third module output data to obtain the fourth module output data; S15: Multiply the FasterNet input data and the fourth module output data element by element to obtain the FasterNet output data.
[0030] The lightweight FasterNet module structure is mainly a PConv module (Partial Convolution), where the PConv structure only performs conventional convolution on some input channels, fully utilizes the information of all channels, extracts spatial features while keeping the sizes of the remaining channels unchanged. During conventional memory access, if the number of channels of the input and output feature maps is the same, then the first or last continuous channel is used as the representative of the entire feature map during calculation.
[0031] The computational complexity of PConv is only 1 / 16 of that of ordinary convolution. In addition, PConv has low memory access, and its computational complexity is as shown in the formula: where, is the channel height, is the channel width, is the continuous network channel, is the filter. The replacement of different network model architectures usually brings additional data operations, and the running time of these data operations has a great impact on the improvement of lightweight models. The addition of the FasterNet backbone network is to improve the speed of the object detection task while maintaining lightweight. Improving the design of PConv ensures no channel loss and exhibits good performance at low latency. The present invention proposes a collaborative optimization scheme of dynamic grouped partial convolution (DG-PConv) and gated linear unit (Convolutional GLU) to reconstruct the FasterNet backbone network, and the specific implementation is as follows: According to the channel correlation of the input feature map, dynamically divide the channel groups participating in convolution through a lightweight gating network (each group of channels independently performs 3×3 convolution), and the remaining channels maintain identity mapping The input features are divided into two groups through 1×1 convolution. One group, after 3×3 depthwise separable convolution (DWConv) and non-linear activation, is multiplied element-wise with the other group to achieve dynamic feature selection.
[0032] Specifically, the Dim-SimAM module includes: S21: Solve the intermediate feature map of the input feature map and the minimized energy function according to the SimAM algorithm; S22: Solve the dynamic parameters according to the input feature map and the minimized energy function; S23: Update according to the dynamic parameters and the intermediate feature map to obtain the updated SimAM algorithm; S24: Calculate the output feature map according to the input feature map and the updated SimAM algorithm.
[0033] Normally, neurons with rich information will exhibit different firing patterns from surrounding neurons, and firing neurons generally inhibit surrounding neurons, which is called spatial location inhibition. Therefore, neurons with spatial location inhibition should be given higher importance. The simplest way to search for such neurons is to measure the linear separability between the target neuron and other neurons. The importance of each neuron can be represented by a corresponding energy function as shown in the following formula: where, is the energy function, is the weight of the target neuron, is the bias coefficient, is the total number of energy functions, is the input of the neuron, is the neuron ordinal number, , is the input of the feature map on a single channel, is the dynamic parameter.
[0034] Its analysis is as follows where, is the average value of the single-channel data, is the variance of the single-channel data, is the mean value of all neurons.
[0035] Solve the minimized energy function of the th neuron in the SimAM module The algorithm is shown as follows: Among them, is the variance of all neurons, is the input on a single channel of the feature map corresponding to the minimized energy function, is the weight of the target neuron corresponding to the minimized energy function, is the bias coefficient corresponding to the minimized energy function. The lower it is, the greater the importance and difference between neuron and its surrounding neurons. Quantify the similarity between different positions on the feature map, and calculate the energy function to determine the feature map with significant differences in both the channel domain and the spatial domain. The refinement of the output feature map by the scaling operator is shown as follows: Among them, is the intermediate feature map, is the activation function, is the value of the energy function, is the input feature map, is element-wise multiplication.
[0036] In particular, the closed-form solution of the energy function is used to enhance the effective information output of neurons, where groups all in the channel and spatial dimensions. Aiming at the problem that the value of is fixed and cannot adapt to the requirements of different network layers or complex scenarios, the present invention introduces learnable dynamic parameters to adaptively adjust the sensitivity of the energy function according to the statistical characteristics of the input feature map. The formula is as follows: Among them, is the learnable parameter, is the global average pooling function.
[0037] At the same time, SimAM only acts on a single-layer feature map and lacks the ability of cross-scale information fusion. The present invention designs multi-scale pyramid attention (MS-Pyramid), which fuses feature maps of different resolutions through upsampling operations, performs multi-scale upsampling (such as 1 / 2, 1 / 4) on the input feature map, independently applies SimAM at each scale to generate attention weights, aligns the resolutions through bilinear interpolation upsampling, and weighted-fuses the multi-scale attention maps.
[0038] Among them, is the multi-layer energy function value, is the learnable scale weight coefficient, is the upsampling function, is the function for obtaining the single-layer energy function value.
[0039] The parameter-free attention SimAM is a general attention mechanism that is not limited to a specific network and has extremely high flexibility. By introducing the SimAM module into the model network, the representation ability of the convolutional layer is greatly improved. At the same time, most operations are based on the selection of the defined energy function, avoiding excessive structural adjustments.
[0040] Specifically, the joint pruning-optimization module includes the following steps: S31: Calculate the sparse loss of the lightweight object detection model, including: S311: Sparsify the BN layer coefficients of the lightweight object detection model using regularization to obtain the pruned lightweight object detection model, and calculate the layer sparse parameters; Use L1 regularization to sparsify the BN layer coefficients of the object detection model to enhance the model's adjustment towards a sparser structure. After sparse training, prune the channels according to a preset pruning rate to obtain an object detection model that occupies less storage space, achieving a balance between model size and accuracy. The pruning operation includes removing unimportant channels to lightweight the model. Channel pruning helps generate a more compact and efficient object detection model by removing unimportant channels. The pruning process is carried out by introducing a scaling factor and analyzing the statistical information of the BN layer to retain the channels that have a greater impact on the model accuracy and minimize the storage and computational costs of the model. The calculation formula of the BN layer is as follows: Among them, is the intermediate value of the BN layer, is the input of the BN layer, is the mean of the BN layer, is the variance of the BN layer, is the parameter to prevent the denominator from being zero, is the output of the BN layer, is the first parameter of the BN layer, is the second parameter of the BN layer. Multiply the parameter of each channel by the output of that channel, and then jointly optimize the weights of the original network and the parameter during the training process. When is close to 0, the output is independent of the input. Subsequently, for Sort and filter the parameters, directly removing those channels that are less than the set global threshold, thereby reducing the number of model parameters and the computational amount. The channel pruning algorithm utilizes the parameters in the BN layer As an important indicator for network pruning to measure the importance of each channel.
[0041] S312: Calculate the sparse loss according to the layer sparse parameters : Among them, is the first parameter of the th BN layer, is the layer ordinal number after pruning, , is the total number of layers after pruning, is the first-order norm.
[0042] Since the accuracy of the pruned model will decrease to some extent, it is necessary to retrain through fine-tuning to adapt to the updated network structure, thereby restoring the performance level of the model. For the problem of possible decrease in model accuracy during lightweight processing, the knowledge distillation method can be used. This method extracts knowledge from a model with high accuracy but large number of parameters and computational amount, and then transfers it to a smaller model to achieve model compression, so as to reduce the model size while maintaining high accuracy.
[0043] S32: Calculate the distillation loss of the lightweight object detection model, including: S321: Perform knowledge distillation on the lightweight object detection model to obtain a teacher model and a student model; After processing the results of the teacher and student networks through Softmax, calculate the loss. Finally, multiply the two loss results by the multiple of the coefficient value and then add them to obtain the total loss for training. During prediction, directly put the trained weights into the student network to directly perform prediction. To enhance the inter-class association, during data annotation, it is divided into Soft label and Hard label. Hard label is the original label, and Soft label carries more information to improve the generalization ability. The calculated probability output by the Softmax layer is used as the Soft target. The Softmax function is shown as follows: Among them, is the Soft label data, is the Hard label data, is the distillation temperature. The larger the distillation temperature, the smaller the numerical discrimination of the Soft label.
[0044] The present invention designs an offline distillation method. The large model teacher network is the improved YOLO-Fast model, and the small model student network is the pruned model. In the knowledge distillation training algorithm, the training accuracy of the teacher model is higher than that of the student model. The more significant the difference, the more obvious the distillation effect. Usually, the parameters of the teacher model are fixed to ensure that the knowledge of the teacher model will not be changed when training the student model. The distillation loss function calculates the output prediction difference between the teacher model and the student model and uses this difference as the loss of the student model, which is combined with the overall training loss to improve the performance and accuracy of the student model through gradient update, ultimately achieving the collaborative optimization of model compression and accuracy recovery.
[0045] S322: Calculate the dynamic response of the teacher model and adjust the pruning threshold according to the dynamic response : Wherein, is an adjustable coefficient, is the mean value function, is the first parameter mean of the BN layer of the teacher network at the th layer. The iterative pruning distillation mechanism is adopted, and forward propagation is alternately executed within each training cycle to calculate the joint loss, and the weights of the student network are updated by backpropagation to dynamically guide the channel importance evaluation through the feature response.
[0046] S323: Calculate the distillation loss according to the pruning threshold, the teacher model and the student model : Wherein, is the feature layer alignment set of the teacher-student network, is the loss of the teacher model Fn at the th layer, is the loss of the student model Fn at the th layer, is the second-order norm.
[0047] S33: Calculate the task loss of the lightweight object detection model, and the task loss is the sum of the box regression loss and the object confidence loss of the YOLO model; S34: Calculate the total loss , and the total loss is the sum of the task loss, the sparse loss and the distillation loss: Wherein, is the distillation loss parameter, is the pruning loss parameter, is the task loss parameter, is the task loss.
[0048] When the total loss converges, the training of the lightweight object detection model is completed.
[0049] As Figure 3 shown, Figure 3 The following is the parameter information of the lightweight model trained on the VisDrone dataset. Figure 3 In it, train and val represent the data of the training set and the validation set respectively. The blue dots represent the actual results result, and the orange dashed line represents the smoothed curve smooth. The training set will directly participate in the complete processes such as the forward inference of the model, loss calculation, backpropagation, and weight update. The validation set is only used for the forward inference of the model, and then the model is evaluated according to the predetermined evaluation metrics, which can be used to reflect the performance of the current model on the validation set. Among them, as Figure 3 shown in (a) in Figure 3 and (f) in Figure 3 the box regression loss (box_loss) is used to measure the positional difference between the predicted box and the ground truth box; as Figure 3 shown in (b) in Figure 3 and (g) in Figure 3 the class classification loss (cls_loss) is used to measure the accuracy of classifying different classes; as Figure 3 shown in (c) in Figure 3 and (h) in Figure 3 the error distance loss (dfl_loss) is used to measure the distance error between the predicted box and the calibrated box. In addition, as Figure 3 shown in (d) in metrics / precision is the proportion of samples where the predicted bounding box of the model coincides with the ground truth bounding box during the training process, that is, the proportion of true positive samples among the predicted positive samples. As shown in (e) in metrics / recall is the proportion of all true samples that the model can find during the training process, that is, the proportion of correctly predicted positive examples among the true positive samples. As . Compared with the original model YOLOv8, the mAP and FP are increased by 3.03% and 2.2% respectively. Then, through model sparsification training and channel pruning, the computational load of the model is compressed to 68.2%.
[0050] As Figure 4 shown, the YOLO-Fast algorithm proposed in the present invention is compared with some mainstream algorithms, including centerNet (Center Point Network), FCOS (Fully Convolutional One-Stage Object Detection), DETR (Detection Transformer), etc., in terms of mAP and inference speed (Inferences) on the Atlas 200I platform. It can be seen from the experimental results that the accuracy of the YOLO-Fast proposed in this paper reaches 91.17% on the mAP@0.5 of the VisDrone test set, and the inference time on the Atlas 200I reaches 42ms, with the highest detection accuracy among all models. Compared with the original model, both the detection accuracy and the inference speed are improved. Compared with the Yolov7, Yolov6 and Yolov5 models, the YOLO-Fast model proposed in the present invention is the first choice in terms of detection accuracy and speed.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0052] It should be noted that the embodiments of the present disclosure can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic: the software part can be stored in a memory and executed by a suitable instruction execution system such as a microprocessor or dedicated designed hardware. Those skilled in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in the processor control code, for example, such code is provided on a programmable memory or a data carrier such as an optical or electronic signal carrier.
[0053] In addition, although the operations of the methods of the present disclosure are depicted in the drawings in a particular order, this is not a requirement or implication that the operations must be performed in that particular order, or that all of the illustrated operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart may be altered in their order of execution. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and executed, and / or one step may be decomposed into multiple steps and executed. It should also be noted that the features and functions of two or more devices according to the present disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.
[0054] Although the present disclosure has been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A lightweight target detection method based on YOLO, characterized in that: include: S1: Acquire and preprocess an image to obtain a preprocessed image; S2: A lightweight target detection model is established based on the YOLOv8 framework, wherein the lightweight target detection model includes a backbone network, a feature extraction network, and a joint pruning-optimization module; the backbone network includes a lightweight FasterNet module, The operation steps of the lightweight FasterNet module are: S11: Divide the FasterNet input data into channels to obtain channel group data, and perform PConv convolution on the channel group data by channel to obtain the first module output data; S12: Performing depth-wise separable convolution on the output data of the first module to obtain output data of the second module; S13: Perform GLU stacking on the output data of the second module, and activate it using the ReLU function to obtain output data of the third module; S14: performing convolution on the output data of the third module to obtain output data of the fourth module; S15: multiplying the FasterNet input data and the fourth module output data element by element to obtain FasterNet output data; The feature extraction network includes a Dim-SimAM module; S3: setting training parameters, and using the preprocessed images to train the lightweight object detection model to obtain a lightweight object detection training model; S4: Input the required detection image into the lightweight object detection training model to obtain the detection result.
2. A lightweight target detection method based on YOLO according to claim 1, characterized in that: The lightweight FasterNet module is used to replace the feature extraction module in the backbone network.
3. A lightweight target detection method based on YOLO according to claim 1, characterized in that: The Dim-SimAM module replaces the RepNCSPELAN4 module of the feature extraction network.
4. A lightweight target detection method based on YOLO according to claim 1, characterized in that: The Dim-SimAM module includes: S21: Solve the intermediate feature map of the input feature map and minimize the energy function according to the SimAM algorithm; S22: solving dynamic parameters according to the input feature map and the minimized energy function; S23: updating the dynamic parameters and the intermediate feature map to obtain an updated SimAM algorithm; S24: Calculate the input feature map according to the updated SimAM algorithm to obtain an output feature map.
5. A YOLO-based lightweight target detection method according to claim 4, characterized in that: The dynamic parameter adjusts the benchmark of the energy function by adding it to the variance of the intermediate feature map.
6. A YOLO-based lightweight target detection method according to claim 4, characterized in that: When the input feature map is a multi-layer feature map, downsampling and upsampling operations fuse different feature maps, perform multi-scale downsampling on the input feature map, independently apply SimAM at each scale to generate attention weights, align the resolution through bilinear interpolation upsampling, and weightedly fuse the input feature maps to obtain the output feature map.
7. A YOLO-based lightweight target detection method according to claim 4, characterized in that: The joint pruning-optimization module comprises the following steps: S31: Calculating the sparse loss of the lightweight object detection model; S32: Calculate the distillation loss of the lightweight object detection model; S33: Calculate the task loss of the lightweight target detection model, where the task loss is the sum of the box regression loss and the target confidence loss of the YOLO model; S34: Calculate the total loss, which is the sum of the task loss, sparse loss and distillation loss. When the total loss converges, the training of the lightweight object detection model is completed.
8. A YOLO-based lightweight target detection method according to claim 7, characterized in that: The sparse loss calculation step includes: S311: Use regularization to sparse the BN layer coefficients of the lightweight object detection model to obtain a pruned lightweight object detection model, and calculate the layer sparse parameters; S312: Calculate the sparse loss according to the layer sparse parameters.
9. A YOLO-based lightweight target detection method according to claim 8, characterized in that: The distillation loss calculation step comprises: S321: performing knowledge distillation on the lightweight object detection model to obtain a teacher model and a student model; S322: Calculate the dynamic response of the teacher model, and adjust the pruning threshold according to the dynamic response; S323: Calculate the distillation loss according to the pruning threshold, the teacher model and the student model.
Citation Information
Patent Citations
SAR optical image mapping model lightweight method based on conditional generative adversarial network
CN114202017A
Crankshaft internal defect detection method based on grading knowledge distillation and detection system thereof
CN115774851A
Bridge construction progress intelligent identification method based on improved YOLOV5S
CN116363586A
Underwater target identification method based on quantitative distillation
CN116524341A
Cross-architecture video action recognition method and device based on knowledge distillation
CN118172705A
Cited By
Smoke identification method and system suitable for comprehensive pipe gallery fire
CN121259962A
Endoscope image detection method, device, equipment, medium and product
CN121481930A