Server side lightweight method based on YOLOv8

The YOLOv8 model is lightweighted through LAMP pruning and response knowledge distillation methods, which solves the computational burden problem of real-time drone detection on edge platforms, significantly reduces the model size and computational burden, and meets real-time detection needs.

CN120688579APending Publication Date: 2025-09-23HANGZHOU HUICUI INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510618456.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies have difficulty achieving real-time and efficient drone target detection on edge platforms with weak computing power, especially due to the high computational burden and storage requirements of the model.

Method used

The YOLOv8 model was lightweighted using layer-wise adaptive pruning (LAMP) and response-based knowledge distillation, with a pruning ratio of 1:1.7. The model was then deployed on the Jetson AGX Orin 32G development kit using knowledge distillation technology.

Benefits of technology

It significantly reduces the model size by 88%, reduces the computational burden by 40%, and achieves real-time detection on the edge platform with an inference speed of only 14.7 milliseconds, outperforming the baseline model YOLOv8n.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688579A_ABST
    Figure CN120688579A_ABST
Patent Text Reader

Abstract

The invention discloses a server-side lightweight method based on YOLOv8, which comprises model pruning and knowledge distillation, specifically, interlayer adaptive pruning based on amplitude and knowledge distillation based on response, and a student model is trained by using output layer response of a teacher model. According to the method, the YOLOv8 model is pruned and distilled, so that the model which is higher in speed and smaller in size is obtained, and meanwhile, the performance loss is reduced as much as possible. In addition, an OpenCV library and a detection model on an edge platform are accelerated by using CUDA and TensorRT, and real-time unmanned aerial vehicle detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model lightweighting, and relates to a server-side lightweighting method based on YOLOv8. Background Technique

[0002] With the development of deep learning technology, vision-based detection methods have become the research focus in the field of server-side lightweighting due to their advantages of high accuracy, wide applicability, and reduced manual intervention. For example, some researchers have adopted advanced object detection networks such as the YOLO series, R-CNN series, SSD, and DETR for server-side lightweighting. Although deep learning methods have been widely applied in server-side lightweighting, most research focuses on image or video processing, with relatively less attention paid to real-time stream processing and edge deployment. By summarizing the research status in the introduction, the main challenges of real-time server-side lightweighting are as follows:

[0003] (1) How to design an accurate and fast detection algorithm to cope with the characteristics of small targets and high speeds of drones.

[0004] (2) How to effectively deploy the algorithm on edge platforms with computing power far lower than that of servers. Summary of the Invention

[0005] To solve the above problems, the present invention provides a server-side lightweighting method based on YOLOv8, including model pruning and knowledge distillation.

[0006] Preferably, the model pruning is specifically layer-wise adaptive pruning based on magnitude.

[0007] Preferably, the model pruning specifically includes:

[0008] In a feed-forward neural network with a depth of d, each layer has a corresponding weight tensor W (1) ,...,W (d) . To uniformly calculate the LAMP score, these weight tensors are flattened into one-dimensional vectors, and the LAMP score is used to evaluate the importance of connections; in the flattened one-dimensional vector, when sorting the weights in ascending order according to a given index, for any u < v, there is a weight tensor |W[u] ≤ W[v]|, then the LAMP score at the u-th index is defined as:

[0009]

[0010] where the weight mapped by the index u from the tensor W is W[u].

[0011] Preferably, the pruning ratio of the model pruning is 1:1.7.

[0012] Preferably, the knowledge distillation is specifically response-based knowledge distillation, which uses the output layer response of the teacher model to train the student model.

[0013] Preferably, the model after model pruning and knowledge distillation is deployed on the Jetson AGX Orin 32G development kit and evaluated on the test set.

[0014] The beneficial effects of the present invention include at least:

[0015] (1) An effective server-side lightweighting method is provided, including model pruning and knowledge distillation. These measures together achieve significant model lightweighting and acceleration, reducing the optimized model size by 88%, reducing the computational burden by 40%, and outperforming the baseline model YOLOv8n.

[0016] (2) We achieved lightweight deployment on the embedded side and used CUDA-accelerated OpenCV and TensorRT to accelerate the model. On the Jetson AGX Orin 32GB development board, the real-time processing time of the optimized model was only 14.7 milliseconds, fully meeting the real-time detection requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of model pruning based on the server-side lightweight method of YOLOv8 in the present invention;

[0018] Figure 2 Schematic diagram of knowledge distillation based on the YOLOv8 server-side lightweight method of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0020] On the contrary, the present invention covers any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention as defined by the claims. Furthermore, to facilitate a better understanding of the present invention, certain specific details are described in detail below in the detailed description of the present invention. Those skilled in the art will be able to fully understand the present invention without these details.

[0021] The server-side lightweight method based on YOLOv8 of the present invention includes model pruning and knowledge distillation.

[0022] In the specific embodiment, see Figure 1, the improvement in lightweight achieved by model pruning is incomparable to changing the model structure. To ensure that the improved YOLOv8 model can achieve better real-time performance on edge platforms while maintaining accuracy, the present invention adopts Layer-Adaptive Magnitude-based Pruning (LAMP). LAMP is an unstructured pruning method that determines whether to remove a connection by evaluating the importance of the connection, usually using the L1 or L2 norm to measure. However, LAMP introduces a new importance metric called the LAMP score, which effectively balances sparsity and performance. This method eliminates the need to manually adjust hyperparameters during the sparse training process, making the pruning process more efficient and automated, thereby improving the inference speed and storage efficiency of the model on edge devices.

[0023] In a feedforward neural network with depth d, each layer has a corresponding weight tensor W (1) ,...,W (d) , to uniformly calculate the LAMP score, these weight tensors are flattened into one-dimensional vectors, and the LAMP score is used to evaluate the importance of the connections; in the flattened one-dimensional vector, when sorting the weights in ascending order according to a given index, for any u < v, there is |W[u] ≤ W[v]| for the weight tensors, then the LAMP score at the u-th index is defined as:

[0024]

[0025] where the weight mapped by the index u from the tensor W is W[u].

[0026] To obtain a model with fewer parameters, lower computational complexity, and minimal accuracy loss, a series of comparative experiments were conducted to determine the appropriate pruning ratio. The results are shown in Table 1. LAMP pruning significantly reduces the number of model parameters, with the reduction ranging from 1 / 15 to 1 / 5, which is particularly important for edge platforms with limited storage. In addition, the reduction in parameters and computational complexity leads to varying degrees of decline in performance metrics. However, when the pruning ratio reaches 1:1.7, the precision (P) of the model reaches the optimal value, and the decline in recall rate (R), mAP50, and mAP50:95 is relatively small. Although a larger pruning ratio can result in fewer parameters and computations, it also causes excessive accuracy loss.

[0027] To strike a balance between reducing computational complexity and minimizing accuracy loss, a model with a pruning ratio of 1:1.7 was selected as the final pruned model. Compared to the unpruned version, the pruned model improved precision (P) by 0.9%, decreased recall (R) by 0.8%, decreased mAP by 0.9%, and only decreased mAP50.95 by 0.4%. The accuracy drop for all metrics remained below 1%, which is acceptable. Simultaneously, the number of parameters was reduced by 88%, and the number of GFLOPs was reduced by over 40%, significantly reducing the computational burden and parameter count. Further increasing the pruning ratio to 1:1.8, P and R decreased further by 0.8%, and mAP50 and mAP50.95 decreased by 0.6% and 0.5%, respectively, resulting in a significant performance loss. These results demonstrate that using the LAMP pruning method with a pruning ratio of 1:1.7 achieves effective lightweighting while maintaining model performance, providing a good foundation for real-time applications on edge platforms.

[0028] Table 1 Performance of the model under different pruning ratios

[0029]

[0030]

[0031] See also Figure 2 Knowledge distillation is the process of transferring knowledge from a large model to a smaller model for deployment under real-world constraints. This is essentially a form of model compression. When applied after model pruning, knowledge distillation can be considered an effective fine-tuning technique. In this study, response-based knowledge distillation was employed to restore the original performance of the pruned model.

[0032] The core idea of ​​response-based knowledge distillation is to use the output layer responses of the teacher model to train the student model. Its main advantage is that it can effectively transfer the implicit knowledge in the teacher model, which is not directly captured by labels, thereby improving the performance of the student model when handling complex tasks. In the knowledge distillation process, a student model and a teacher model are required. The present invention uses the pruned model as the student model and uses YOLOv8s as the teacher model. In this way, the pruned model can better capture the subtle features in complex tasks, further improving the application effect of the model on the edge platform and ensuring efficient and accurate drone detection performance.

[0033] Since knowledge distillation does not introduce additional parameters or computational overhead, we use it to fine-tune the pruned model to recover the accuracy lost during the pruning process. The results are shown in Table 2, demonstrating the effectiveness of this approach.

[0034] Table 2 Performance of different knowledge distillation methods

[0035]

[0036] Common knowledge distillation methods include response-based knowledge distillation, feature-based knowledge distillation, and a combination of the two. After comparing these three methods, it was found that response-based distillation performed best, so it was selected as the final distillation technique. Compared with other methods, response-based distillation improved the precision (P) of the pruned model by 0.2%, mAP by 0.2%, and mAP50:95 by 0.4%. Although the improvement in each metric is small, the key advantage of knowledge distillation is that it does not increase computational or storage costs, making it an excellent method for fine-tuning model performance in the final stage. This effectiveness further enhances the application potential of the YOLO-Drone model on edge platforms.

[0037] To verify the real-time performance of the lightweight model on edge devices, it was deployed on the Jetson AGX Orin32G development kit and evaluated on a test set. The accuracy of the lightweight model on the server and edge device did not differ significantly, but after using the TensorRT accelerated inference framework, the accuracy showed slight fluctuations, remaining within 1%.

[0038] In terms of inference speed, the embedded side is significantly slower than the server side. In addition, thanks to the over 50% acceleration provided by TensorRT, although the inference speed of edge devices cannot compare with that of servers, it can still meet the needs of real-time inference.

[0039] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A server-side lightweight method based on YOLOv8, characterized in that: Including model pruning and knowledge distillation.

2. A server-side lightweight method based on YOLOv8 according to claim 1, characterized in that: The model pruning is specifically amplitude-based inter-layer adaptive pruning.

3. A server-side lightweight method based on YOLOv8 according to claim 1, characterized in that: The model pruning specifically includes: In a feedforward neural network with depth d, each layer has a corresponding weight tensor W (1) ,...,W (d) . To uniformly calculate the LAMP score, which is used to evaluate the importance of connections, these weight tensors are flattened into one-dimensional vectors. In the flattened one-dimensional vector, when the weights are sorted in ascending order according to a given index, for any u < v, there is |W[u] ≤ W[v]| for the weight tensors. Then the LAMP score at the u-th index is defined as: Among them, the weight of tensor W mapped by index u is W[u].

4. A server-side lightweight method based on YOLOv8 according to claim 1, characterized in that: The pruning ratio of the model pruning is 1:1.

7.

5. A server-side lightweight method based on YOLOv8 according to claim 1, characterized in that: The knowledge distillation is specifically response-based knowledge distillation, which uses the output layer response of the teacher model to train the student model.

6. A server-side lightweight method based on YOLOv8 according to claim 1, characterized in that: It also includes deploying the model after model pruning and knowledge distillation on the Jetson AGX Orin 32G development kit and evaluating it on the test set.

Citation Information

Patent Citations

  • Knowledge distillation method based on model after neov5 pruning and application of knowledge distillation method

    CN116384438A

  • Lightweight target detection method for aerial image of unmanned aerial vehicle based on YOLOv7-tiny

    CN118262255A

  • Small target detection tracking method and system based on improved YOLOv8

    CN118675070A