End side infrared image target detection method based on knowledge distillation

By using knowledge distillation technology in infrared image object detection, students' models are trained to imitate the teacher's model, solving the problem of high consumption of edge equipment computing resources, achieving high-precision infrared object detection, and reducing the model's parameter volume and memory usage.

CN120032104APending Publication Date: 2025-05-23SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082417.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing infrared image object detection algorithm consumes a lot of computing resources on edge devices, making it difficult to achieve high-precision detection. At the same time, the resolution of civil thermal imaging cameras is not high, resulting in low resolution and poor details of infrared images, making it difficult to achieve high-precision object detection.

Method used

The end-side infrared image object detection method based on knowledge distillation is adopted. By training the student model to imitate the teacher model, the key features of the teacher model are distilled into the student model, reducing the model's parameter amount and memory usage, while improving the detection accuracy.

Benefits of technology

Without increasing the parameter amount and memory usage, the accuracy and detection accuracy of the infrared object detection model are improved, and the performance of the infrared object detection model at the edge devices is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032104A_ABST
    Figure CN120032104A_ABST
Patent Text Reader

Abstract

The invention relates to an end-side infrared image target detection method based on knowledge distillation, and the method comprises the following steps: converting an existing high-resolution infrared image data set into a 640 * 640 resolution, converting a labeling file format into a yolov8 training format, and dividing a training set and a test set; the model is based on a yov8 network structure, and a deep neural network model is constructed and comprises a backbone network feature extraction module, a neck feature fusion module and a feature detection head module. Training a deep neural network model of a teacher: training a teacher model by adopting a yolov8-x network and a yolov8-s network respectively; training a deep neural network model of students: adopting a minimum-scale network yov8-n as a student model, distilling key features of the teacher model from the teacher model to the student model for training, and storing the model; and obtaining a detection result of the infrared image according to the output of the detection head module of the student model. According to the method, the accuracy of infrared image detection can be improved and the use efficiency of the end side equipment can be improved without increasing the computing power and memory occupation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and infrared image processing technology, and specifically to an end-side infrared image target detection method based on knowledge distillation, which is used to improve the infrared image target detection accuracy while reducing the computing resource consumption of the end-side device as much as possible. Background Art

[0002] With the continuous development of science and technology, infrared target detection technology is increasingly widely used in military, aviation, medical and other fields. Infrared target detection is an important research direction in the field of computer vision. It is a method of detecting targets in infrared images using infrared thermal imaging technology.

[0003] Although the optical cameras used in daily life can flexibly obtain high-resolution images from optical sensors, the imaging performance will drop sharply at night and in complex natural environments such as smoke, mist, and dust, making it extremely difficult to identify target objects. In the above environments, the imaging effect of infrared detectors is more stable, so thermal imaging cameras can be used to detect the thermal radiation emitted by objects in the long-wavelength infrared (LWIR) spectrum range. However, due to the low resolution of general civilian thermal imaging cameras, the resolution of infrared images is not high and the details are poor, which poses a challenge to target detection. Therefore, it is crucial to achieve high-precision detection of infrared targets.

[0004] With the support of various deep learning methods in recent years, deep neural networks have made great breakthroughs in the accuracy of target detection tasks, basically replacing traditional target detection algorithms. For example, the R-CNN series of convolutional neural networks and the YOLO (You Only Look Once) series of networks, which are representative of single-stage networks, have achieved performance that exceeds traditional methods in target detection tasks.

[0005] The more advanced the deep neural network, the more computing power it requires and the more memory it occupies, which limits its use in edge devices such as self-driving cars and drones. The YOLO series of algorithms uses a single neural network to complete target detection end-to-end, which has the advantages of real-time performance and small memory usage, and is suitable for edge devices with limited computing power. In addition, in order to solve the above problems, researchers have proposed many model acceleration and compression technologies, including pruning, quantization, lightweight model design, and knowledge distillation, to help neural network models be deployed on edge devices.

[0006] Knowledge distillation, also known as teacher-student learning, is an effective model compression and model accuracy improvement technology. Its purpose is to train the student model to imitate the teacher model and transfer the parameterized teacher knowledge to the lightweight student model. By training the student to imitate the teacher's logits or features, the student can inherit knowledge from the teacher to achieve higher accuracy. Due to the simplicity and effectiveness of knowledge distillation, it is widely used in model compression and model accuracy improvement.

[0007] In response to the problems encountered when deploying deep neural networks on edge devices, knowledge distillation technology is used to improve the target detection algorithm of infrared images. The student model with smaller memory occupancy and computing power consumption can learn the knowledge from the already trained teacher model with larger scale and higher accuracy. The student model can search for a model that is closer to the actual distribution in the hypothesis space, thereby improving the accuracy of the student model and reducing the number of parameters, and ultimately improving the performance of the infrared target detection model on edge devices. Summary of the invention

[0008] The purpose of the present invention is to address the shortcomings of the existing technology and propose a new end-side infrared image target detection method based on knowledge distillation, which can further reduce the number of parameters and memory usage of the model while improving the accuracy of infrared image target detection, and ultimately improve the performance of the infrared target detection model on the edge device.

[0009] The method of the present invention first converts the existing high-resolution infrared image data set into a resolution of 640×640, standardizes the annotation file format, converts it into the yolov8 training format, divides the training set and the test set, and then, for a given infrared image target detection network, randomly selects one of the following data enhancement methods for image enhancement according to the set probability: Mosaic, Mixup, RandomPerspective, Scale and Flip, etc. for the training set of the input network. Subsequently, a deep neural network model is constructed: including a backbone network feature extraction module, a neck feature fusion module and a feature detection head module. According to the model structure, the hyperparameters are set, and the teacher model Teacher is trained using the yolov8-x and yolov8-s networks respectively, and the model is saved. Finally, the hyperparameters are set, and yolov8-n is used as the student model Student, and the key features of the teacher model are distilled from the teacher model to the student model Student for training, and the student model network weight is updated by the back propagation algorithm until the set maximum number of iterations is reached, and the student network model is saved, and the detection result of the infrared image is obtained according to the output of the detection head module of the student model Student. The parameter comparison of the teacher model Teacher and the student model Student is shown in Table 1:

[0010]

[0011] Table 1 Comparison of parameters of yolov8 teacher and student models

[0012] In the above method, preparing the infrared image training data set includes: HR Image resized to 640×640 resolution using Letter Box Resize L s R . Standardize the original COCO annotation format of the dataset, convert it into yolov8 format, and divide it into training set and test set. The process can be described as:

[0013]

[0014] in, Indicates a letterbox zoom operation.

[0015] In the above methods, for infrared image target detection networks of different scales, According to a certain probability, one of the following centralized data enhancement methods is randomly selected for image enhancement to improve the nonlinear fitting ability of the network: Mosaic, Mixup, RandomPerspective, Scale and Flip. The process can be expressed as:

[0016]

[0017] in, represents the selected data augmentation operation, Represents the enhanced image.

[0018] In the above method, a deep neural network model is constructed: it includes a backbone network feature extraction module, a neck feature fusion module and a feature detection head module. The module structure diagram is shown in Figure 2As shown in the figure. The network model uses CSP (Cross Stage Partial) Darknet as the backbone network, including CLBS (Convolution + Linear Deformable Convolution + Batch Normalization + SiLU) modules and C2f (Cross Stage Partial and 2-Fold Aggregation) modules; the CLBS module is an optimization of the original CBS module (Convolution + Batch Normalization + SiLU) in the yolov8 network model. After the input of the CBS module passes through a standard convolution block, the original standard convolution block is replaced by a linear deformable convolution module (LDConv). There are two defects in the standard convolution operation. On the one hand, the convolution operation is limited to the local window and cannot capture information from other locations; on the other hand, the size of the convolution kernel is fixed to k*k, that is, a square shape, and the number of parameters increases quadratically with the increase of size. Although the deformable convolution solves the fixed sampling problem of the standard convolution, the number of its parameters still increases quadratically, and the impact of different initial sampling shapes on network performance has not been explored. LDConv provides convolution kernels with an arbitrary number of parameters and arbitrary sampling shapes, thus providing richer choices for the trade-off between network overhead and performance. With appropriate parameter selection, it can reduce the number of network parameters while maintaining the original performance. In LDConv, a new coordinate generation algorithm is defined to generate different initial sampling positions to accommodate convolution kernels of arbitrary sizes. In order to adapt to changing targets, an offset is introduced at each position to adjust the sampling shape, which can effectively correct the square growth trend of the number of parameters in standard convolution and deformable convolution to linear growth. The infrared image is processed by CLBS module to obtain a first CLBS result, and the first CLBS result is extracted by C2f module to obtain a first feature, a second feature and a third feature; the first feature, the second feature and the third feature are processed by CLBS to obtain a first fusion feature, a second fusion feature and a third fusion feature; the first fusion feature, the second fusion feature and the third fusion feature are classified and located respectively by the detection head network to obtain a first target detection result, a second target detection result and a third target detection result respectively; the first target detection result, the second target detection result and the third target detection result are formed into a detection result set to obtain a target detection result.

[0019] In the above method, the teacher model is trained: Yolov8-x and Yolov8-s are used as the teacher model Teacher respectively in contrast to the constructed deep neural network model. After determining the optimizer, learning rate and maximum number of iterations used for training, start training the teacher network; after each forward propagation, calculate the loss between the training result and the true label (Ground Truth), and then update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached, thus completing the teacher model Teacher training.

[0020] In the above method, the student model is trained: the module structure diagram is as follows Figure 3 As shown in the figure, compared with the constructed neural network structure, the network with the minimum number of parameters yolov8-n is used as the student model Student, and yolov8-x and yolov8-s trained in the previous step are used as the teacher model Teacher. Before the training starts, the teacher model is passed in, and the strategy of knowledge distillation is adopted to guide the student model to train. During the training process, the gradient of the detection loss with respect to the feature is used to indicate the degree of influence of the feature on the final detection result. For features with larger gradients, they have a greater impact on the decision-making process, so more attention should be paid to them in the knowledge distillation process. For the kth feature map (Feature Map) of the lth layer of the neural network, the weight matrix is ​​defined as follows:

[0021]

[0022] in, represents the total detection loss, including bounding box regression loss and classification loss, is the single activation value at position index (i, j) in the k-th feature map of layer l, W represents the width of the feature map, and H represents the height of the feature map. First, calculate Relative to the feature The gradient of the back-propagated gradient is then globally averaged pooled in the width and height dimensions to obtain the weight matrix of the feature channel use To weight the kth activation map:

[0023]

[0024] A k l represents the kth gradient weighted activation map of the lth layer. Perform linear combination in the channel dimension and then normalize to obtain the final target feature map for distillation:

[0025]

[0026] Among them, C represents the number of channels and Norm represents the normalization function. By using the gradient to weight the features, the features that have a greater impact on the overall detection loss can be effectively highlighted. Applying the above process to the teacher model and the student model, the target feature maps of the teacher model and the student model are obtained: The total difference between the target feature maps of the teacher model and the student model is:

[0027]

[0028] Among them, L is the total number of intermediate layers used for distillation, l represents an intermediate layer, and the goal of knowledge distillation is to minimize the difference between the target feature maps In order to enhance the knowledge transfer between different scales, the output layer of the Spatial Pyramid Pooling Fast module in the neural network model is selected as the target layer of knowledge distillation.

[0029] In the above method, the student model is saved by solidifying the network weights of the group with the highest evaluation index in the evaluation process of the two groups of student network training and saving them as the final student network model. represents the student model trained by yolov8-x distillation, Represents the student model trained by yolov8-s distillation. Compare the accuracy and parameter quantity of the two models and select the better one as the final model.

[0030] Compared with the prior art, the method of the present invention has the following obvious outstanding substantive features and significant technical progress:

[0031] 1) The data enhancement methods such as Mosaic, Mixup, RandomPerspective, Scale and Flip are used to enhance the data of infrared images, which more effectively enhances the nonlinear fitting ability of the data set and the network.

[0032] 2) Letterbox scaling was used to adjust the resolution of high-resolution infrared images to 640×640, maintaining the original aspect ratio of the image and avoiding image distortion caused by rough scaling.

[0033] 3) The standard convolution in the neural network model is replaced by linear deformable convolution, which can reduce the number of parameters and the number of operations of the network model while maintaining the performance unchanged.

[0034] 4) The knowledge distillation method is used to train the student network model, so that the student model can improve the accuracy and detection accuracy of the model without increasing the number of parameters and memory usage, and ultimately improve the performance of the infrared target detection model on edge devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A flowchart of the method of the present invention;

[0036] Figure 2 A diagram of the deep neural network structure of the method of the present invention;

[0037] Figure 3 Flow chart of knowledge distillation of the method of the present invention. DETAILED DESCRIPTION

[0038] The preferred embodiments of the present invention are described in detail as follows in conjunction with the accompanying drawings:

[0039] See also Figure 1 The present invention provides a method for detecting target in infrared images on the edge based on knowledge distillation. The specific operation steps are as follows:

[0040] 1) Generate training dataset: The existing Flir Thermal Dataset high-resolution infrared images are resized to 640×640 resolution through letter box resize. The original COCO annotation format of the dataset is standardized, converted into yolov8 format, and divided into training set and test set.

[0041] 2) Data enhancement: For infrared image target detection networks of different scales, one of the following enhancement methods is randomly selected for image enhancement according to a certain probability to improve the nonlinear fitting ability of the network: Mosaic, Mixup, RandomPerspective, Scale and Flip.

[0042] 3) Build a deep neural network model: The network model structure diagram is as follows Figure 2As shown in the figure, it includes a backbone network feature extraction module, a neck feature fusion module and a feature detection head module. CSPDarknet is used as the backbone network, including CLBS (Convolution + Linear Deformable Convolution + Batch Normalization + SiLU) module and C2f (Cross Stage Partial and 2-Fold Aggregation) module; the CLBS module is an optimization of the original CBS module (Convolution + Batch Normalization + SiLU) in the yolov8 network model. After the input of the CBS module passes through a standard convolution block, the original standard convolution block is replaced by a linear deformable convolution module (LinearDeformable Convolution, LDConv). There are two defects in the standard convolution operation. On the one hand, the convolution operation is limited to the local window and cannot capture information from other positions; on the other hand, the size of the convolution kernel is fixed to k*k, that is, a square shape, and the number of parameters increases quadratically with the increase of size. Although Deformable Convolution solves the fixed sampling problem of standard convolution, its number of parameters still grows quadratically, and the impact of different initial sampling shapes on network performance has not been explored. LDConv provides an arbitrary number of parameters and arbitrary sampling shapes for the convolution kernel, thus providing a richer choice for the trade-off between network overhead and performance. With appropriate parameter selection, it can reduce the number of network parameters while maintaining the original performance. In LDConv, a new coordinate generation algorithm is defined to generate different initial sampling positions to accommodate convolution kernels of arbitrary sizes. In order to adapt to the changing target, an offset is introduced at each position to adjust the sampling shape, which can effectively correct the quadratic growth trend of the number of parameters in standard convolution and deformable convolution to linear growth. The infrared image is processed by CLBS module to obtain a first CLBS result, and the first CLBS result is extracted by C2f module to obtain a first feature, a second feature and a third feature; the first feature, the second feature and the third feature are processed by CLBS to obtain a first fusion feature, a second fusion feature and a third fusion feature; the first fusion feature, the second fusion feature and the third fusion feature are classified and located respectively by the detection head network to obtain a first target detection result, a second target detection result and a third target detection result respectively; the first target detection result, the second target detection result and the third target detection result are formed into a detection result set to obtain a target detection result.

[0043] 4) Train the teacher model: Compare the constructed deep neural network model, use yolov8-x and yolov8-s as the teacher model Teacher. Determine the optimizer, learning rate and maximum number of iterations used for training, and start training the teacher network; after each forward propagation, calculate the loss between the training result and the true label (Ground Truth), and then update the network parameters through the back propagation algorithm; repeat the above steps until the preset maximum number of iterations is reached, and the teacher model Teacher is completed. x and Teacher s train.

[0044] 5) Train the student model: Figure 3 As shown in the figure, compared with the constructed neural network structure, the minimum parameter network yolov8-n is used as the student model Student, and the Teacher trained in the previous step is used respectively. x and Teacher s As the teacher model Teacher. Before the training starts, the teacher model Teacher is passed in, and the knowledge distillation strategy is adopted to guide the student model to train. During the training process, the gradient of the detection loss with respect to the feature is used to indicate the degree of influence of the feature on the final detection result. For features with larger gradients, they have a greater impact on the decision process, so more attention should be paid to them during the knowledge distillation process. For the kth feature map (Feature Map) of the lth layer of the neural network, the weight matrix is ​​defined as follows:

[0045]

[0046] in, represents the total detection loss, including bounding box regression loss and classification loss, is the single activation value at position index (i, j) in the k-th feature map of layer l, W represents the width of the feature map, and H represents the height of the feature map. First, calculate Relative to the feature The gradient of the back-propagated gradient is then globally averaged pooled in the width and height dimensions to obtain the weight matrix of the feature channel use To weight the kth activation map:

[0047]

[0048] represents the kth gradient weighted activation map of the lth layer. Perform linear combination in the channel dimension and then normalize to obtain the final target feature map for distillation:

[0049]

[0050] Among them, C represents the number of channels and Norm represents the normalization function. By using the gradient to weight the features, the features that have a greater impact on the overall detection loss can be effectively highlighted. Applying the above process to the teacher model and the student model, the target feature maps of the teacher model and the student model are obtained: The total difference between the target feature maps of the teacher model and the student model is:

[0051]

[0052] Among them, L is the total number of intermediate layers used for distillation, l represents an intermediate layer, and the goal of knowledge distillation is to minimize the difference between the target feature maps In order to enhance the knowledge transfer between different scales, the output layer of the Spatial Pyramid Pooling Fast module in the neural network model is selected as the target layer of knowledge distillation.

[0053] 6) Save the network: Solidify the network weights of the group with the highest evaluation index in the evaluation process of the two groups of student network training and save them as the final student network model represents the student model trained by yolov8-x distillation, Represents the student model trained by yolov8-s distillation. Compare the accuracy and parameter quantity of the two models and select the better one as the final model. The evaluation results are shown in Table 2:

[0054]

[0055]

[0056] Table 2 The evaluation results are shown in Table 2 compared with Table 1. The model trained with linear variability convolution and knowledge distillation strategy Both have higher accuracy than the YOLOv8-n model trained alone, which quantitatively demonstrates the effectiveness of the present invention, and the number of parameters of the student model is reduced by 8%, and the number of floating-point operations is reduced by 9.8%. Comparing the distillation results of the two teacher models, the YOLOv8-n student model trained by distillation of the YOLOv8-s teacher model has higher detection accuracy while maintaining the same number of parameters and calculations, and can be used as the final end-side deployment model. This result also shows that a teacher model with a large gap between the number of parameters and the amount of calculation and the student model is not necessarily suitable as a teacher model.

[0057] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A method for edge-side infrared image target detection based on knowledge distillation, characterized in that: The following steps are involved: S1) Convert the existing infrared image dataset into a resolution of 640×640, convert the annotation file format into the yolov8 training format, and divide it into a training set and a test set; S2) For a given infrared image target detection network, a data enhancement method among Mosaic, Mixup, RandomPerspective, Scale and Flip is randomly selected according to a set probability for the training set of the input network to perform image enhancement; S3) constructing a deep neural network model, wherein the deep neural network model includes a backbone network feature extraction module, a neck feature fusion module and a feature detection head module; S4) training the teacher's deep neural network model, including setting hyperparameters, using yolov8-x and yolov8-s networks to train the teacher model respectively, and saving the teacher model; S5) training the student's deep neural network model, including setting hyperparameters, using the minimum parameter network yolov8-n as the student model, distilling the key features of the teacher model from the teacher model to the student model for training, and updating the student model network weights by a back propagation algorithm until a set maximum number of iterations is reached; S6) Save the student network model, and obtain the detection result of the infrared image according to the output of the detection head module of the student model.

2. The method for edge-side infrared image target detection based on knowledge distillation according to claim 1, characterized in that: The step S1 also includes: performing a high-resolution infrared image I HR Image resized to 640×640 resolution using letterbox scaling Convert the original COCO annotation format of the dataset into yolov8 format and divide it into training set and test set. The process can be described as follows: in, Indicates a letterbox zoom operation.

3. The method for edge-side infrared image target detection based on knowledge distillation according to claim 1, characterized in that: The step S2 also includes randomly selecting one of the enhancement methods of Mosaic, Mixup, RandomPerspective, Scale and Flip for image enhancement according to a certain probability. The process can be described as follows: in, represents the selected data augmentation operation, Represents the enhanced image.

4. The method for edge-side infrared image target detection based on knowledge distillation according to claim 1, characterized in that: The deep neural network model in step S3 uses CSP Darknet as a backbone network, and the CSP Darknet includes a CLBS module and a C2f module; wherein the input of the CLBS module passes through a standard convolution block, and then the standard convolution block is replaced by a linear deformable convolution module; the infrared image is subjected to CLBS processing by the CLBS module to obtain a first CLBS result, and the first CLBS result is subjected to feature extraction by the C2f module to obtain a first feature, a second feature, and a third feature; the first feature, the second feature, and the third feature are subjected to CLBS processing to obtain a first fusion feature, a second fusion feature, and a third fusion feature; the first fusion feature, the second fusion feature, and the third fusion feature are respectively classified and located by the detection head network, and a first target detection result, a second target detection result, and a third target detection result are correspondingly obtained, and the first target detection result, the second target detection result, and the third target detection result are formed into a detection result set to obtain a target detection result.

5. The method for edge-side infrared image target detection based on knowledge distillation according to claim 1, characterized in that: The step S4 adopts the neural network structure constructed in step S3, and uses yolov8-x and yolov8-s as teacher models respectively; determines the optimizer, learning rate and maximum number of iterations used in training, and starts training the teacher network; After each forward propagation is completed, the loss between the training result and the true label is calculated, and then the network parameters are updated through the back-propagation algorithm; Repeat the above steps until the maximum number of iterations is reached to complete the teacher model Teacher x and Teacher s train.

6. The method for detecting target in edge-side infrared images based on knowledge distillation according to claim 5, characterized in that: The step S5 adopts the neural network structure constructed in step S3, uses the minimum parameter network yolov8-n as the student model, and uses the Teacher trained in the previous step. x and Teacher s As a teacher model; before the training starts, the step teacher model is passed in, and the strategy of knowledge distillation is adopted to guide the student model to train; During the training process, the gradient of the detection loss with respect to the feature is used to represent the degree of influence of the feature on the final detection result. For the kth feature map of the lth layer of the neural network, the weight matrix is ​​defined as follows: in, represents the total detection loss, including bounding box regression loss and classification loss, is a single activation value at position index (i, j) in the k-th feature map of layer l, and W and H represent the width and height of the feature map, respectively.

7. The method for detecting target in edge-side infrared images based on knowledge distillation according to claim 6, characterized in that: The weight matrix is ​​first calculated Relative to the feature The gradient of the back-propagated gradient is then globally averaged pooled in the width and height dimensions to obtain the weight matrix of the feature channel use To weight the k-th activation map: represents the kth gradient weighted activation map of the lth layer; then, Perform linear combination in the channel dimension and then normalize to obtain the final target feature map for distillation: Where C represents the number of channels and Norm represents the normalization function. Apply the above process to the teacher model and the student model to obtain the target feature maps of the teacher model and the student model: and The total difference between the target feature maps of the teacher model and the student model is: Among them, L is the total number of intermediate layers used for distillation, l represents an intermediate layer, and the goal of knowledge distillation is to minimize the difference between the target feature maps The output layer of the spatial pyramid pooling module in the neural network model is selected as the target layer for knowledge distillation.

8. The method for edge-side infrared image target detection based on knowledge distillation according to claim 1, characterized in that: The step S6 solidifies the network weights of the group with the highest evaluation index in the evaluation process of the two groups of student network training respectively, and saves them as the final student network model and represents the student model trained by yolov8-x distillation, Represents the student model trained by yolov8-s distillation. Compare the accuracy and parameter quantity of the two models and select the better one as the final model.