Lightning arrester identification method and device based on improved YOLOv3-Tiny and knowledge distillation
By improving the YOLOv3-Tiny network structure and introducing knowledge distillation technology, the dilemma between recognition accuracy and model complexity of the lightning arrester infrared image recognition model is solved, and efficient deployment on edge devices and improvement of all-weather recognition capabilities are achieved.
Patent Information
- Application Number
- CN202510230130.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
The existing infrared image recognition model of lightning arrester has a dilemma between recognition accuracy and model complexity, and it is difficult to effectively deploy on edge devices and cannot meet the identification needs of power equipment in all weather conditions.
By improving the YOLOv3-Tiny network structure, the Fuse-MBConv module and the adaptive feature recalibration mechanism were introduced, and combined with knowledge distillation technology, YOLOv11 is used as the teacher model and the improved YOLOv3-Tiny is used as the student model to achieve efficient compression and performance improvement of the model.
It significantly improves the accuracy and inference speed of infrared image recognition of lightning arresters, reduces the model's computing resource usage, and realizes efficient deployment on edge devices, meeting the identification needs of power equipment in all weather conditions.
Smart Images

Figure CN120071011A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of lightning arrester identification methods, and in particular relates to a lightning arrester identification method and device based on improved YOLOv3-Tiny and knowledge distillation. Background Art
[0002] As a key protection device for the stability and reliability of the power system, lightning arresters ensure the safe operation of power equipment by quickly conducting current and limiting overvoltage. They are an indispensable and important inspection object in power inspections. However, my country's transmission lines are mostly distributed in harsh environments such as mountain canyons and desert Gobi. Traditional manual inspections not only face huge safety risks, but also have many problems such as low efficiency. For this reason, intelligent inspection methods represented by drone inspections and online monitoring have emerged. Drones are equipped with a variety of sensor equipment such as infrared thermal imagers, which can collect images and operating data of power equipment in all directions. Even in complex terrain and extreme weather conditions, they can effectively ensure the continuity and reliability of inspection work, and provide strong support for the safe operation of the power grid. Therefore, accurate recognition of lightning arrester image targets has become a key issue that needs to be solved in power intelligent inspections, which is of great significance to improving the efficiency of drone intelligent inspections.
[0003] In recent years, with the rapid development of computer vision and deep learning technologies, in the field of power line inspection, object detection algorithms based on deep learning have made remarkable progress in the field of image recognition. The literature NGUYEN V N, JENSSEN R and ROVERSO D. Intelligent monitoring and inspection of power line components powered by UAVs and deep learning[J]. IEEE Power and Energy Technology Systems Journal, 2019, 6(1): 11-21 combines UAV technology with deep learning to achieve the detection and intelligent monitoring of power line components; the literature LIU Xinyu, MIAO Xiren, JIANG Hao, et al. Data analysis in visual power line inspection: an in-depth review of deep learning for component detection and fault diagnosis[J]. Annual Reviews in Control, 2020, 50: 253-277 proposes a fault diagnosis method based on image recognition through in-depth analysis of the visual images of power equipment; the literature HUANG Xiaoning, ZHANG Zhenliang. A method to extract insulator image from aerial image of helicopter patrol[J]. Power System Technology, 2010, 34(1): 194-197 focuses on insulation detection, combines the images collected by UAVs and deep learning algorithms, and constructs an intelligent insulator fault recognition system.The above-mentioned literature conducts target detection on the visible light images of power equipment. However, visible light detection is only applicable during the day with good weather conditions and is difficult to meet the requirements of all-weather power training. In the current research on infrared image recognition of power equipment, the vast majority of models used are from the YOLO series. However, compared with visible light images, infrared images have disadvantages such as low resolution, high noise, and blurred or missing edge information. It is difficult to achieve accurate target detection of arrester infrared images using basic YOLO series models. Therefore, it is necessary to make targeted improvements to the existing target detection models to improve their detection performance for arrester infrared images.
[0004] At the same time, since the arrester infrared image target detection model needs to be deployed on edge inspection devices such as drones, limited by the computing resources of edge devices, it is difficult to deploy target detection models with complex structures. Lightweighting can solve the problem of difficult deployment of models with complex structures, but the current research content only modifies the YOLO series of target recognition models, and the degree of model compression is limited. Therefore, in view of the above problems, it is necessary to study a lightweight recognition model for arrester infrared images to deploy it on the edge devices of the drone inspection system and achieve lightweight recognition of arrester infrared images. Summary of the Invention
[0005] To solve the problems existing in the background technology, the present invention adopts the following technical solutions:
[0006] An arrester recognition method based on improved YOLOv3-Tiny and knowledge distillation, the method comprising the following steps:
[0007] Improve YOLOv3-Tiny, the improvement includes improving YOLOv3-Tiny based on the Fuse-MBConv module and improving the training strategy of YOLOv3-Tiny; the Fuse-MBConv module includes a fused convolution operation and an SE module, and by fusing depthwise separable convolution and expansion convolution into a conventional convolution operation, the amount of calculation and the number of parameters are reduced; when improving the training strategy of YOLOv3-Tiny, a progressive learning strategy is used to dynamically adjust the image size and regularization strength to optimize the training efficiency while improving the model performance;
[0008] Establish an arrester infrared image recognition knowledge distillation model with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain an arrester infrared image recognition knowledge distillation model based on YOLOv11 and the improved YOLOv3-Tiny;
[0009] Use the knowledge distillation model for lightning arrester infrared image recognition based on YOLOv11 and improved YOLOv3-Tiny to recognize the collected lightning arrester infrared images.
[0010] Furthermore, when improving YOLOv3-Tiny based on the Fuse-MBConv module, in terms of the network structure, by introducing the Fuse-MBConv module after the feature extraction layer, its efficient feature fusion mechanism is used to enhance the network's feature extraction ability; at the same time, the SE module is integrated at different scales of the feature pyramid network to optimize the feature representation through the adaptive channel attention mechanism.
[0011] Furthermore, the method for improving YOLOv3-Tiny based on the Fuse-MBConv module is as follows:
[0012] Let the input feature map be where c is the number of channels, and h and w are the height and width of the feature map;
[0013] The fused convolution operation processes the input features: First, a 3×3 convolutional kernel is used to extract spatial features of the input features; subsequently, the batch normalization layer is passed through to stabilize the feature distribution and accelerate the training process; then the SiLU activation function is used to introduce non-linear transformation; finally, the 1×1 convolution is used to adjust the feature channel dimension to achieve the recombination of cross-channel information; its mathematical expression is:
[0014] F fused (X) = Conv 1×1 (SiLU(BN(Conv 3×3 (X))))(1)
[0015] In the formula, Conv is the convolution operation; BN is the batch normalization; SiLU is the activation function;
[0016] The SE module performs adaptive feature recalibration by explicitly modeling the interdependence between channels: First, the features in the spatial dimension are compressed into a channel descriptor through global average pooling to capture global context information; subsequently, two fully connected layers and non-linear activation functions are used to construct the non-linear relationship between channels to generate the importance weights of each feature channel; finally, these weights are multiplied by the original features to achieve channel-level adaptive feature recalibration; its mathematical expression is:
[0017]
[0018] In the formula, z c is the statistical value of the c-th channel; u c represents the feature map of the c-th channel.
[0019] s = F ex(z, W) = σ(W 2 δ(W 1 z))(3)
[0020] where s is the channel weight; is the weight of the first fully connected layer; is the weight of the second fully connected layer; δ represents the ReLU activation function; r is the dimensionality reduction ratio;
[0021]
[0022] where s c is the weight of the c-th channel, is the feature map of the c-th channel after recalibration.
[0023] Furthermore, the method for improving the YOLOv3-Tiny training strategy is as follows:
[0024] During training, the image size gradually increases: at the beginning of training, smaller-sized images are used, and as training progresses, the image size is gradually increased until the target size is reached; the adjustment function for the image size is formula (5):
[0025]
[0026] where d t is the image size at the t-th step; d min is the minimum image size; d max is the maximum image size; T d is the total number of steps for size increase; t is the current training step;
[0027] During training, the regularization strength gradually increases: at the beginning of training, weak regularization is used, and as the image size increases, the regularization strength is increased synchronously; the adjustment function for the regularization strength is formula (6):
[0028]
[0029] where λ t is the regularization strength at the t-th step; λ min is the minimum regularization strength; λ max is the maximum regularization strength; T λ is the total number of steps for regularization adjustment.
[0030] Furthermore, before establishing a knowledge distillation model for identifying infrared images of lightning arresters with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, it is necessary to design the loss function of the knowledge distillation model. The designed loss function includes a classification loss function, a regression loss function, and a feature map distillation loss function, and the total loss function of knowledge distillation is obtained by weighted summation.
[0031] Furthermore, the classification loss function is as follows: The student model needs to learn the true labels and at the same time imitate the classification probability distribution of the teacher model:
[0032]
[0033] In the formula, L CE is the cross-entropy loss of the student model for the true labels; T is the distillation temperature, which is used to smooth the probability distribution; z t and z s are the classification output logits of the teacher model and the student model respectively; α is the weight of the classification loss and the distillation loss; KL is the Kullback-Leibler divergence calculation; is the probability distribution obtained by scaling z t by temperature and then passing through the sigmoid function (σ); is the probability distribution obtained by scaling z s by temperature and then passing through the sigmoid function (σ);
[0034] The regression loss function is as follows: The bounding box regression of the student model should be close to the bounding box predicted by the teacher model
[0035]
[0036] In the formula, and are the predictions of the student and teacher models for the i-th bounding box respectively; N is the number of bounding boxes; SmoothL1 is the smooth L1 loss function used for the regression error.
[0037] The feature map distillation loss function is as follows: The intermediate layer features of the student network need to imitate the features of the corresponding layer in the teacher network
[0038]
[0039] In the formula, and are the feature maps of the student and teacher models at the i-th position respectively.
[0040] The above three loss functions are weighted and summed to obtain the total loss function:
[0041] L total= λ 1 ·L cls + λ 2 ·L reg + λ 3 ·L feat (10)
[0042] Wherein, λ 1 , λ 2 and λ 3 respectively represent the weights of the classification loss function, the regression loss function, and the feature distillation loss function.
[0043] Furthermore, a method for establishing a knowledge distillation model for identifying infrared images of lightning arresters with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, training the knowledge distillation model, and obtaining the knowledge distillation model for identifying infrared images of lightning arresters based on YOLOv11 and the improved YOLOv3-Tiny is as follows:
[0044] ①Precisely label the positions of lightning arresters in the collected infrared image dataset, and input the labeled images into the teacher model and the student model respectively to carry out model training and target recognition;
[0045] ②First, pre-train the teacher model, and use the original image data with labeled positions to train the YOLOv11 model until convergence; then fix the parameters of the teacher model and generate soft labels including the classification probability distribution, the bounding box regression value, and the intermediate feature map;
[0046] ③After initializing the student model, start training. The improved YOLOv3-Tiny model is simultaneously supervised by both soft labels and original hard labels during the training process;
[0047] ④During the training process, evaluate the convergence of the model by monitoring the loss function of the knowledge distillation model: if the convergence condition is not reached, adjust the network parameters and continue training; once the model converges, terminate the training and save the parameters of the current knowledge distillation model.
[0048] A lightning arrester identification device based on the improved YOLOv3-Tiny and knowledge distillation includes:
[0049] The YOLOv3-Tiny improvement module is used to improve YOLOv3-Tiny. The improvements include improving YOLOv3-Tiny based on the Fuse-MBConv module and improving the YOLOv3-Tiny training strategy. The Fuse-MBConv module includes a fused convolution operation and an SE module. By fusing depthwise separable convolution and expanded convolution into a conventional convolution operation, the amount of computation and the number of parameters are reduced. When improving the YOLOv3-Tiny training strategy, a progressive learning strategy is used to dynamically adjust the image size and regularization strength, optimizing the training efficiency while improving the model performance.
[0050] The lightning arrester infrared image recognition knowledge distillation model acquisition module is used to establish a lightning arrester infrared image recognition knowledge distillation model with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain a lightning arrester infrared image recognition knowledge distillation model based on YOLOv11 and the improved YOLOv3-Tiny.
[0051] The lightning arrester infrared image recognition module is used to recognize the collected lightning arrester infrared images by using the lightning arrester infrared image recognition knowledge distillation model based on YOLOv11 and the improved YOLOv3-Tiny.
[0052] A non-transitory computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the lightning arrester recognition method based on the improved YOLOv3-Tiny and knowledge distillation as described above.
[0053] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the lightning arrester recognition method based on the improved YOLOv3-Tiny and knowledge distillation as described above.
[0054] The beneficial technical effects of the present invention are as follows:
[0055] The present invention innovatively improves the structure and optimizes the training strategy of the YOLOv3-Tiny network. While keeping the model scale basically unchanged, it significantly improves the recognition accuracy and inference speed of the model. To further optimize the model performance, the knowledge distillation technology is introduced. With the high-performance YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, efficient compression of the model is achieved. The lightweight lightning arrester infrared image recognition method based on the improved YOLOv3-Tiny and knowledge distillation of the present invention has good practical value and technical feasibility. Description of the Drawings
[0056] Figure 1 It is the improvement strategy diagram of the YOLOv3-Tiny model in Embodiment 1 of the present invention;
[0057] Figure 2 It is the flow chart of arrester infrared image recognition based on knowledge distillation in Embodiment 1 of the present invention;
[0058] Figure 3 It is the structure of the Fuse-MBConv module in Embodiment 1 of the present invention;
[0059] Figure 4 It is the mAP and Loss of four distillation models in Embodiment 1 of the present invention;
[0060] Figure 5 It is the recognition effect of the model proposed in Embodiment 1 of the present invention on the arrester infrared image. Detailed implementation manners
[0061] The following further clearly and completely describes the arrester recognition method and device based on improved YOLOv3-Tiny and knowledge distillation provided by the present invention with reference to the accompanying drawings:
[0062] Embodiment 1
[0063] As Figure 1 、 2 described, the arrester recognition method based on improved YOLOv3-Tiny and knowledge distillation includes the following steps:
[0064] Improve YOLOv3-Tiny, and the improvement includes improving YOLOv3-Tiny based on the Fuse-MBConv module and improving the training strategy of YOLOv3-Tiny; the Fuse-MBConv module includes a fused convolution operation and an SE module, and by fusing the depthwise separable convolution and the expansion convolution into a conventional convolution operation, the amount of calculation and the number of parameters are reduced; when improving the training strategy of YOLOv3-Tiny, use the progressive learning strategy to dynamically adjust the image size and regularization strength, and optimize the training efficiency while improving the model performance;
[0065] Establish a knowledge distillation model for arrester infrared image recognition with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain a knowledge distillation model for arrester infrared image recognition based on YOLOv11 and the improved YOLOv3-Tiny;
[0066] Use the knowledge distillation model for arrester infrared image recognition based on YOLOv11 and the improved YOLOv3-Tiny to recognize the collected arrester infrared images;
[0067] Specifically, when improving YOLOv3-Tiny, it includes:
[0068] 1. Improve YOLOv3-Tiny based on the Fuse-MBConv module;
[0069] The Fuse-MBConv module is a new network structure unit introduced in EfficientNetv2. It is an improvement on the traditional MBConv (mobile inverted bottleneck convolution) module. By fusing depthwise separable convolution and expansion convolution into a conventional convolution operation, the amount of computation and the number of parameters are reduced. The structure of the Fuse-MBConv module is as Figure 3 shown. The Fuse-MBConv module is divided into two parts: a fused convolution operation and an SE module:
[0070] Let the input feature map be where c is the number of channels, and h and w are the height and width of the feature map.
[0071] ① The fused convolution operation processes the input features in the following order: First, a 3×3 convolutional kernel is used to extract spatial features of the input features, which can effectively capture local spatial information; subsequently, it passes through a Batch Normalization (BN) layer to stabilize the feature distribution and accelerate the training process; then the SiLU activation function is used to introduce a non-linear transformation, which has a smoother characteristic compared to the traditional ReLU; finally, the feature channel dimension is adjusted through a 1×1 convolution to achieve cross-channel information recombination; its mathematical expression can be formalized as formula (1):
[0072] F fused (X) = Conv 1×1 (SiLU(BN(Conv 3×3 (X))))(1)
[0073] In the formula, Conv is the convolution operation; BN is batch normalization; SiLU is the activation function;
[0074] ②SE (Squeeze-and-Excitation) module: The SE module performs adaptive feature recalibration by explicitly modeling the interdependencies between channels; its operation process includes two key steps: squeezing and excitation. First, the features in the spatial dimension are compressed into channel descriptors through global average pooling to capture global context information. Subsequently, two fully connected layers and a non-linear activation function are used to construct non-linear relationships between channels, generating the importance weights for each feature channel. Finally, these weights are multiplied by the original features to achieve channel-level adaptive feature recalibration. Through the operation of the SE module, the features of important channels can be enhanced; its mathematical expression can be formalized as formula (2):
[0075]
[0076] In the formula, z c is the statistical value of the c-th channel; u c represents the feature map of the c-th channel;
[0077] s = F ex (z, W) = σ(W 2 δ(W 1 z)) (3)
[0078] In the formula, s is the channel weight; is the weight of the first fully connected layer; is the weight of the second fully connected layer; δ represents the ReLU activation function; r is the dimensionality reduction ratio;
[0079]
[0080] In the formula, s c is the weight of the c-th channel, is the feature map of the re-calibrated c-th channel;
[0081] To improve the performance of YOLOv3-Tiny in the recognition of arrester infrared images, an improved scheme based on Fuse-MBConv is proposed; in terms of network structure, Fuse-MBConv modules are introduced after the key feature extraction layers (the 3rd, 4th, and 5th convolutional layers), and its efficient feature fusion mechanism is used to enhance the network's feature extraction ability; at the same time, SE modules are integrated at different scales of the feature pyramid network to optimize feature representation through an adaptive channel attention mechanism; this improvement strategy, while maintaining the network's light weight, effectively improves the network's detection ability for arrester infrared images in complex environmental scenarios through the synergistic effect of feature fusion and attention mechanism, making the improved network show better performance in arrester recognition.
[0082] 2. Improve the training strategy of YOLOv3-Tiny
[0083] Use a progressive learning strategy to dynamically adjust the image size and regularization strength, improving the training efficiency while enhancing the model performance; the method of this training strategy is as follows:
[0084] ① During the training process, the image size increases progressively: use smaller-sized images at the beginning of training, and gradually increase the image size as training progresses until the target size is reached; the adjustment function for the image size is formula (5):
[0085]
[0086] where d t is the image size at the t-th step; d min is the minimum image size; d max is the maximum image size; T d is the total number of steps for size increase; t is the current training step;
[0087] ② During the training process, the regularization strength increases progressively: use weaker regularization at the beginning of training, and synchronously increase the regularization strength as the image size increases; the adjustment function for the regularization strength is formula (6):
[0088]
[0089] where λ t is the regularization strength at the t-th step; λ min is the minimum regularization strength; λ max is the maximum regularization strength; T λ is the total number of steps for regularization adjustment.
[0090] The adopted progressive training strategy enables the model to quickly learn the basic features of the input image at the beginning of training, and focus on optimizing the recognition ability of complex features in the later stage; this training method can not only significantly improve the detection accuracy and small target recognition performance of the model, but also effectively reduce the consumption of training time and computing resources.
[0091] This training strategy highly matches the characteristics of the YOLOv3-Tiny lightweight object detection model: it can quickly obtain the basic feature information of the arrester infrared image at the beginning of training, and focus on optimizing the recognition ability of areas with less obvious features in the later stage, especially for accurately locating abnormal features at different positions and sizes; through this progressive feature learning and targeted optimization method, the training efficiency and target recognition ability of the model are significantly improved, making it show stronger performance and higher reliability in practical applications.
[0092] It should be noted that before establishing a knowledge distillation model for the infrared image recognition of lightning arresters with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, it is necessary to design the loss function of the knowledge distillation model. The design of the knowledge distillation model loss with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model is divided into three parts: classification loss, regression loss, and feature map distillation loss, and finally the total loss function of knowledge distillation is obtained; the design of the loss function is as follows:
[0093] ① Classification loss: The student model needs to learn the true labels and at the same time imitate the classification probability distribution of the teacher model:
[0094]
[0095] In the formula, L CE is the cross-entropy loss of the student model for the true labels; T is the distillation temperature, which is used to smooth the probability distribution; z t and z s are the classification output logits of the teacher model and the student model respectively; α is the weight of the classification loss and the distillation loss; KL is the Kullback-Leibler divergence calculation; is the probability distribution obtained by scaling z t by temperature and then passing through the sigmoid function (σ); is the probability distribution obtained by scaling z s by temperature and then passing through the sigmoid function (σ);
[0096] ② Regression loss: The bounding box regression of the student model should be close to the bounding box predicted by the teacher model
[0097]
[0098] In the formula, and are the predictions of the student and teacher models for the i-th bounding box respectively; N is the number of bounding boxes; SmoothL1 is the smooth L1 loss function used for the regression error;
[0099] ③ Feature map distillation loss: The intermediate layer features of the student network need to imitate the features of the corresponding layer in the teacher network
[0100]
[0101] In the formula, and are the feature maps of the student and teacher models at the i-th position respectively;
[0102] The above three loss functions are weighted and summed to obtain the total loss function:
[0103] L total = λ 1 ·L cls + λ 2 ·L reg + λ 3 ·L feat (10)
[0104] In the formula, λ 1 , λ 2 and λ 3 respectively represent the weights of the classification loss function, the regression loss function, and the feature map distillation loss function.
[0105] The method for establishing a knowledge distillation model for identifying infrared images of lightning arresters with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, training the knowledge distillation model, and obtaining the knowledge distillation model for identifying infrared images of lightning arresters based on YOLOv11 and the improved YOLOv3-Tiny is as follows:
[0106] ①Accurately label the positions of lightning arresters in the collected infrared image dataset, and input the labeled images into the teacher model and the student model respectively to carry out model training and target recognition;
[0107] ②First, pre-train the teacher model, and use the original image data with labeled positions to train the YOLOv11 model until convergence; then fix the parameters of the teacher model and generate soft labels including the classification probability distribution, the bounding box regression value, and the intermediate feature map;
[0108] ③After initializing the student model, start training. The improved YOLOv3-Tiny model is supervised by both soft labels and original hard labels during the training process;
[0109] ④During the training process, evaluate the model convergence situation by monitoring the loss function of the knowledge distillation model: if the convergence condition is not reached, adjust the network parameters and continue training; once the model converges, terminate the training and save the parameters of the current knowledge distillation model.
[0110] Thus, a knowledge distillation model for identifying infrared images of lightning arresters based on YOLOv11 and the improved YOLOv3-Tiny is obtained, which can accurately identify the infrared images collected by the drone;
[0111] The method for identifying infrared images of lightning arresters in this embodiment can effectively solve two key problems in the current field of intelligent drone inspection: the insufficient accuracy of identifying infrared images of lightning arresters and the problem that complex models are difficult to deploy on edge devices.
[0112] For example, in this embodiment, the effects of the method of the present invention are compared and verified from different aspects:
[0113] 1. Verification of the effectiveness of model improvement:
[0114] To verify the improvement effect of the Fuse-MBConv module based on EfficientNetv2 and the training strategy optimization on YOLOv3-Tiny, three models were compared and analyzed: the original YOLOv3-Tiny, the YOLOv3-Tiny improved only by the Fuse-MBConv module, and the YOLOv3-Tiny model with the comprehensive improvement scheme. As shown in Table 1, when only the Fuse-MBConv module was introduced, the number of model parameters and size increased by 3.4% and 4.2% respectively, while the mAP increased by 6.7%. This indicates that the Fuse-MBConv module can effectively enhance the feature extraction ability of the network. After further introducing the training strategy optimization, although the number of parameters and model size increased slightly, the mAP was significantly improved to 95.1%, fully demonstrating that the comprehensive improvement scheme proposed in this paper can effectively improve the recognition ability of YOLOv3-Tiny for arrester infrared images.
[0115] In addition, it can be found from the data comparison in Table 1 that the floating-point operation amount of the improved model decreased, indicating that the improvement scheme further promoted the lightweight of the model while improving the performance. At the same time, it can be seen from the FPS index that the recognition speed of the model has been significantly improved, which shows that the improvement scheme proposed in this paper not only reduces the computational resource occupancy of the model, but also improves the recognition efficiency of the model. This double improvement in accuracy and efficiency makes the improved model more suitable for the needs of scenarios with limited computing resources.
[0116] Table 1 Comparison of model parameters and performance
[0117]
[0118] 2. Verification of the effectiveness of introducing knowledge distillation for model compression:
[0119] To verify the effectiveness of introducing the knowledge distillation method for model compression, two large models with leading performance, RT-DETR (Real-Time Detection Transformer) and YOLOv11, and two lightweight models, YOLOv3-Tiny and the improved YOLOv3-Tiny in this paper, were selected for comprehensive evaluation from three key indicators: mAP, floating-point operation amount, and FPS; Table 2 shows the verification results of these four models on Raspberry Pi-4B. Since all models use the same arrester infrared image dataset, the data loading time is not compared separately.
[0120] It can be seen from the experimental results that although the original YOLOv3-Tiny model has a relatively high inference speed of 135.1 FPS and a low floating-point operation volume, its mAP of 87.7% cannot meet the requirements for accurate recognition and positioning of arrester infrared images. Although the RT-DETR and YOLOv11 models achieve relatively high recognition accuracies, their huge floating-point operation volumes result in low FPS, making them not suitable for deployment on edge devices with limited computing resources. The improved YOLOv3-Tiny model, while maintaining low computing resource occupancy and fast recognition speed, still has room for improvement in recognition accuracy.
[0121] Table 2 Comparison of the performance of four models
[0122]
[0123] Taking RT-DETR and YOLOv11 as teacher models and YOLOv3-Tiny and the improved YOLOv3-Tiny as student models respectively, two knowledge distillation networks were constructed and their performances were evaluated on the same arrester infrared image dataset. The results are shown in Table 3. The experimental results show that through knowledge distillation, while maintaining the original computing resource occupancy and recognition speed, the mAP of the two lightweight models has been significantly improved. At the same time, compared with directly using large models, the knowledge distillation method saves more than 80% of the computing resources. This fully demonstrates the effectiveness of knowledge distillation in model compression, which can not only ensure the improvement of model recognition accuracy but also achieve the optimization of computing resources and the improvement of inference speed.
[0124] Table 3 Comparison of the performance of two knowledge distillation models
[0125]
[0126] 3. Performance comparison of different knowledge distillation networks:
[0127] To verify the superiority of the proposed knowledge distillation model in the recognition of arrester infrared images, four knowledge distillation models were constructed using the above four object detection models, namely RT-DETR+YOLOv3-Tiny, RT-DETR+YOLOv3-Tiny, YOLOv11+YOLOv3-Tiny, and YOLOv11+the improved YOLOv3-Tiny. All four models were trained for 50 epochs.
[0128] The experimental results are as Figure 4As shown, the knowledge distillation model proposed by the present invention based on YOLOv11 and improved YOLOv3-Tiny shows excellent training effects. The model quickly reaches an mAP of over 91% within the first 10 training cycles, then continues to improve and finally stabilizes at a high-precision level of over 98%. Compared with the other three comparison models, this model not only shows a faster convergence speed but also achieves a higher recognition accuracy. These experimental results fully verify the excellent performance of this knowledge distillation model in the task of identifying arrester infrared images.
[0129] Select several representative arrester infrared images in the dataset. Figure 5 The detection effect of the knowledge distillation model based on YOLOv11 and improved YOLOv3-Tiny proposed in this patent is shown. It can be seen from the figure that in a scene with a relatively simple background environment, the proposed algorithm can accurately identify the position of the arrester. In a scene with multiple arresters, the proposed model can also accurately identify the positions of all arresters. Combining the experimental data and the specific recognition effect diagrams can verify that the knowledge distillation model proposed in this paper can effectively and accurately identify arrester infrared images.
[0130] Verified by experiments, this knowledge distillation model has achieved remarkable results in the task of identifying arrester infrared images collected by drones: the average precision (mAP) reaches 98.7%, and at the same time, the computing resources required by the model are reduced by 80% compared with before compression.
[0131] Embodiment 2
[0132] This embodiment provides an arrester identification device based on improved YOLOv3-Tiny and knowledge distillation, including:
[0133] The YOLOv3-Tiny improvement module is used to improve YOLOv3-Tiny. The improvement includes improving YOLOv3-Tiny based on the Fuse-MBConv module and improving the training strategy of YOLOv3-Tiny. The Fuse-MBConv module includes a fused convolution operation and an SE module. By fusing the depthwise separable convolution and the expansion convolution into a conventional convolution operation, the amount of computation and the number of parameters are reduced. When improving the training strategy of YOLOv3-Tiny, a progressive learning strategy is used to dynamically adjust the image size and regularization strength to optimize the training efficiency while improving the model performance.
[0134] An arrester infrared image recognition knowledge distillation model acquisition module, which is used to establish an arrester infrared image recognition knowledge distillation model with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain an arrester infrared image recognition knowledge distillation model based on YOLOv11 and the improved YOLOv3-Tiny;
[0135] An arrester infrared image recognition module, which is used to recognize the collected arrester infrared images by using the arrester infrared image recognition knowledge distillation model based on YOLOv11 and the improved YOLOv3-Tiny.
[0136] A non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the arrester recognition method based on the improved YOLOv3-Tiny and knowledge distillation as described above is realized.
[0137] Furthermore, the present invention adopts the following technical solutions:
[0138] An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the arrester recognition method based on the improved YOLOv3-Tiny and knowledge distillation as described above is realized.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that the facilities of the present invention can be realized by means of software plus a necessary general hardware platform. The embodiments of the present invention can be realized by using existing processors, or by dedicated processors used for this purpose or other purposes in a suitable system, or by a hardwired system. The embodiments of the present invention also include a non-transitory computer-readable storage medium, which includes a machine-readable medium for carrying or having machine-executable instructions or data structures stored thereon; such a machine-readable medium can be any available medium accessible by a general or special computer or other machine with a processor. For example, such a machine-readable medium can include RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disc memories, magnetic disk memories or other magnetic storage devices, or any other medium that can be used to carry or store the required program code in the form of machine-executable instructions or data structures and can be accessed by a general or special computer or other machine with a processor. When information is transmitted or provided to a machine through a network or other communication connection (hardwired, wireless, or a combination of hardwired and wireless), this connection is also regarded as a machine-readable medium.
[0140] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation, characterized in that: The method comprises the following steps: Improve YOLOv3-Tiny, including improving YOLOv3-Tiny based on Fuse-MBConv module and improving YOLOv3-Tiny training strategy; the Fuse-MBConv module includes a fused convolution operation and an SE module, which reduces the amount of calculation and the amount of parameters by fusing the depthwise separable convolution and the extended convolution into a conventional convolution operation; when improving the YOLOv3-Tiny training strategy, a progressive learning strategy is used to dynamically adjust the image size and regularization strength to optimize the training efficiency while improving the model performance; Establish a knowledge distillation model for infrared image recognition of lightning arresters with YOLOv11 as the teacher model and improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain a knowledge distillation model for infrared image recognition of lightning arresters based on YOLOv11 and improved YOLOv3-Tiny. The captured infrared images of lightning arresters are recognized using the knowledge distillation model for lightning arrester infrared image recognition based on YOLOv11 and improved YOLOv3-Tiny.
2. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 1 is characterized in that: When improving YOLOv3-Tiny based on the Fuse-MBConv module, in terms of network structure, the Fuse-MBConv module is introduced after the feature extraction layer, and its efficient feature fusion mechanism is used to enhance the feature extraction capability of the network; at the same time, the SE module is integrated at different scales of the feature pyramid network, and the feature expression is optimized through an adaptive channel attention mechanism.
3. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 2 is characterized in that: The method for improving YOLOv3-Tiny based on the Fuse-MBConv module is: Assume the input feature map is Where c is the number of channels, h and w are the height and width of the feature map; The fused convolution operation processes the input features: First, a 3×3 convolution kernel is used to extract spatial features from the input features; then, a batch normalization layer is used to stabilize the feature distribution and accelerate the training process; then the SiLU activation function is used to introduce nonlinear transformations; finally, a 1×1 convolution is used to adjust the feature channel dimension to achieve cross-channel information reorganization; its mathematical expression is: F fused (X)=Conv 1×1 (SiLU(BN(Conv 3×3 (X)))) (1) In the formula, Conv is the convolution operation; BN is batch normalization; SiLU is the activation function; The SE module performs adaptive feature recalibration by explicitly modeling the inter-channel dependencies: first, the spatial dimension features are compressed into channel descriptors through global average pooling to capture global context information; Then, two fully connected layers and nonlinear activation functions are used to build nonlinear relationships between channels and generate importance weights for each feature channel. Finally, these weights are multiplied with the original features to achieve adaptive feature recalibration at the channel level. Its mathematical expression is: In the formula, z c is the statistical value of the cth channel; u c Represents the feature map of the c-th channel. s=F ex (z,W)=σ(W2δ(W1z)) (3) Where s is the channel weight; is the weight of the first fully connected layer; is the weight of the second fully connected layer; δ represents the ReLU activation function; r is the dimensionality reduction ratio; In the formula, s c is the weight of the cth channel, is the recalibrated feature map of the cth channel.
4. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 1 is characterized in that: The method for improving the YOLOv3-Tiny training strategy is: During the training process, the image size increases gradually: a smaller image size is used at the beginning of the training, and as the training progresses, the image size is gradually increased until the target size is reached. The image size adjustment function is formula (5): Where, d t is the image size at step t; d min is the minimum image size; d max is the maximum image size; T d is the total number of steps in which the size increases; t is the current training step number; During the training process, the regularization strength is gradually increased: a weaker regularization is used at the beginning of the training, and as the image size increases, the regularization strength is increased synchronously; the regularization strength adjustment function is formula (6): In the formula, λ t is the regularization strength at step t; min is the minimum regularization strength; λ max is the maximum regularization strength; T λ is the total number of regularization steps.
5. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 1 is characterized in that: Before establishing a knowledge distillation model for arrester infrared image recognition with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, it is necessary to design the loss function of the knowledge distillation model. The designed loss functions include classification loss function, regression loss function and feature map distillation loss function, and the total loss function of knowledge distillation is obtained by weighted sum.
6. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 5 is characterized in that: The classification loss function is: the student model needs to learn the true label while imitating the classification probability distribution of the teacher model: Where, L CE is the cross entropy loss of the student model for the true label; T is the distillation temperature, which is used to smooth the probability distribution; z t and z s are the classification output logits of the teacher model and the student model respectively; α is the weight of classification loss and distillation loss; KL is the Kullback-Leibler divergence calculation; is to t After temperature scaling, the probability distribution obtained by the sigmoid function (σ); is to s After temperature scaling, the probability distribution obtained by the sigmoid function (σ); The regression loss function is: the bounding box regression of the student model should be close to the bounding box predicted by the teacher model In the formula, and are the predictions of the i-th bounding box by the student and teacher models respectively; N is the number of bounding boxes; SmoothL1 is the smoothed L1 loss function used for the regression error. The feature map distillation loss function is: the intermediate layer features of the student network need to mimic the features of the corresponding layer in the teacher network In the formula, and F t i are the feature maps of the i-th position of the student and teacher models respectively; The above three loss functions are weighted and the total loss function is obtained: L total =λ1·L cls +λ2·L reg +λ3·L feat (10) Where λ1, λ2 and λ3 represent the weights of the classification loss function, regression loss function and feature distillation loss function, respectively.
7. The lightning arrester identification method based on improved YOLOv3-Tiny and knowledge distillation according to claim 1 is characterized in that: A knowledge distillation model for infrared image recognition of lightning arresters is established with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model. The knowledge distillation model is trained to obtain the knowledge distillation model for infrared image recognition of lightning arresters based on YOLOv11 and improved YOLOv3-Tiny. The method is as follows: ① Accurately mark the arrester positions in the collected infrared image data set, and input the marked images into the teacher model and student model respectively to carry out model training and target recognition; ② First, pre-train the teacher model and use the original image data with labeled positions to train the YOLOv11 model until convergence; then fix the teacher model parameters and generate soft labels including classification probability distribution, bounding box regression values and intermediate feature maps; ③ After initializing the student model, training begins. The improved YOLOv3-Tiny model receives dual supervision from both soft labels and original hard labels during training; ④ During the training process, the model convergence is evaluated by monitoring the loss function of the knowledge distillation model: if the convergence conditions are not met, the network parameters are adjusted to continue training; once the model converges, the training is terminated and the parameters of the current knowledge distillation model are saved.
8. A lightning arrester identification device based on improved YOLOv3-Tiny and knowledge distillation, characterized in that: include: A YOLOv3-Tiny improvement module, used for improving YOLOv3-Tiny, wherein the improvement includes improving YOLOv3-Tiny based on a Fuse-MBConv module and improving a YOLOv3-Tiny training strategy; the Fuse-MBConv module includes a fused convolution operation and an SE module, which reduces the amount of calculation and the amount of parameters by fusing a depthwise separable convolution and an extended convolution into a conventional convolution operation; when improving the YOLOv3-Tiny training strategy, a progressive learning strategy is used to dynamically adjust the image size and regularization strength, thereby optimizing the training efficiency while improving the model performance; The module for acquiring the knowledge distillation model for infrared image recognition of lightning arresters is used to establish a knowledge distillation model for infrared image recognition of lightning arresters with YOLOv11 as the teacher model and the improved YOLOv3-Tiny as the student model, train the knowledge distillation model, and obtain the knowledge distillation model for infrared image recognition of lightning arresters based on YOLOv11 and the improved YOLOv3-Tiny. The arrester infrared image recognition module is used to recognize the collected arrester infrared image by using the arrester infrared image recognition knowledge distillation model based on YOLOv11 and improved YOLOv3-Tiny.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the arrester identification method based on improved YOLOv3-Tiny and knowledge distillation as described in any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the arrester identification method based on improved YOLOv3-Tiny and knowledge distillation is implemented as described in any one of claims 1 to 7.