A Transmission Line Defect Detection Method and System Based on Knowledge Distillation

The knowledge distillation method enhances defect detection in power transmission lines by optimizing student models with shallow and deep feature distillation, addressing computational and resource constraints while maintaining high accuracy.

CN119941714BActive Publication Date: 2025-07-15WENZHOU ELECTRIC POWER BUREAU +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510412959.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-15
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing transmission line defect detection methods have shortcomings in terms of computational complexity, model volume and training time, especially in resource-constrained scenarios, and deep learning-based methods have limitations in adapting to complex scenarios and target changes.

Method used

The teacher-student training method based on knowledge distillation is adopted, and a multi-level feature distillation and knowledge transfer mechanism is introduced between the teacher model and the student model through shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation, and the training process of the student model is optimized, including regional mask technology, attention mechanism and error measurement method, to achieve multi-level transfer of knowledge.

Benefits of technology

While reducing the computational complexity, the student model retains the high detection accuracy of the teacher model to the greatest extent, and can be efficiently deployed on terminals or edge devices, achieving large-scale, real-time, and lightweight defect detection, and improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941714B_ABST
    Figure CN119941714B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of transmission line defects, and discloses a transmission line defect detection method and system based on knowledge distillation, including obtaining a to-be-detected image of a transmission line; inputting the to-be-detected image into a pre-constructed defect detection model to obtain a defect detection result of the transmission line, the defect detection model is constructed based on a convolutional neural network and is trained by a teacher-student training method based on knowledge distillation, and the knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model, realizes multi-level knowledge transfer from the teacher model to the student model, effectively avoids the neglect of key details caused by the simplification of the student model, has high detection accuracy while reducing the computational complexity of the student model, and thus significantly improves the overall efficiency of transmission line defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of transmission line defect detection, and particularly to a transmission line defect detection method and system based on knowledge distillation. Background Art

[0002] In the power system, the transmission line is an important channel for electric energy transmission, and it is prone to various defects due to natural and human factors. In recent years, with the development of digital image processing and analysis technologies in scenarios such as security monitoring and drone cruising, the object detection technology based on digital image processing has gradually been applied to the defect detection of transmission lines in the power system.

[0003] Currently, the methods for object detection of transmission line defects are mainly divided into two categories: methods based on handcrafted features and methods based on deep learning. Among them, the methods based on handcrafted features mainly include the sliding window method, the template matching method, and the methods using features such as SIFT and HOG for classification; the methods based on deep learning mainly include automatically extracting image features through a neural network model and combining classification and regression tasks to achieve the detection of transmission line defects. However, both of these have certain limitations. On the one hand, the methods based on handcrafted features have achieved certain results in early object detection, but due to relying on manually designed features, they have poor adaptability to complex scenarios and target changes and are difficult to meet the actual application requirements; on the other hand, although the methods based on deep learning can significantly improve the performance of transmission line defect object detection, there are still deficiencies in terms of computational complexity, model size, and training time, especially in resource-constrained scenarios, it is difficult to directly deploy. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a transmission line defect detection method and system based on knowledge distillation, so as to solve the deficiencies of the existing object detection models in terms of computational complexity, model size, and training time, and enable the model to have high-precision detection capabilities while maintaining lightweight.

[0005] In a first aspect, the present invention provides a transmission line defect detection method based on knowledge distillation, and the method includes:

[0006] Obtain a to-be-detected image of a transmission line, where the to-be-detected image includes image data and video frame data;

[0007] Input the to-be-detected image into a pre-constructed defect detection model to obtain a defect detection result of the transmission line, where the defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation;

[0008] Among them, the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of transmission lines. The network architectures of the teacher model and the student model are the same, but the network depths are different;

[0009] The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation;

[0010] When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss according to the shallow feature differences in different regions;

[0011] When using deep feature knowledge distillation for model training, based on the attention mechanism, associate the candidate boxes of the teacher model with the deep features of the student model to generate an attention mask, and calculate the deep feature distillation loss;

[0012] When using response knowledge distillation for model training, use the detection results of the teacher model as soft labels, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss.

[0013] Furthermore, when using at least two knowledge distillation methods among shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation for model training, perform a weighted sum of the distillation losses corresponding to the used knowledge distillation methods to construct a multi-task optimization framework and optimize the training of the student model.

[0014] Furthermore, the steps of calculating the shallow feature differences between the teacher model and the student model in different regions according to the region masking technique and the error metric method, and calculating the shallow feature knowledge distillation loss according to the shallow feature differences in different regions include:

[0015] Input the image data in the dataset into the teacher model and the student model respectively for shallow feature extraction to obtain the shallow feature map of the teacher model and the shallow feature map of the student model;

[0016] According to the set of annotation boxes corresponding to the input image data, divide the image data into regions, and generate a region mask according to the divided regions;

[0017] According to the region mask and the error metric method, calculate the shallow feature differences between the shallow feature map of the teacher model and the shallow feature map of the student model in different regions;

[0018] Perform a weighted sum of the shallow feature differences in different regions to obtain the shallow feature knowledge distillation loss.

[0019] Further, the step of dividing the image data into regions according to the set of annotation boxes corresponding to the input image data and generating a region mask according to the divided regions includes:

[0020] Regarding the regions covered by each annotation box in the set of annotation boxes corresponding to the image data as foreground regions, and regarding the regions outside each annotation box as background regions;

[0021] Marking the mask value corresponding to each pixel point in the foreground region as 1, and marking the mask value corresponding to each pixel point in the background region as 0, to obtain a region mask.

[0022] Further, the step of calculating the shallow feature differences of the shallow feature maps of the teacher model and the student model in different regions according to the region mask and the error metric method includes:

[0023] Calculating the error loss of the shallow feature maps of the teacher model and the student model in the foreground region according to the region mask and the mean square error;

[0024] Calculating the error loss of the shallow feature maps of the teacher model and the student model in the background region according to the inverted value of the region mask and the mean square error.

[0025] Further, the shallow feature knowledge distillation loss is represented by the following formula:

[0026]

[0027] In the formula, represents the region mask at the coordinate point (h, w), represents the eigenvalue at the coordinate point (h, w) of the c-th channel of the shallow feature map of the -th layer of the student model, represents the eigenvalue at the coordinate point (h, w) of the c-th channel of the shallow feature map of the -th layer of the teacher model, represents the total number of pixels in the foreground region, represents the foreground region weight loss coefficient, represents the total number of pixels in the background region, represents the background region weight loss coefficient.

[0028] Further, the step of associating the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism, generating an attention mask, and calculating the deep feature distillation loss includes:

[0029] In the teacher model and the student model, feature extraction is respectively performed on their respective shallow feature maps to obtain the deep feature map of the teacher model and the deep feature map of the student model;

[0030] The teacher model performs object bounding on the deep feature map of the teacher model to obtain a set of candidate bounding boxes of the teacher model;

[0031] According to the set of candidate bounding boxes, a candidate bounding box guidance region is generated, and the candidate bounding box guidance region is mapped to the spatial dimension of the deep feature map of the student model through a linear transformation to generate an attention mask;

[0032] According to the attention mask, the deep feature map of the teacher model and the deep feature map of the student model are aligned to calculate the deep feature distillation loss.

[0033] Further, the step of the teacher model performing object bounding on the deep feature map of the teacher model to obtain a set of candidate bounding boxes of the teacher model includes:

[0034] The teacher model performs object bounding on the deep feature map of the teacher model to obtain multiple candidate bounding boxes, and forms a set of candidate bounding boxes;

[0035] Calculate the intersection over union (IoU) between each candidate bounding box in the set of candidate bounding boxes and each annotation bounding box in the set of annotation bounding boxes corresponding to the input image data;

[0036] According to the comparison relationship between the IoU and the IoU threshold, multiple high-quality candidate bounding boxes are filtered out from the set of candidate bounding boxes to form a set of high-quality candidate bounding boxes, and the set of candidate bounding boxes is updated according to the set of high-quality candidate bounding boxes.

[0037] Further, the step of generating a candidate bounding box guidance region according to the set of candidate bounding boxes, and mapping the candidate bounding box guidance region to the spatial dimension of the deep feature map of the student model through a linear transformation to generate an attention mask includes:

[0038] Merge each candidate bounding box in the set of candidate bounding boxes to generate a candidate bounding box guidance region;

[0039] The candidate bounding box guidance region is mapped to a query matrix through a weight matrix, and the deep feature map of the student model is mapped to a key matrix and a value matrix;

[0040] According to the attention mechanism, the query matrix and the key matrix are dot-producted and normalized to obtain an attention mask.

[0041] Further, the step of aligning the deep feature map of the teacher model and the deep feature map of the student model according to the attention mask to calculate the deep feature distillation loss includes:

[0042] Calculate the deep feature error loss between the deep feature maps of the teacher model and the student model based on the mean square error, and calculate the inner product between the attention mask and the deep feature error loss;

[0043] Calculate the deep feature distillation loss according to the number of deep feature layers of the student model and the inner product.

[0044] Furthermore, it is characterized in that the deep feature knowledge distillation loss is expressed by the following formula:

[0045]

[0046] In the formula, represents the deep feature error loss, represents the attention mask of the r-th layer of deep features, represents the inner product, and T represents the number of deep feature layers.

[0047] Furthermore, before the step of performing object bounding on the deep feature map of the teacher model by the teacher model to obtain the candidate box set of the teacher model, it further includes:

[0048] Perform an alignment operation on the deep feature map of the student model and the deep feature map of the teacher model, and the alignment operation includes size alignment and resolution alignment.

[0049] Furthermore, the step of using the detection result of the teacher model as a soft label and calculating the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on an error metric method to obtain the response knowledge distillation loss includes:

[0050] Perform classification and regression predictions on their respective candidate boxes through the teacher model and the student model to obtain the detection result of the teacher model and the detection result of the student model. The detection result of the teacher model includes the classification probability distribution of the teacher model and the regression result of the teacher model, and the detection result of the student model includes the classification probability distribution of the student model and the regression result of the student model;

[0051] Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on the KL divergence;

[0052] Calculate the regression response knowledge loss between the regression result of the teacher model and the regression result of the student model based on the mean square error;

[0053] Perform weighted summation on the classification response knowledge loss and the regression response knowledge loss to obtain the response knowledge distillation loss.

[0054] Further, the step of performing classification and regression prediction on the respective candidate boxes through the teacher model and the student model to obtain the detection results of the teacher model and the detection results of the student model includes:

[0055] Taking the candidate boxes of the teacher model as the candidate boxes of the student model, and performing classification and regression prediction on the respective candidate boxes through the teacher model and the student model to obtain the detection results of the teacher model and the detection results of the student model.

[0056] Further, the response knowledge distillation loss is expressed by the following formula:

[0057]

[0058] In the formula, L kl represents the KL divergence, represents the i-th candidate box, N represents the total number of candidate boxes, represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, represents the regression result of the i-th candidate box in the teacher model, represents the regression result of the i-th candidate box in the student model, represents the weight coefficient of the classification response knowledge loss, represents the weight coefficient of the regression response knowledge loss.

[0059] Further, both the teacher model and the student model are constructed using a fast convolutional neural network, and the teacher model uses ResNet-101 as the backbone network, while the student model uses ResNet-50 as the backbone network.

[0060] Further, the network architectures of both the teacher model and the student model include a shallow feature extraction network, a deep feature extraction network, a region proposal network, and a detection head;

[0061] Among them, the shallow feature extraction network is the first three stages of the backbone network, and is used for extracting shallow features from the input image data;

[0062] The deep feature extraction network is the last two stages of the backbone network, and is used for extracting deep features from the shallow features output by the shallow feature extraction network;

[0063] The region proposal network is used for selecting target boxes from the deep features output by the deep feature extraction network to generate candidate boxes;

[0064] The detection head is used for performing classification and regression prediction on the candidate boxes output by the region proposal network and outputting detection results.

[0065] Further, the SGD optimizer is used to solve the multi-task optimization framework to obtain the optimal parameters of the student model.

[0066] In a second aspect, the present invention provides a transmission line defect detection system based on knowledge distillation. The system includes:

[0067] An image acquisition module, configured to acquire an image to be detected of a transmission line, where the image to be detected includes image data and video frame data;

[0068] A defect detection module, configured to input the image to be detected into a pre-constructed defect detection model to obtain a defect detection result of the transmission line. The defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation;

[0069] Among them, the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of transmission lines. The network architectures of the teacher model and the student model are the same, but the network depths are different;

[0070] The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation;

[0071] When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss based on the shallow feature differences in different regions;

[0072] When using deep feature knowledge distillation for model training, associate the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism to generate an attention mask, and calculate the deep feature distillation loss;

[0073] When using response knowledge distillation for model training, use the detection result of the teacher model as a soft label, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss.

[0074] The present invention provides a method and system for detecting transmission line defects based on knowledge distillation. By designing shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation, the present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model, realizing multi-level knowledge transfer from the teacher model to the student model, effectively avoiding the neglect of key details caused by the simplification of the student model, while reducing the computational complexity of the student model, enabling the student model to retain the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on terminals or edge devices, and still have excellent detection effects even in resource-constrained scenarios. Through the defect detection method provided by the present invention, the overall efficiency of transmission line defect detection is significantly improved, providing an effective solution for large-scale, real-time, and lightweight defect detection applications, and having good practicability and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is a schematic flowchart of the method for detecting transmission line defects based on knowledge distillation in an embodiment of the present invention;

[0076] Figure 2 is another schematic flowchart of the method for detecting transmission line defects based on knowledge distillation in an embodiment of the present invention;

[0077] Figure 3 is a schematic structural diagram of the system for detecting transmission line defects based on knowledge distillation in an embodiment of the present invention;

[0078] Figure 4 is an internal structural diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0080] Please refer to Figure 1 , a method for detecting transmission line defects based on knowledge distillation proposed in the first embodiment of the present invention, which includes steps S10 to S20:

[0081] Step S10, obtain a to-be-detected image of a transmission line, where the to-be-detected image includes image data and video frame data;

[0082] Step S20: Input the image to be detected into a pre-constructed defect detection model to obtain the defect detection result of the transmission line. The defect detection model is constructed based on a convolutional neural network and trained using a teacher-student training method based on knowledge distillation.

[0083] In the present invention, the defect detection model is used to perform object detection on the image of the transmission line, so as to determine whether there are defects in the transmission line in the image. Among them, the data input into the defect detection model is a set of image data or video frame data containing the object. Before inputting into the model, the image to be detected also needs to be preprocessed, including uniformly adjusting the size of the image, normalizing the pixel values in the image, etc. The defect detection model is constructed using a convolutional neural network model, and a teacher-student training method based on knowledge distillation is adopted during the model training. Through the trained defect detection model, the predicted object type and bounding box information can be obtained.

[0084] In model training, the teacher-student training method is a technology for model compression and acceleration based on knowledge distillation. Its core idea is to transfer the knowledge of a trained complex model (teacher model) to a smaller and more easily deployable model (student model). The teacher model usually has a high accuracy rate but a large computational cost; the student model is relatively simple and has high computational efficiency, but its performance may not be as good as that of the teacher model when directly trained. Through knowledge distillation, the student model can learn the "knowledge" of the teacher model, so as to achieve a performance close to that of the teacher model while maintaining a small model size. That is to say, the teacher-student training method guides the learning of the student model through the feature knowledge, response knowledge, and global knowledge of the teacher model, so as to significantly reduce the number of parameters and computational overhead of the student model while maintaining the detection accuracy of the model and improving the performance of the student model.

[0085] Currently, the conventional teacher-student training method only considers deep feature distillation during the knowledge distillation process. In fact, in the teacher model, its shallow features, candidate boxes, etc. also contain important information, but this information has not been effectively introduced into the knowledge distillation process, resulting in the performance of the student model not meeting the high-precision requirements. To solve this problem, the present invention provides a teacher-student training method based on knowledge distillation to train the student model. Among them, knowledge distillation includes various knowledge distillation methods such as shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation. When training the model, any one or more of the various knowledge distillation methods can be selected according to the actual situation such as the requirements of the model for knowledge distillation. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model to avoid ignoring key details due to model simplification, so that the student model can retain the high detection accuracy of the teacher model to the greatest extent.

[0086] The three knowledge distillation methods provided by the present invention are respectively applied to different stages of model training. Among them, shallow feature knowledge distillation is applied to the shallow feature extraction stage of the model. According to the region masking technique and the error metric method, the shallow feature differences between the teacher model and the student model in different regions are calculated, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions.

[0087] Deep feature knowledge distillation is applied to the deep feature extraction stage of the model. Based on the attention mechanism, the candidate boxes of the teacher model are associated with the deep features of the student model to generate an attention mask, and the deep feature distillation loss is calculated.

[0088] Response knowledge distillation is applied to the candidate box classification and regression prediction stage of the model. The detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error metric method to obtain the response knowledge distillation loss.

[0089] The reason why the present invention provides three knowledge distillation methods is that the existing knowledge distillation methods only consider deep feature distillation. In fact, the shallow features of the teacher model contain rich underlying visual information such as edges and textures, and this underlying visual information is not distilled to the student model in the existing methods; at the same time, the candidate boxes of the teacher model contain the model's attention information for the key feature regions, and this attention is not distilled to the student model in the existing methods; in addition, the existing methods often lack attention to the quality of the candidate boxes in the object detection results and cannot effectively solve the problem of poor quality of the candidate boxes extracted by the student model.

[0090] To solve the above problems, the present invention provides three knowledge distillation methods applied to different stages. These three knowledge distillation methods can be applied alone, jointly, or in combination with the existing knowledge distillation methods. For the convenience of description, the model training method provided by the present invention is described below through the joint application of the three knowledge distillation methods.

[0091] Please refer to Figure 2 , the knowledge distillation in the present invention can be divided into three stages: shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation. For the convenience of description, the following settings are made for the teacher model, the student model, and the dataset for model training:

[0092] The teacher model is a pre-trained model for defect detection of transmission lines, and the student model is a defect detection model. Both the teacher model and the student model are constructed using the Faster Region-based Convolutional Neural Network (Faster-RCNN) model, and the Residual Network (ResNet) is used as the backbone network for feature extraction. Among them, the backbone network of the teacher model uses the relatively complex ResNet-101, while the backbone network of the student model uses the lightweight ResNet-50.

[0093] The training dataset for the object detection task includes an image data set, as well as the corresponding class label set and bounding box set. Among them, the bounding box contains the upper left and lower right coordinates of each object. At the same time, the size of each image data is uniformly adjusted, and its pixel values are normalized to unify the input format. It should be noted that the teacher model and the student model in the present invention can also be constructed using other neural network models with object detection functions. Only a preferred construction method is given here, rather than a specific limitation.

[0094] Based on the above settings, when training the defect detection model using shallow feature knowledge distillation, the specific knowledge distillation steps are as follows:

[0095] Input the image data in the dataset into the teacher model and the student model respectively for shallow feature extraction to obtain the shallow feature maps of the teacher model and the student model.

[0096] According to the bounding box set corresponding to the input image data, divide the image data into regions, and generate region masks according to the divided regions.

[0097] According to the region masks and the error metric method, calculate the shallow feature differences between the shallow feature maps of the teacher model and the student model in different regions respectively.

[0098] Perform weighted summation on the shallow feature differences in different regions to obtain the shallow feature knowledge distillation loss.

[0099] In this embodiment, the shallow feature knowledge distillation is to perform knowledge distillation on the shallow features of the teacher model and the student model, and the goal is to align the shallow features (such as edges and textures) of the student model with those of the teacher model. Therefore, it is necessary to first obtain the shallow features of the teacher model and the student model: input the image data in the dataset into the teacher model and the student model respectively for shallow feature extraction, so as to obtain the shallow feature maps of the teacher model and the student model.

[0100] Based on the above model settings, the network architectures of the teacher model and the student model both include four parts, namely the shallow feature extraction network, the deep feature extraction network, the region proposal network, and the detection head. Since the backbone networks of the teacher model and the student model adopt the residual network, and the network structure of the residual network can be divided into five stages. Among them, the first stage is the input layer, and the input layer includes a convolutional layer and a pooling layer. The second stage is the first residual block group, the third stage is the second residual block group, the fourth stage is the third residual block group, and the fifth stage is the fourth residual block group. Each residual block group contains multiple residual blocks. Whether it is the ResNet-101 model or the ResNet-50 model, it can be divided into these five stages. The difference lies in the network depth, and the number of residual blocks contained in the residual block group will be different.

[0101] Based on the five stages of the residual network, in this embodiment, the first three stages are used as the shallow feature extraction network, and the last two stages are used as the deep feature extraction network. In addition, since both the teacher model and the student model are constructed using the Faster-RCNN model, after the backbone network of the model, there is also a region proposal network PRN and a detection head. Of course, if the teacher model and the student model are constructed using other neural network models, they can be divided into stages according to the above four functions of shallow feature extraction, deep feature extraction, candidate box selection, and result detection, and will not be elaborated here one by one.

[0102] Based on the above architecture, in the teacher model, for the input image data, the shallow feature extraction network is used to capture the underlying visual information in the image, such as texture, edges, etc., so as to generate a shallow feature map. :

[0103]

[0104] where I q represents the input image data, and Z1, Z2, and Z3 respectively represent the first stage, the second stage, and the third stage of the backbone network of the teacher model. represents a real number matrix with dimensions of H (height) × W (width) × C (number of channels), which is used to describe the shape and data type of the feature map.

[0105] In the student model, similarly, the image data is input into the shallow feature extraction network of the student model to obtain the shallow feature map of the student model. :

[0106]

[0107] where Z1 s 、Z2 s and Z3 srespectively represent the first stage, the second stage, and the third stage of the backbone network of the student model.

[0108] Before performing knowledge distillation on the shallow feature maps of the above-mentioned teacher model and the shallow feature maps of the student model, the present invention divides the image data based on the annotation boxes of the input images and generates corresponding region masks. The specific steps include:

[0109] Take the regions covered by each annotation box in the set of annotation boxes corresponding to the image data as the foreground regions, and take the regions outside each annotation box as the background regions;

[0110] Mark the mask values corresponding to each pixel point in the foreground region as 1, and mark the mask values corresponding to each pixel point in the background region as 0 to obtain a region mask.

[0111] In this embodiment, based on the set of annotation boxes in the dataset, the foreground region and the background region are divided. Among them, the foreground region refers to the region containing the target. Therefore, the parts covered by each annotation box in the set of annotation boxes are taken as the foreground regions, and the parts outside the annotation boxes are taken as the background regions. Then, based on the divided regions, a binary region mask is generated. Among them, the mask values corresponding to each pixel point in the foreground region are marked as 1, and the mask values corresponding to each pixel point in the background region are marked as 0, so as to generate a region mask M for the input image data:

[0112]

[0113] In the formula, represents the coordinate value of the pixel point, i represents the row coordinate, and j represents the column coordinate, , and G represents the set of annotation boxes.

[0114] Then, based on the divided regions, an error metric method is used to calculate the differences between the teacher model and the student model in the shallow features within each region respectively. The specific steps include:

[0115] Calculate the error loss between the shallow feature map of the teacher model and the shallow feature map of the student model in the foreground region according to the region mask and the mean square error;

[0116] Calculate the error loss between the shallow feature map of the teacher model and the shallow feature map of the student model in the background region according to the inverted value of the region mask and the mean square error.

[0117] In this embodiment, by calculating the mean square error between the shallow feature map of the teacher model and the shallow feature map of the student model, the shallow feature difference is characterized. At the same time, according to the region mask, the shallow feature difference is represented by region. Among them, the mean square error loss in the foreground region can be expressed as:

[0118]

[0119] The mean square error in the background region is calculated by inverting the region mask:

[0120]

[0121] In the formula, represents the region mask at the coordinate point (h, w), represents the eigenvalue at the coordinate point (h, w) in channel c of the shallow feature map of the student model in the -th layer, represents the eigenvalue at the coordinate point (h, w) in channel c of the shallow feature map of the teacher model in the -th layer, represents the total number of pixels in the foreground region, represents the total number of pixels in the background region.

[0122] Finally, the weighted sum of the shallow feature differences in the background region and the shallow feature differences in the foreground region is calculated to obtain the shallow feature knowledge distillation loss:

[0123] +

[0124] Where, represents the shallow feature knowledge distillation loss in the -th layer, represents the foreground region weight loss coefficient, represents the background region weight loss coefficient.

[0125] It should be noted that using the mean square error to characterize the difference between features in the present invention is only a preferred form of difference characterization. Other error measurement methods such as the mean absolute error, Huber loss, squared absolute error, logarithmic cosine loss, etc. can also be used to characterize the feature difference. There is no excessive limitation here. That is to say, and in the above formula can be calculated by various error measurement methods.

[0126] In this embodiment, the importance of different regions is adjusted by the loss weight coefficients of the foreground region and the background region, where the weight loss coefficient of the foreground region is greater than that of the background region. It can be seen that the shallow knowledge distillation loss in this embodiment includes two parts. In the multi-task loss of knowledge distillation, the influence of different parts may be unbalanced. Therefore, in a preferred embodiment, the above shallow feature knowledge distillation loss is normalized to avoid a certain region dominating the entire loss due to an excessive number of elements, ensuring that the contribution of the loss term of each region to the total loss is appropriate. Taking the feature difference characterization based on the mean square error as an example, the formula for the normalized distillation loss can be expressed as:

[0127]

[0128] The denominator is added to the weight loss coefficient in the above formula, so that when calculating the mean square error, each region is normalized separately, making the loss values of the foreground and background regions comparable, avoiding a certain term dominating the total loss due to the difference in region size. At the same time, in backpropagation, the gradient of the loss value is proportional to the derivative of the loss function with respect to the parameters. The gradient of a large region (such as the number of background pixels being much more than that of the foreground) may dominate the optimization direction, causing the model to over-focus on the background and ignore the target region. Therefore, through the above formula, the gradient amplitude of each region can be reduced to make the optimization process more stable. Of course, other weight analysis methods can also be used to determine the weight loss coefficient, and the specific weight loss coefficient can be flexibly set according to the actual situation, which will not be elaborated here one by one.

[0129] In this embodiment, the region mask technology is adopted to distinguish the target region (foreground region) and the non-target region (background region) in the image data, and the shallow feature differences between the student model and the teacher model in different regions are calculated respectively, so as to optimize the shallow feature representation of the student model. Through the shallow feature knowledge distillation of the present invention, the student model can be effectively guided to learn the detailed information in the shallow features, solve the problem of detail loss caused by the lightweight of the student model, and thus improve the detail perception ability of the student model.

[0130] Furthermore, aiming at the problem that the existing methods do not make full use of the candidate box information output by the teacher model, the present invention provides a deep feature knowledge distillation for deep feature alignment between the teacher model and the student model, providing the candidate box information and attention mechanism of the teacher model to strengthen the semantic feature expression ability of the student model for key regions. The specific steps include:

[0131] Feature extraction is respectively performed on the shallow feature maps of the teacher model and the student model to obtain the deep feature map of the teacher model and the deep feature map of the student model;

[0132] Perform object bounding on the deep feature maps of the teacher model to obtain a set of candidate boxes of the teacher model;

[0133] Generate a candidate box guidance region based on the set of candidate boxes, and map the candidate box guidance region to the spatial dimension of the deep feature maps of the student model through a linear transformation to generate an attention mask;

[0134] Align the deep feature maps of the teacher model and the deep feature maps of the student model according to the attention mask, and calculate the deep feature distillation loss.

[0135] In this embodiment, the purpose of deep feature knowledge distillation is to align the deep features of the teacher model and the student model, such as semantic information, and combine the attention mechanism to focus on key regions. During the deep feature knowledge distillation process, it is first necessary to obtain the deep features of the teacher model and the student model. Still taking the above settings as an example, in the teacher model, the shallow feature maps are input into the deep feature extraction network to capture the high-level semantic information of the target region, thereby obtaining the deep feature maps of the teacher model :

[0136]

[0137] In the formula, Z4 and Z5 respectively represent the fourth stage and the fifth stage of the backbone network.

[0138] Similarly, in the student model, the shallow feature maps are input into the deep feature extraction network of the student model to obtain the deep feature maps of the student model :

[0139]

[0140] In the formula, Z4 s and Z5 s respectively represent the fourth stage and the fifth stage of the backbone network of the student model.

[0141] After obtaining the deep feature maps of the teacher model and the student model, it is also necessary to perform an alignment operation on the deep feature maps of the student model to ensure that the deep feature maps of the teacher model and the transformed deep feature maps of the student model are consistent in size and resolution for subsequent processing.

[0142] After the alignment operation, the region candidate network of the teacher model performs object bounding on the deep feature map of the teacher model, so as to obtain a candidate box set composed of multiple candidate boxes output by the teacher model. This candidate box set will be used as a basis and combined with the attention mechanism to calculate the correlation between the candidate box region of the teacher model and the deep feature map of the student model, thereby optimizing the student model's learning of deep features.

[0143] In order to improve the quality of the candidate boxes output by the teacher model, in a preferred embodiment, the steps of performing object bounding on the deep feature map by the teacher model to obtain the candidate box set of the teacher model include:

[0144] Performing object bounding on the deep feature map of the teacher model by the teacher model to obtain multiple candidate boxes and forming a candidate box set;

[0145] Calculating the intersection over union (IoU) between each candidate box in the candidate box set and each annotation box in the annotation box set corresponding to the input image data;

[0146] According to the comparison relationship between the IoU and the IoU threshold, screening out multiple high-quality candidate boxes from the candidate box set to form a high-quality candidate box set, and updating the candidate box set according to the high-quality candidate box set.

[0147] In this embodiment, first, the deep feature map of the teacher model is sent into the region proposal network (RPN) for object bounding to select multiple candidate boxes, thereby generating a candidate box set B:

[0148]

[0149] where the candidate box is represented by the top-left coordinate and the bottom-right coordinate, that is, Figure 2 the coordinates (x i , y i ) and the coordinates (x' i , y' i ) in the candidate box information of

[0150] Then calculate the intersection over union (IoU) between each candidate box in the candidate box set and each annotation box in the annotation box set in the dataset. According to the comparison relationship between the IoU and the IoU threshold, screen out multiple high-quality candidate boxes from the candidate box set to form a high-quality candidate box set :

[0151]

[0152] where represents the IoU threshold, represents the annotation box, and G represents the annotation box set.

[0153] Finally, the candidate box set is updated using the high-quality candidate box set, thereby improving the quality of the candidate boxes output by the teacher model. Since during the deep feature knowledge distillation process, the candidate box information and the attention mechanism are used to guide and optimize the deep features of the student model, improving the quality of the candidate boxes can enhance the optimization effect of the deep feature knowledge distillation.

[0154] After obtaining the high-quality candidate boxes, the attention mechanism is used to associate the candidate box regions of the teacher model with the deep features of the student model to generate an attention mask. The specific steps are as follows:

[0155] Merge each candidate box in the candidate box set to generate a candidate box guidance region;

[0156] Through the weight matrix, map the candidate box guidance region to a query matrix, and map the deep feature map of the student model to a key matrix and a value matrix;

[0157] According to the attention mechanism, perform dot product and normalization on the query matrix and the key matrix to obtain the attention mask.

[0158] In this embodiment, first, based on the candidate box set filtered by the teacher model , perform region merging to generate a candidate box guidance region , and its merging strategy can be expressed as:

[0159]

[0160] Among them, the candidate box guidance region is a matrix. represents the pixel coordinates in the image data, , .

[0161] Then, take the candidate box guidance region matrix as the query in the attention mechanism, and take the r-th layer deep feature map of the student model as the key and value respectively, and generate the corresponding query matrix , key matrix and value matrix through weight function mapping. The specific formulas are as follows:

[0162]

[0163] Among them, , , respectively represent the query, key, and value weight functions of the attention mechanism.

[0164] Finally, combine the query matrix and the key matrix of the r-th layer of the student model, and through normalization, calculate the association between the candidate boxes of the teacher model and the deep features of the r-th layer of the student model, so as to obtain the attention mask :

[0165]

[0166] where U is a scalar used to scale the result.

[0167] After obtaining the attention mask, based on the attention mask, align the deep feature maps of the teacher model and the student model, and calculate the deep feature distillation loss. The specific steps include:

[0168] Calculate the deep feature error loss between the deep feature maps of the teacher model and the student model based on the mean square error, and calculate the inner product between the attention mask and the deep feature error loss;

[0169] Calculate the deep feature distillation loss according to the number of deep feature layers of the student model and the inner product.

[0170] In this embodiment, the mean square error is used to characterize the difference between the teacher model and the student model in terms of deep features. At the same time, the importance of the candidate box region of the teacher model is characterized by the attention mask. Therefore, first calculate the deep feature error loss between the deep feature maps of the teacher model and the student model based on the mean square error , then calculate the inner product between the attention mask and the deep feature error loss , and finally, based on the number of deep feature layers T of the student model, calculate the deep feature knowledge distillation loss:

[0171]

[0172] In the formula, represents the deep feature error loss, represents the attention mask of the r-th layer of deep features, represents the inner product, and T represents the number of deep feature layers. It should be noted that the representation method of the mean square error used for the deep feature difference in this embodiment is only a preference rather than a specific limitation. Similar to the shallow feature difference part, the deep feature difference can also be characterized by other error measurement methods such as mean absolute error, Huber loss, squared absolute error, logarithmic cosine loss, etc.

[0173] Based on the candidate boxes generated by the teacher model, this embodiment combines the attention mechanism to calculate the correlation between the candidate box regions of the teacher model and the deep features of the student model, generates an attention mask to optimize the student model's learning of the deep features, and effectively improves the model performance by making full use of the candidate box information to guide the student model to focus on the key regions in the features.

[0174] Furthermore, to solve the problem of the limited performance of the detection head of the student model in the existing methods, the present invention designs response knowledge distillation, which aims to align the detection results of the teacher model and the student model. During the response knowledge distillation process, the detection results of the teacher model are used as soft labels, and the difference analysis is respectively performed on the classification probabilities and the bounding box regression results in the detection results of the teacher model and the student model, so as to obtain the response knowledge distillation loss. The specific steps include:

[0175] Through the teacher model and the student model, perform classification regression prediction on their respective candidate boxes to obtain the detection results of the teacher model and the detection results of the student model. The detection results of the teacher model include the classification probability distribution of the teacher model and the regression results of the teacher model, and the detection results of the student model include the classification probability distribution of the student model and the regression results of the student model;

[0176] Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on the KL divergence;

[0177] Calculate the regression response knowledge loss between the regression results of the teacher model and the regression results of the student model based on the mean square error;

[0178] Perform weighted summation on the classification response knowledge loss and the regression response knowledge loss to obtain the response knowledge distillation loss.

[0179] In this embodiment, for the detection results of the teacher model and the student model, feature alignment is performed. Based on the above model settings, the detection results of the model are based on the detection head in the model to perform classification regression prediction on the input candidate boxes, so as to obtain the classification probability distribution and the regression results. Then, knowledge distillation is performed on the detection results of the teacher model and the student model.

[0180] In the existing knowledge distillation methods, there is a lack of attention to the quality of the input candidate boxes in the detection head part of the object detection model. Therefore, the quality of the candidate boxes extracted by the student model is poor. Due to the lack of guidance for the candidate box quality, it is difficult for the student model to obtain high-quality input at the detection head stage, resulting in limited optimization ability of the detection head and further limiting the improvement of the final detection accuracy. To solve this problem, in a preferred embodiment, the present invention uses the candidate boxes of the teacher model as the candidate boxes of the student model, that is, the set of high-quality candidate boxes output by the teacher model As the candidate boxes of the student model, so as to realize the guidance of the quality of the candidate boxes of the student model.

[0181] Then, the respective candidate boxes are respectively input into the detection heads of the teacher model and the student model for classification regression prediction:

[0182]

[0183] Among them, represents the detection head of the model.

[0184] Thus, the classification probability distribution and the regression result output by the teacher model are obtained, as well as the classification probability distribution and the regression result

[0185] output by the student model. , , , , where N represents N detected targets.

[0186] Then, the KL divergence is used to calculate the classification response knowledge loss between the classification probability distributions: :

[0187]

[0188] In the formula, L kl represents the KL divergence, represents the i-th candidate box, , represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, and N represents the total number of candidate boxes.

[0189] The mean square error is used to calculate the regression response knowledge loss between the regression results of the teacher model and the student model :

[0190]

[0191] In the formula, represents the regression result of the i-th candidate box in the teacher model, represents the regression result of the i-th candidate box in the student model.

[0192] Finally, the classification response knowledge loss and the regression response knowledge loss are weighted and summed to obtain the response knowledge distillation loss :

[0193]

[0194] That is:

[0195]

[0196] In the formula, represents the weight coefficient of the classification response knowledge loss, represents the weight coefficient of the regression response knowledge loss.

[0197] Among them, the weights corresponding to the classification response knowledge loss and the regression response knowledge loss can be flexibly set based on the actual situation. When the weights of the classification response knowledge loss and the regression response knowledge loss are the same and both are 1, the response knowledge distillation loss can be expressed as:

[0198]

[0199] It should be noted that using KL divergence and mean squared error to characterize the difference in the knowledge distillation process at this stage is only a preferred method rather than a specific limitation. For KL divergence, cross-entropy loss or Hinge loss and other error measurement methods can also be used for replacement. For mean squared error, other error measurement methods such as mean absolute error, Huber loss, squared absolute error, and logarithmic cosine loss can be used for replacement and characterization, which will not be elaborated here one by one.

[0200] This embodiment uses the candidate boxes and detection head outputs generated by the teacher model to guide the student model to complete the classification and bounding box regression tasks. By comparing the prediction distribution of the teacher model and the bounding box regression results, the performance of the detection head of the student model is optimized. And through effective guidance on the candidate boxes of the student model, the optimization ability and final detection accuracy of the detection head of the student model are effectively improved.

[0201] The shallow feature knowledge distillation, deep feature knowledge distillation, and response feature knowledge distillation provided by the present invention can solve the problems existing in the existing knowledge distillation technology at different stages of object detection. These three knowledge distillation methods can be used jointly, or any one or more of them can be selected based on the actual situation for use, and at the same time, they can also be combined with the existing knowledge distillation technology for use.

[0202] When using at least two knowledge distillation methods for model training, in a preferred embodiment, the present invention also provides a joint optimization strategy. By weighted summing the knowledge distillation losses corresponding to the selected multiple knowledge distillation methods, a multi-task optimization framework is constructed. For the multi-task optimization framework, preferably, the SGD optimizer in the Pytorch framework is used to solve the parameters in the model.

[0203] Taking the selection of three knowledge distillation methods as an example, the loss function of the multi-task optimization framework can be expressed as:

[0204]

[0205] In the formula, represents the weight coefficient of the shallow feature knowledge distillation loss, represents the weight coefficient of the deep feature knowledge distillation loss, represents the weight coefficient of the response knowledge distillation loss. These weight coefficients can be flexibly adjusted according to the different phased focuses of the model, so as to balance the distillation optimization at different stages.

[0206] In this embodiment, by integrating the losses of all feature alignment modules, end-to-end optimization is achieved, and the importance of different alignment tasks is balanced through weight coefficients, so as to ensure that the student model can still maintain high detection accuracy after being lightweight.

[0207] A transmission line defect detection method based on knowledge distillation provided in this embodiment aims at the problems existing in traditional methods, such as the lack of shallow feature distillation, the failure to fully utilize candidate box information, and the limited performance of the detection head of the student model. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model. By designing shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation, multi-level knowledge transfer from the teacher model to the student model is realized. Among them, shallow feature knowledge distillation optimizes the alignment of underlying visual features by using the foreground-background region mask technology, improving the student model's ability to perceive details; deep feature knowledge distillation combines candidate box information with the attention mechanism, enhancing the feature expression ability of key regions and effectively improving the performance of the student model; response knowledge distillation guides the optimization of the detection head through the soft labels of classification and regression results, improving the optimization ability of the student model's detection head and the final detection accuracy; in addition, the present invention also provides a joint optimization strategy. By integrating the loss functions of multi-level feature distillation, candidate box knowledge guidance, and response knowledge distillation, a multi-task optimization framework is constructed to form a full-process knowledge transfer from features to decisions, improving the overall performance of the student model.

[0208] The method provided by the present invention can effectively avoid ignoring key details due to the simplification of the student model. While reducing the computational complexity of the student model, the student model retains the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on terminal or edge devices, and still has excellent detection effects even in resource-constrained scenarios. Through the defect detection method provided by the present invention, the overall efficiency of transmission line defect detection is significantly improved, providing an effective solution for large-scale, real-time, and lightweight defect detection applications, and having good practicability and scalability.

[0209] Please refer to Figure 3

[0210] An image acquisition module 10, configured to acquire a to-be-detected image of a transmission line, where the to-be-detected image includes image data and video frame data;

[0211] A defect detection module 20, configured to input the to-be-detected image into a pre-constructed defect detection model to obtain a defect detection result of the transmission line, where the defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation;

[0212]

[0213] Among them, in the teacher-student training method, the student model is the defect detection model, and the teacher model is a pre-trained model for defect detection of the transmission line. The network architectures of the teacher model and the student model are the same, but the network depths are different; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation;

[0214] When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss according to the shallow feature differences in different regions;

[0215] When using deep feature knowledge distillation for model training, based on the attention mechanism, associate the candidate boxes of the teacher model with the deep features of the student model to generate an attention mask, and calculate the deep feature distillation loss;

[0216] When using response knowledge distillation for model training, use the detection result of the teacher model as a soft label, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss.

[0217] Further, the defect detection module 20 is further configured to, when using at least two knowledge distillation methods among shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation for model training, perform a weighted sum of the distillation losses corresponding to the used knowledge distillation methods to construct a multi-task optimization framework and optimize and train the student model.

[0218]

[0219] ​​The shallow feature knowledge distillation module is used to input the image data in the dataset into the teacher model and the student model respectively for shallow feature extraction, obtaining the shallow feature map of the teacher model and the shallow feature map of the student model;

[0220] According to the set of annotation boxes corresponding to the input image data, the image data is divided into regions, and region masks are generated according to the divided regions;

[0221] According to the region masks and the error metric method, the shallow feature differences of the shallow feature maps of the teacher model and the student model in different regions are calculated respectively;

[0222] The shallow feature differences in different regions are weighted and summed to obtain the shallow feature knowledge distillation loss.

[0223] Furthermore, the defect detection module 20 includes a deep feature knowledge distillation module;

[0224] The deep feature knowledge distillation module is used to perform feature extraction on the respective shallow feature maps in the teacher model and the student model to obtain the deep feature map of the teacher model and the deep feature map of the student model;

[0225] The teacher model performs target box selection on the deep feature map of the teacher model to obtain the set of candidate boxes of the teacher model;

[0226] According to the set of candidate boxes, a candidate box guidance region is generated, and the candidate box guidance region is mapped to the spatial dimension of the deep feature map of the student model through linear transformation to generate an attention mask;

[0227] According to the attention mask, the deep feature maps of the teacher model and the student model are aligned to calculate the deep feature distillation loss.

[0228] Furthermore, the defect detection module 20 includes a response knowledge distillation module;

[0229] The response knowledge distillation module is used to perform classification regression prediction on the respective candidate boxes through the teacher model and the student model to obtain the detection results of the teacher model and the detection results of the student model. The detection results of the teacher model include the classification probability distribution of the teacher model and the regression result of the teacher model, and the detection results of the student model include the classification probability distribution of the student model and the regression result of the student model;

[0230] Based on the KL divergence, the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model is calculated;

[0231] Based on the mean square error, the regression response knowledge loss between the regression result of the teacher model and the regression result of the student model is calculated;

[0232] The classification response knowledge loss and the regression response knowledge loss are weighted and summed to obtain the response knowledge distillation loss.

[0233] The technical features and technical effects of the transmission line defect detection system based on knowledge distillation proposed in the embodiments of the present invention are the same as those of the method proposed in the embodiments of the present invention, and will not be elaborated here. Each module in the above-mentioned transmission line defect detection system based on knowledge distillation can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0234] In addition, an embodiment of the present invention also proposes a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0235] Please refer to Figure 4 , the internal structure diagram of the computer device in one embodiment. The computer device may specifically be a terminal or a server. The computer device includes a processor, a memory, a network interface, a display, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the method for detecting transmission line defects based on knowledge distillation. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0236] Those of ordinary skill in the art can understand that Figure 4 the structure shown in

[0237] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0238] In summary, a transmission line defect detection method and system based on knowledge distillation proposed in an embodiment of the present invention. The method includes obtaining a to-be-detected image of a transmission line, where the to-be-detected image includes image data and video frame data; inputting the to-be-detected image into a pre-constructed defect detection model to obtain a defect detection result of the transmission line. The defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation. Among them, the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of the transmission line. The network architectures of the teacher model and the student model are the same, but the network depths are different. The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation. When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss based on the shallow feature differences in different regions. When using deep feature knowledge distillation for model training, based on the attention mechanism, associate the candidate boxes of the teacher model with the deep features of the student model to generate an attention mask, and calculate the deep feature distillation loss. When using response knowledge distillation for model training, use the detection result of the teacher model as a soft label, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss. By designing shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation, the present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model, realizes multi-level knowledge transfer from the teacher model to the student model, effectively avoids ignoring key details due to the simplification of the student model, while reducing the computational complexity of the student model, enables the student model to retain the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on terminals or edge devices, and still has excellent detection effects even in resource-constrained scenarios. Through the defect detection method provided by the present invention, the overall efficiency of transmission line defect detection is significantly improved, providing an effective solution for large-scale, real-time, and lightweight defect detection applications, and having good practicability and generalizability.

[0239] The various embodiments in this specification are described in a progressive manner. For parts that are the same or similar in the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, reference can be made to the corresponding parts in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.

[0240] The above-described embodiments merely represent several preferred embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the protection scope of the claimed rights.

Claims

1. A transmission line defect detection method based on knowledge distillation, characterized in that Including: Obtain a to-be-detected image of a power transmission line, where the to-be-detected image includes image data and video frame data; Input the to-be-detected image into a pre-constructed defect detection model to obtain a defect detection result of the power transmission line. The defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; Among them, the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of power transmission lines. The network architectures of the teacher model and the student model are the same, but the network depths are different; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation; When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss according to the shallow feature differences in different regions; When using deep feature knowledge distillation for model training, based on the attention mechanism, associate the candidate boxes of the teacher model with the deep features of the student model to generate an attention mask, and calculate the deep feature knowledge distillation loss; When using response knowledge distillation for model training, use the detection result of the teacher model as a soft label, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss; Among them, the step of associating the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism to generate an attention mask and calculate the deep feature knowledge distillation loss includes: In the teacher model and the student model, respectively extract features from their respective shallow feature maps to obtain the deep feature map of the teacher model and the deep feature map of the student model; Perform target box selection on the deep feature map of the teacher model through the teacher model to obtain a set of candidate boxes of the teacher model; According to the set of candidate boxes, generate a candidate box guidance region, and map the candidate box guidance region to the spatial dimension of the deep feature map of the student model through a linear transformation to generate an attention mask; According to the attention mask, align the deep feature map of the teacher model and the deep feature map of the student model, and calculate the deep feature knowledge distillation loss, including: Calculate the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model based on the mean square error, and calculate the inner product between the attention mask and the deep feature error loss; Calculate the deep feature knowledge distillation loss according to the number of deep feature layers of the student model and the inner product.

2. The method for detecting transmission line defects based on knowledge distillation according to claim 1, wherein When using at least two knowledge distillation methods among shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation for model training, perform weighted summation on the distillation losses corresponding to the used knowledge distillation methods to construct a multi-task optimization framework and optimize and train the student model.

3. The method for detecting transmission line defects based on knowledge distillation according to claim 1, characterized in that, The steps of calculating the shallow feature differences between the teacher model and the student model in different regions according to the region mask technology and the error metric method, and calculating the shallow feature knowledge distillation loss based on the shallow feature differences in different regions are as follows: Input the image data in the dataset into the teacher model and the student model respectively for shallow feature extraction, obtaining the shallow feature maps of the teacher model and the shallow feature maps of the student model; According to the set of annotation boxes corresponding to the input image data, divide the image data into regions, and generate a region mask according to the divided regions; According to the region mask and the error metric method, calculate the shallow feature differences between the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions respectively; Perform weighted summation on the shallow feature differences in different regions to obtain the shallow feature knowledge distillation loss.

4. The method for detecting transmission line defects based on knowledge distillation according to claim 3, wherein The steps of dividing the image data into regions according to the set of annotation boxes corresponding to the input image data, and generating a region mask according to the divided regions are as follows: Take the regions covered by each annotation box in the set of annotation boxes corresponding to the image data as the foreground regions, and take the regions outside each annotation box as the background regions; Mark the mask values corresponding to each pixel point in the foreground regions as 1, and mark the mask values corresponding to each pixel point in the background regions as 0, obtaining the region mask.

5. The method for detecting transmission line defects based on knowledge distillation according to claim 4, wherein The steps of calculating the shallow feature differences between the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions respectively according to the region mask and the error metric method are as follows: According to the region mask and the mean squared error, calculate the error loss between the shallow feature maps of the teacher model and the shallow feature maps of the student model in the foreground regions; According to the inverted value of the region mask and the mean squared error, calculate the error loss between the shallow feature maps of the teacher model and the shallow feature maps of the student model in the background regions.

6. The method for detecting transmission line defects based on knowledge distillation according to claim 5, characterized in that The shallow feature knowledge distillation loss is represented by the following formula: In the formula, represents the regional mask at the coordinate point (h, w), represents the eigenvalue of the shallow feature map of the student model at the coordinate point (h, w) in channel c of the th layer, represents the total number of pixels in the foreground region, represents the foreground region weight loss coefficient, represents the total number of pixels in the background region, represents the background region weight loss coefficient.

7. The method for detecting transmission line defects based on knowledge distillation according to claim 1, wherein The steps of performing object bounding on the deep feature map of the teacher model through the teacher model to obtain the set of candidate boxes of the teacher model are as follows: Perform object bounding on the deep feature map of the teacher model through the teacher model to obtain multiple candidate boxes, and form a set of candidate boxes; Calculate the intersection over union (IoU) between each candidate box in the set of candidate boxes and each annotation box in the set of annotation boxes corresponding to the input image data; According to the comparison relationship between the IoU and the IoU threshold, filter out multiple high-quality candidate boxes from the set of candidate boxes to form a set of high-quality candidate boxes, and update the set of candidate boxes according to the set of high-quality candidate boxes.

8. The method for detecting transmission line defects based on knowledge distillation according to claim 1, wherein The steps of generating a candidate box guidance region according to the set of candidate boxes, and mapping the candidate box guidance region to the spatial dimension of the deep feature map of the student model through a linear transformation to generate an attention mask are as follows: Merge each candidate box in the set of candidate boxes to generate a candidate box guidance region; Map the candidate box guidance region to a query matrix through a weight matrix, and map the deep feature map of the student model to a key matrix and a value matrix; According to the attention mechanism, perform dot product and normalization on the query matrix and the key matrix to obtain the attention mask.

9. The method for detecting transmission line defects based on knowledge distillation according to claim 1, wherein, The deep feature knowledge distillation loss is represented by the following formula: In the formula, represents the deep feature error loss, represents the attention mask of the r-th layer deep feature, represents the inner product, and T represents the number of deep feature layers.

10. The method for detecting transmission line defects based on knowledge distillation according to claim 1, characterized in that, Before the step of performing object bounding on the deep feature map of the teacher model to obtain the set of candidate bounding boxes of the teacher model, it further includes: Align the deep feature map of the student model with the deep feature map of the teacher model, and the alignment operation includes size alignment and resolution alignment.

11. The method for detecting transmission line defects based on knowledge distillation according to claim 1, characterized in that, The step of using the detection result of the teacher model as a soft label and calculating the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model based on an error metric method to obtain the response knowledge distillation loss includes: Through the teacher model and the student model, perform classification and regression predictions on their respective candidate bounding boxes to obtain the detection result of the teacher model and the detection result of the student model. The detection result of the teacher model includes the classification probability distribution of the teacher model and the regression result of the teacher model. The detection result of the student model includes the classification probability distribution of the student model and the regression result of the student model; Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on the KL divergence; Calculate the regression response knowledge loss between the regression result of the teacher model and the regression result of the student model based on the mean square error; Perform weighted summation on the classification response knowledge loss and the regression response knowledge loss to obtain the response knowledge distillation loss.

12. The method for detecting transmission line defects based on knowledge distillation according to claim 11, wherein The step of performing classification and regression predictions on their respective candidate bounding boxes through the teacher model and the student model to obtain the detection result of the teacher model and the detection result of the student model includes: Use the candidate bounding boxes of the teacher model as the candidate bounding boxes of the student model, and through the teacher model and the student model, perform classification and regression predictions on their respective candidate bounding boxes to obtain the detection result of the teacher model and the detection result of the student model.

13. The method for detecting transmission line defects based on knowledge distillation according to claim 11, characterized in that, The response knowledge distillation loss is expressed by the following formula: where L kl represents the KL divergence, represents the i-th candidate box, N represents the total number of candidate boxes, represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, represents the regression result of the i-th candidate box in the teacher model, represents the regression result of the i-th candidate box in the student model, represents the weight coefficient of the classification response knowledge loss, represents the weight coefficient of the regression response knowledge loss.

14. The method for detecting transmission line defects based on knowledge distillation according to claim 1, wherein Both the teacher model and the student model are constructed using a fast convolutional neural network. The teacher model uses ResNet-101 as the backbone network, and the student model uses ResNet-50 as the backbone network.

15. The method for detecting transmission line defects based on knowledge distillation according to claim 14, characterized in that, The network architectures of the teacher model and the student model both include a shallow feature extraction network, a deep feature extraction network, a region candidate network, and a detection head; Among them, the shallow feature extraction network is the first three stages of the backbone network, which is used to perform shallow feature extraction on the input image data; The deep feature extraction network is the last two stages of the backbone network, which is used to perform deep feature extraction on the shallow features output by the shallow feature extraction network; The region candidate network is used to perform object bounding on the deep features output by the deep feature extraction network to generate candidate bounding boxes; The detection head is used to perform classification and regression predictions on the candidate bounding boxes output by the region candidate network and output the detection result.

16. The method for detecting transmission line defects based on knowledge distillation according to claim 2, wherein Use the SGD optimizer to solve the multi-task optimization framework to obtain the optimal parameters of the student model.

17. A transmission line defect detection system based on knowledge distillation, characterized in that It includes: An image acquisition module, which is used to acquire the image to be detected of the transmission line. The image to be detected includes image data and video frame data; A defect detection module, which is used to input the image to be detected into a pre-constructed defect detection model to obtain the defect detection result of the transmission line. The defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; Among them, the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of transmission lines. The network architectures of the teacher model and the student model are the same, but the network depths are different; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation; When using shallow feature knowledge distillation for model training, according to the region masking technique and the error metric method, calculate the shallow feature differences between the teacher model and the student model in different regions, and calculate the shallow feature knowledge distillation loss based on the shallow feature differences in different regions; When using deep feature knowledge distillation for model training, based on the attention mechanism, associate the candidate boxes of the teacher model with the deep features of the student model to generate an attention mask, and calculate the deep feature knowledge distillation loss; When using response knowledge distillation for model training, use the detection results of the teacher model as soft labels, and calculate the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error metric method to obtain the response knowledge distillation loss; Among them, the associating the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism, generating an attention mask, and calculating the deep feature knowledge distillation loss includes: In the teacher model and the student model, respectively extract features from their respective shallow feature maps to obtain the deep feature map of the teacher model and the deep feature map of the student model; Use the teacher model to perform object bounding on the deep feature map of the teacher model to obtain a set of candidate boxes of the teacher model; According to the set of candidate boxes, generate a candidate box guidance region, and map the candidate box guidance region to the spatial dimension of the deep feature map of the student model through a linear transformation to generate an attention mask; According to the attention mask, align the deep feature map of the teacher model and the deep feature map of the student model, and calculate the deep feature knowledge distillation loss, including: Calculate the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model based on the mean square error, and calculate the inner product between the attention mask and the deep feature error loss; Calculate the deep feature knowledge distillation loss according to the number of deep feature layers of the student model and the inner product.

Citation Information

Patent Citations

  • SAR (Synthetic Aperture Radar) image target detection method combined with high-credibility knowledge distillation

    CN115761511A

  • Knowledge distillation-based lightweight power transmission line fault detection method

    CN117746083A

  • Steel plate surface defect detection method and system based on knowledge distillation

    CN118096768A