Power transmission line defect detection method and system based on knowledge distillation
By using knowledge distillation technology in transmission line defect detection, the multi-level knowledge transfer between teacher models and student models is achieved, and the shortcomings of existing methods in computing complexity and resource-constrained scenarios are solved, and efficient and lightweight defect detection effect is achieved.
Patent Information
- Application Number
- CN202510412959.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing transmission line defect detection methods have shortcomings in terms of calculation complexity, model volume and training time, and are especially difficult to deploy directly in resource-constrained scenarios.
Using a knowledge distillation method, through shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation, the multi-level knowledge transfer of teacher models to student models is realized, reducing the computational complexity of student models, while maintaining high detection accuracy.
While keeping the model lightweight, it improves the overall efficiency of transmission line defect detection, can be efficiently deployed on terminals or edge devices, and has excellent detection results. It is suitable for large-scale, real-time, and lightweight defect detection applications.
Smart Images

Figure CN119941714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power transmission line defect detection, and in particular to a power transmission line defect detection method and system based on knowledge distillation. Background Art
[0002] In the power system, transmission lines are important channels for power transmission, which are easily affected by natural and human factors and produce various defects. In recent years, with the development of digital image processing and analysis technology in security monitoring, drone cruising and other scenarios, target detection technology based on digital image processing has gradually been applied to power system transmission line defect detection.
[0003] At present, the methods for target detection of transmission line defects are mainly divided into two categories: methods based on manual features and methods based on deep learning. Among them, the methods based on manual features mainly include sliding window method, template matching method and classification method using features such as SIFT and HOG; the methods based on deep learning mainly include automatic extraction of image features through neural network models and combined with classification and regression tasks to realize the detection of transmission line defects. However, both of them have certain limitations. On the one hand, the methods based on manual features have achieved certain results in early target detection, but due to the reliance on artificially designed features, they have poor adaptability to complex scenes and target changes and are difficult to meet the needs of actual applications; on the other hand, although the methods based on deep learning can significantly improve the performance of transmission line defect target detection, they still have shortcomings in terms of computational complexity, model volume and training time, especially in resource-constrained scenarios. It is difficult to deploy directly. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a transmission line defect detection method and system based on knowledge distillation, so as to solve the shortcomings of the existing target detection model in terms of computational complexity, model volume and training time, so that the model can maintain lightweight while also having high-precision detection capabilities.
[0005] In a first aspect, the present invention provides a method for detecting power transmission line defects based on knowledge distillation, the method comprising: Acquire an image to be detected of a power transmission line, wherein the image to be detected includes image data and video frame data; Inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and trained using a teacher-student training method based on knowledge distillation; The student model in the teacher-student training method is the defect detection model, the teacher model is a pre-trained model for defect detection of power transmission lines, the teacher model and the student model have the same network architecture, but different network depths; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation; When shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated based on the regional masking technique and the error measurement method, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions; When deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; When response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
[0006] Furthermore, when at least two knowledge distillation methods among shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation are used for model training, the distillation losses corresponding to the adopted knowledge distillation methods are weighted summed up, and a multi-task optimization framework is constructed to optimize the student model.
[0007] Furthermore, the step of calculating the shallow feature differences between the teacher model and the student model in different regions according to the regional mask technology and the error measurement method, and calculating the shallow feature knowledge distillation loss according to the shallow feature differences in different regions includes: The image data in the data set are input into the teacher model and the student model for shallow feature extraction, and the shallow feature map of the teacher model and the shallow feature map of the student model are obtained; According to the set of annotation boxes corresponding to the input image data, the image data is divided into regions, and a regional mask is generated according to the divided regions; According to the regional mask and error measurement method, the shallow feature differences of the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions are calculated respectively; The shallow feature differences in different regions are weighted and summed to obtain the shallow feature knowledge distillation loss.
[0008] Furthermore, the step of dividing the image data into regions according to the set of annotation boxes corresponding to the input image data, and generating a regional mask according to the divided regions includes: The area covered by each annotation box in the annotation box set corresponding to the image data is regarded as the foreground area, and the area outside each annotation box is regarded as the background area; The mask value corresponding to each pixel point in the foreground area is marked as 1, and the mask value corresponding to each pixel point in the background area is marked as 0, so as to obtain a regional mask.
[0009] Furthermore, the step of respectively calculating the shallow feature differences of the shallow feature map of the teacher model and the shallow feature map of the student model in different regions according to the regional mask and the error measurement method includes: According to the regional mask and mean square error, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the foreground area is calculated; According to the inverted value and mean square error of the region mask, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the background area is calculated.
[0010] Furthermore, the following formula is used to express the shallow feature knowledge distillation loss: In the formula, represents the region mask at the coordinate point (h, w), Indicates The eigenvalue of the shallow feature map of the student model of the layer at the coordinate point (h, w) of channel c, Indicates The eigenvalue of the shallow feature map of the teacher model of the layer at the coordinate point (h, w) of channel c, represents the total number of pixels in the foreground area, represents the weight loss coefficient of the foreground area, represents the total number of pixels in the background area, Represents the background area weight loss coefficient.
[0011] Furthermore, the steps of associating the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism, generating an attention mask, and calculating the deep feature distillation loss include: By extracting features from the shallow feature maps of the teacher model and the student model, respectively, the deep feature maps of the teacher model and the student model are obtained; The teacher model is used to select the target frame of the deep feature map of the teacher model to obtain the candidate frame set of the teacher model; According to the candidate box set, a candidate box guidance region is generated, and the candidate box guidance region is mapped to the spatial dimension of the deep feature map of the student model through linear transformation to generate an attention mask; According to the attention mask, the deep feature map of the teacher model is aligned with the deep feature map of the student model, and the deep feature distillation loss is calculated.
[0012] Furthermore, the step of performing target box selection on the deep feature map of the teacher model through the teacher model to obtain a candidate box set of the teacher model includes: The teacher model is used to select the target frame of the deep feature map of the teacher model, and multiple candidate frames are obtained and formed into a candidate frame set; Calculate the intersection-and-union ratio between each candidate box in the candidate box set and each annotation box in the annotation box set corresponding to the input image data; According to the comparison relationship between the IoU ratio and the IoU ratio threshold, multiple high-quality candidate boxes are screened out from the candidate box set to form a high-quality candidate box set, and the candidate box set is updated according to the high-quality candidate box set.
[0013] Furthermore, the step of generating a candidate box guidance region according to the candidate box set, and mapping the candidate box guidance region to the spatial dimension of the deep feature map of the student model through a linear transformation, and generating an attention mask includes: Merge each candidate box in the candidate box set to generate a candidate box guidance area; Through the weight matrix, the candidate box guidance area is mapped to the query matrix, and the deep feature map of the student model is mapped to the key matrix and the value matrix; According to the attention mechanism, the query matrix and the key matrix are dot-producted and normalized to obtain the attention mask.
[0014] Furthermore, the step of aligning the deep feature map of the teacher model and the deep feature map of the student model according to the attention mask and calculating the deep feature distillation loss includes: Calculate the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model based on the mean squared error, and calculate the inner product between the attention mask and the deep feature error loss; According to the number of deep feature layers of the student model and the inner product, the deep feature distillation loss is calculated.
[0015] Furthermore, it is characterized in that the deep feature knowledge distillation loss is expressed by the following formula: In the formula, represents the deep feature error loss, represents the attention mask of the deep features of the rth layer, represents the inner product, and T represents the number of deep feature layers.
[0016] Furthermore, before the step of performing target box selection on the deep feature map of the teacher model through the teacher model to obtain the candidate box set of the teacher model, it also includes: The deep feature map of the student model is aligned with the deep feature map of the teacher model, and the alignment operation includes size alignment and resolution alignment.
[0017] Furthermore, the step of using the detection result of the teacher model as a soft label, calculating the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error measurement method, and obtaining the response knowledge distillation loss includes: Through the teacher model and the student model, classification regression prediction is performed on the respective candidate boxes to obtain the detection result of the teacher model and the detection result of the student model, wherein the detection result of the teacher model includes the classification probability distribution of the teacher model and the regression result of the teacher model, and the detection result of the student model includes the classification probability distribution of the student model and the regression result of the student model; Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on KL divergence; The regression response knowledge loss between the regression results of the teacher model and the regression results of the student model is calculated based on the mean square error; The classification response knowledge loss and the regression response knowledge loss are weighted summed to obtain the response knowledge distillation loss.
[0018] Furthermore, the step of performing classification regression prediction on respective candidate boxes through the teacher model and the student model to obtain the detection result of the teacher model and the detection result of the student model includes: The candidate boxes of the teacher model are used as the candidate boxes of the student model, and the teacher model and the student model are used to perform classification regression prediction on their respective candidate boxes to obtain the detection results of the teacher model and the detection results of the student model.
[0019] Furthermore, the response knowledge distillation loss is expressed by the following formula: Where, L kl represents the KL divergence, represents the i-th candidate box, N represents the total number of candidate boxes, represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, represents the regression result of the i-th candidate box in the teacher model, represents the regression result of the i-th candidate box in the student model, represents the weight coefficient of the knowledge loss of the classification response, The weight coefficient representing the knowledge loss of the regression response.
[0020] Furthermore, both the teacher model and the student model are constructed using fast convolutional neural networks, and the teacher model uses ResNet-101 as the backbone network, and the student model uses ResNet-50 as the backbone network.
[0021] Furthermore, the network architectures of both the teacher model and the student model include a shallow feature extraction network, a deep feature extraction network, a region candidate network, and a detection head; Among them, the shallow feature extraction network is the first three stages of the backbone network, which is used to extract shallow features of the input image data; The deep feature extraction network is the last two stages of the backbone network, which is used to perform deep feature extraction on the shallow features output by the shallow feature extraction network; The region candidate network is used to select the target frame based on the deep features output by the deep feature extraction network and generate candidate frames; The detection head is used to perform classification regression prediction on the candidate boxes output by the region candidate network and output the detection results.
[0022] Furthermore, the SGD optimizer is used to solve the multi-task optimization framework to obtain the optimal parameters of the student model.
[0023] In a second aspect, the present invention provides a power transmission line defect detection system based on knowledge distillation, the system comprising: An image acquisition module, used to acquire an image to be detected of the power transmission line, wherein the image to be detected includes image data and video frame data; A defect detection module, used for inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; The student model in the teacher-student training method is the defect detection model, the teacher model is a pre-trained model for defect detection of power transmission lines, the teacher model and the student model have the same network architecture, but different network depths; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation; When shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated based on the regional masking technique and the error measurement method, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions; When deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; When response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
[0024] The present invention provides a method and system for power transmission line defect detection based on knowledge distillation. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model by designing shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation, thereby realizing multi-level knowledge migration from the teacher model to the student model, effectively avoiding the neglect of key details due to the simplification of the student model, and reducing the computational complexity of the student model while enabling the student model to retain the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on the terminal or edge device, and still has excellent detection effect even in resource-constrained scenarios. The defect detection method provided by the present invention significantly improves the overall efficiency of power transmission line defect detection, provides an effective solution for large-scale, real-time, lightweight defect detection applications, and has good practicality and scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow chart of a power transmission line defect detection method based on knowledge distillation in an embodiment of the present invention; Figure 2 is another flow chart of a power transmission line defect detection method based on knowledge distillation in an embodiment of the present invention; Figure 3 is a structural schematic diagram of a power transmission line defect detection system based on knowledge distillation in an embodiment of the present invention; Figure 4 It is a diagram of the internal structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] See also Figure 1The first embodiment of the present invention proposes a method for detecting power transmission line defects based on knowledge distillation, which includes steps S10 to S20: Step S10, obtaining an image to be detected of the power transmission line, wherein the image to be detected includes image data and video frame data; Step S20, inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation.
[0028] The present invention performs target detection on images of transmission lines through a defect detection model, so as to determine whether there are defects in the transmission lines in the image. The data input into the defect detection model is a set of image data or video frame data containing targets. Before inputting into the model, data preprocessing is required for the images to be detected, including uniformly adjusting the size of the images, standardizing the pixel values in the images, etc. The defect detection model is constructed by using a convolutional neural network model, and a teacher-student training method based on knowledge distillation is adopted when training the model. The predicted target type and bounding box information can be obtained by using the trained defect detection model.
[0029] In model training, the teacher-student training method is a model compression and acceleration technology based on knowledge distillation. Its core idea is to transfer the knowledge of a trained complex model (teacher model) to a smaller and easier to deploy model (student model). The teacher model usually has a high accuracy, but the computational cost is high; the student model is relatively simple and computationally efficient, but its performance may not be as good as the teacher model when trained directly. Through knowledge distillation, the student model can learn the "knowledge" of the teacher model, so as to achieve performance close to that of the teacher model while maintaining a small model size. In other words, the teacher-student training method guides the learning of the student model through the feature knowledge, response knowledge and global knowledge of the teacher model, thereby significantly reducing the number of parameters and computational overhead of the student model while maintaining the model detection accuracy, and improving the performance of the student model.
[0030] At present, conventional teacher-student training methods only consider deep-level feature distillation in the process of knowledge distillation. In fact, in the teacher model, its shallow features and candidate boxes also contain important information, but this information has not been effectively introduced into the knowledge distillation process, resulting in the performance of the student model not being able to meet the requirements of high precision. In order to solve this problem, the present invention provides a teacher-student training method based on knowledge distillation to train a student model, wherein knowledge distillation includes shallow feature knowledge distillation, deep feature knowledge distillation, response knowledge distillation and other knowledge distillation methods. When training the model, according to different actual conditions such as model requirements, any one or more of the multiple knowledge distillation methods can be selected for knowledge distillation. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model to avoid neglecting key details due to model simplification, so that the student model can retain the high detection accuracy of the teacher model to the greatest extent.
[0031] The three knowledge distillation methods provided by the present invention are respectively applied to different stages of model training, wherein shallow feature knowledge distillation is applied to the model in the shallow feature extraction stage, and the shallow feature differences between the teacher model and the student model in different regions are calculated according to the regional masking technology and the error measurement method, and the shallow feature knowledge distillation loss is calculated according to the shallow feature differences in different regions; Deep feature knowledge distillation is applied to the model in the deep feature extraction stage. Based on the attention mechanism, the candidate boxes of the teacher model are associated with the deep features of the student model, the attention mask is generated, and the deep feature distillation loss is calculated; Response knowledge distillation is applied to the candidate box classification regression prediction stage of the model. The detection results of the teacher model are used as soft labels. The classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
[0032] The reason why the present invention provides three knowledge distillation methods is that the existing knowledge distillation methods only consider deep-level feature distillation, but in fact, the shallow features of the teacher model contain rich underlying visual information such as edges and textures, and these underlying visual information have not been knowledge distilled to the student model in the existing methods; at the same time, the candidate frame of the teacher model contains the model's attention information on the key feature areas, and these attentions have not been knowledge distilled to the student model in the existing methods; in addition, the existing methods often lack attention to the quality of the candidate frames in the target detection results, and cannot effectively solve the problem of poor quality of the candidate frames extracted by the student model.
[0033] In order to solve the above problems, the present invention provides three knowledge distillation methods applied to different stages. These three knowledge distillation methods can be applied individually or in combination, and can also be combined with existing knowledge distillation methods for application. For the sake of convenience, the model training method provided by the present invention is explained below through the combined application of three knowledge distillation methods.
[0034] See also Figure 2 The knowledge distillation in the present invention can be divided into three stages: shallow feature knowledge distillation, deep feature knowledge distillation, and response knowledge distillation. For the convenience of description, the teacher model, student model, and model training data set are set as follows: The teacher model is a pre-trained model for defect detection of transmission lines, and the student model is a defect detection model. Both the teacher model and the student model are constructed using the fast convolutional neural network model Faster-RCNN, and the residual network ResNet is used as the backbone network for feature extraction. Among them, the backbone network of the teacher model uses the more complex ResNet-101, and the backbone network of the student model uses the lightweight ResNet-50.
[0035] The training data set for the target detection task includes an image data set, a corresponding category label set and an annotation box set, wherein the annotation box contains the coordinates of the upper left corner and the lower right corner of each target, and the size of each image data is uniformly adjusted, and its pixel value is standardized to unify the input format. It should be noted that the teacher model and the student model in the present invention can also be constructed using other neural network models with target detection functions. Here, only a preferred construction method is given without specific limitation.
[0036] Based on the above settings, when shallow feature knowledge distillation is used to train the defect detection model, the specific knowledge distillation steps include: The image data in the data set are input into the teacher model and the student model for shallow feature extraction, and the shallow feature map of the teacher model and the shallow feature map of the student model are obtained; According to the set of annotation boxes corresponding to the input image data, the image data is divided into regions, and a regional mask is generated according to the divided regions; According to the regional mask and error measurement method, the shallow feature differences of the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions are calculated respectively; The shallow feature differences in different regions are weighted and summed to obtain the shallow feature knowledge distillation loss.
[0037] In this embodiment, shallow feature knowledge distillation is to perform knowledge distillation on the shallow features of the teacher model and the student model, with the goal of aligning the shallow features (such as edges and textures) of the student model with the teacher model. Therefore, it is necessary to first obtain the shallow features of the teacher model and the student model: input the image data in the data set into the teacher model and the student model respectively for shallow feature extraction, thereby obtaining the shallow feature map of the teacher model and the shallow feature map of the student model.
[0038] Based on the above model settings, the network architecture of the teacher model and the student model includes four parts, namely, a shallow feature extraction network, a deep feature extraction network, a region candidate network and a detection head. Since the backbone networks of the teacher model and the student model adopt residual networks, the network structure of the residual network can be divided into five stages, among which the first stage is the input layer, which includes a convolutional layer and a pooling layer, the second stage is the first residual block group, the third stage is the second residual block group, the fourth stage is the third residual block group, and the fifth stage is the fourth residual block group. Each residual block group contains multiple residual blocks. Whether it is a ResNet-101 model or a ResNet-50 model, it can be divided into these five stages. The difference lies in the different network depths, and the number of residual blocks contained in the residual block group will be different.
[0039] Based on the five stages of the residual network, this embodiment uses the first three stages as a shallow feature extraction network and the last two stages as a deep feature extraction network. In addition, since both the teacher model and the student model are constructed using the Faster-RCNN model, the regional candidate network PRN and the detection head are also included after the backbone network of the model. Of course, if the teacher model and student model are constructed using other neural network models, they can be divided into stages according to the above four functions of shallow feature extraction, deep feature extraction, candidate box selection and result detection, which will not be described one by one here.
[0040] Based on the above architecture, in the teacher model, the input image data is captured through a shallow feature extraction network to capture the underlying visual information in the image, such as texture, edges, etc., thereby generating a shallow feature map : In the formula, I q represents the input image data, Z1, Z2 and Z3 represent the first stage, second stage and third stage of the backbone network of the teacher model respectively, Represents a real matrix of dimension H (height) × W (width) × C (number of channels), which is used to describe the shape and data type of the feature map.
[0041] In the student model, the image data is also input into the shallow feature extraction network of the student model to obtain the shallow feature map of the student model. : Where Z1 s 、Z2 s and Z3 s They represent the first, second and third stages of the backbone network of the student model respectively.
[0042] Before performing knowledge distillation on the shallow feature maps of the teacher model and the shallow feature maps of the student model, the present invention divides the image data into regions based on the annotation box of the input image and generates corresponding region masks. The specific steps include: The area covered by each annotation box in the annotation box set corresponding to the image data is regarded as the foreground area, and the area outside each annotation box is regarded as the background area; The mask value corresponding to each pixel point in the foreground area is marked as 1, and the mask value corresponding to each pixel point in the background area is marked as 0, so as to obtain a regional mask.
[0043] In this embodiment, the foreground area and the background area are divided based on the set of annotation boxes in the data set, wherein the foreground area refers to the area containing the target, so the part covered by each annotation box in the annotation box set is taken as the foreground area, and the part outside the annotation box is taken as the background area, and then based on the divided area, a binary area mask is generated, wherein the mask value corresponding to each pixel point in the foreground area is marked as 1, and the mask value corresponding to each pixel point in the background area is marked as 0, thereby generating a regional mask M for the input image data: In the formula, Represents the coordinate value of the pixel point, i represents the row coordinate, j represents the column coordinate, , G represents the set of annotation boxes.
[0044] Then, based on the divided regions, the error measurement method is used to calculate the difference between the shallow features of the teacher model and the student model in each region. The specific steps include: According to the regional mask and mean square error, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the foreground area is calculated; According to the inverted value and mean square error of the region mask, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the background area is calculated.
[0045] In this embodiment, the shallow feature difference is characterized by calculating the mean square error between the shallow feature map of the teacher model and the shallow feature map of the student model. At the same time, the shallow feature difference is represented by region according to the regional mask, where the mean square error loss in the foreground region can be expressed as: The mean square error in the background area is calculated by inverting the area mask: In the formula, represents the region mask at the coordinate point (h,w), Indicates The eigenvalue of the shallow feature map of the student model of the layer at the coordinate point (h, w) of channel c, Indicates The eigenvalue of the shallow feature map of the teacher model of the layer at the coordinate point (h, w) of channel c, represents the total number of pixels in the foreground area, Represents the total number of pixels in the background area.
[0046] Finally, the shallow feature differences in the background area and the shallow feature differences in the foreground area are weighted summed to obtain the shallow feature knowledge distillation loss: + in, Indicated in The shallow feature knowledge distillation loss of the layer, represents the weight loss coefficient of the foreground area, Represents the background area weight loss coefficient.
[0047] It should be noted that the use of mean square error to characterize the difference between features in the present invention is only a preferred difference characterization form. Other error metrics such as mean absolute error, Huber loss, square absolute error, log cosine loss, etc. can also be used to characterize feature differences. There are no excessive restrictions here. That is to say, in the above formula and It can be calculated using a variety of error metrics.
[0048] This embodiment adjusts the importance of different regions through the loss weight coefficients of the foreground region and the background region, where the weight loss coefficient of the foreground region is greater than the weight loss coefficient of the background region. It can be seen that the shallow knowledge distillation loss in this embodiment includes two parts, and in the multi-task loss in knowledge distillation, the influence of different parts may be unbalanced. Therefore, in a preferred embodiment, the above-mentioned shallow feature knowledge distillation loss is normalized to avoid a certain area dominating the entire loss due to too many elements, ensuring that the loss term of each area contributes appropriately to the total loss. Taking the feature difference representation based on mean square error as an example, the formula of the normalized distillation loss can be expressed as: The denominator is added to the weight loss coefficient of the above formula, so that each area is normalized separately when calculating the mean square error, making the loss values of the foreground and background areas comparable, avoiding the dominance of a certain item in the total loss due to differences in area size. At the same time, in back propagation, the gradient of the loss value is proportional to the derivative of the loss function with respect to the parameter. The gradient of a large area (such as the number of background pixels far exceeds that of the foreground) may dominate the optimization direction, causing the model to pay too much attention to the background and ignore the target area. Therefore, the above formula can reduce the gradient amplitude of each area, making the optimization process more stable. Of course, the weight loss coefficient can also be determined by other weight analysis methods. The specific weight loss coefficient can be flexibly set according to the actual situation, and will not be described here one by one.
[0049] In this embodiment, the regional mask technology is used to distinguish the target area (foreground area) and the non-target area (background area) in the image data, and the shallow feature differences between the student model and the teacher model in different areas are calculated respectively, so as to optimize the shallow feature representation of the student model. Through the shallow feature knowledge distillation of the present invention, the student model can be effectively guided to learn the detail information in the shallow features, solve the problem of detail loss caused by lightweighting of the student model, and thus improve the student model's ability to perceive details.
[0050] Furthermore, in view of the problem that the existing methods do not fully utilize the candidate box information output by the teacher model, the present invention provides a deep feature knowledge distillation for deep feature alignment of the teacher model and the student model, provides the candidate box information and attention mechanism of the teacher model, and enhances the semantic feature expression ability of the student model for key areas. The specific steps include: By extracting features from the shallow feature maps of the teacher model and the student model, respectively, the deep feature maps of the teacher model and the student model are obtained; The teacher model is used to select the target frame of the deep feature map of the teacher model to obtain the candidate frame set of the teacher model; According to the candidate box set, a candidate box guidance region is generated, and the candidate box guidance region is mapped to the spatial dimension of the deep feature map of the student model through linear transformation to generate an attention mask; According to the attention mask, the deep feature map of the teacher model is aligned with the deep feature map of the student model, and the deep feature distillation loss is calculated.
[0051] In this embodiment, the purpose of deep feature knowledge distillation is to align the deep features of the teacher model and the student model, such as semantic information, and combine the attention mechanism to focus on key areas. In the process of deep feature knowledge distillation, it is first necessary to obtain the deep features of the teacher model and the student model. Taking the above setting as an example, in the teacher model, the shallow feature map is input into the deep feature extraction network to capture the high-level semantic information of the target area, thereby obtaining the deep feature map of the teacher model. : Where Z4 and Z5 represent the fourth and fifth stages of the backbone network respectively.
[0052] Similarly, in the student model, the shallow feature map is input into the deep feature extraction network of the student model to obtain the deep feature map of the student model. : Where Z4 s and Z5 s They represent the fourth and fifth stages of the backbone network of the student model respectively.
[0053] After obtaining the deep feature maps of the teacher model and the student model, it is also necessary to Perform alignment operations to ensure that the deep feature maps of the teacher model And the transformed student model deep feature map Keep the size and resolution consistent to facilitate subsequent processing.
[0054] After the alignment operation, the region candidate network of the teacher model is used to select the target box of the deep feature map of the teacher model, so as to obtain a candidate box set composed of multiple candidate boxes output by the teacher model. This candidate box set will be used as the basis and combined with the attention mechanism to calculate the association between the candidate box area of the teacher model and the deep feature map of the student model, so as to optimize the student model's learning of deep features.
[0055] In order to improve the quality of the candidate boxes output by the teacher model, in a preferred embodiment, the step of performing target box selection on the deep feature map through the teacher model to obtain the candidate box set of the teacher model includes: The teacher model is used to select the target frame of the deep feature map of the teacher model, and multiple candidate frames are obtained and formed into a candidate frame set; Calculate the intersection-and-union ratio between each candidate box in the candidate box set and each annotation box in the annotation box set corresponding to the input image data; According to the comparison relationship between the IoU ratio and the IoU ratio threshold, multiple high-quality candidate boxes are screened out from the candidate box set to form a high-quality candidate box set, and the candidate box set is updated according to the high-quality candidate box set.
[0056] In this embodiment, the deep feature map of the teacher model is first sent to the region candidate network RPN for target box selection, and multiple candidate boxes are selected to generate a candidate box set B: In the formula, , candidate box It is represented by the coordinates of the upper left corner and the lower right corner, that is, Figure 2 The coordinates in the candidate box information (x i ,y i ) and coordinates (x` i ,y` i ).
[0057] Then calculate the IOU between each candidate box in the candidate box set and each labeled box in the labeled box set in the data set, the comparison relationship between the IOU and the IOU threshold, and select multiple high-quality candidate boxes from the candidate box set to form a high-quality candidate box set. : In the formula, represents the intersection-over-union ratio threshold, represents a label box, and G represents a set of label boxes.
[0058] Finally, a high-quality candidate box set is used to update the candidate box set, thereby improving the quality of the candidate boxes output by the teacher model. In the process of deep feature knowledge distillation, the candidate box information and attention mechanism are used to guide and optimize the deep features of the student model. Therefore, by improving the quality of the candidate boxes, the optimization effect of the deep feature knowledge distillation can be improved.
[0059] After obtaining high-quality candidate boxes, the attention mechanism is used to associate the candidate box area of the teacher model with the deep features of the student model to generate an attention mask. The specific steps include: Merge each candidate box in the candidate box set to generate a candidate box guidance area; Through the weight matrix, the candidate box guidance area is mapped to the query matrix, and the deep feature map of the student model is mapped to the key matrix and the value matrix; According to the attention mechanism, the query matrix and the key matrix are dot-producted and normalized to obtain the attention mask.
[0060] In this embodiment, firstly, based on the candidate box set filtered by the teacher model , perform region merging to generate candidate box guidance regions , its merging strategy can be expressed as: Among them, the candidate box guidance area is a The matrix of . Represented as pixel coordinates in image data, , .
[0061] Then the candidate box guide region matrix As the query in the attention mechanism, the r-th layer deep feature map of the student model As key and value respectively, the corresponding query matrix is generated through weight function mapping , key matrix Sum Matrix The specific formula is as follows: in, , , They represent the query, key, and value weight functions of the attention mechanism respectively.
[0062] Finally, the query matrix of the rth layer of the student model is and key matrix Combined with normalization, the association between the candidate box of the teacher model and the deep features of the r-th layer student model is calculated to obtain the attention mask : Among them, U is a scalar used to scale the result.
[0063] After obtaining the attention mask, the deep feature map of the teacher model and the deep feature map of the student model are aligned based on the attention mask, and the deep feature distillation loss is calculated. The specific steps include: Calculate the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model based on the mean squared error, and calculate the inner product between the attention mask and the deep feature error loss; According to the number of deep feature layers of the student model and the inner product, the deep feature distillation loss is calculated.
[0064] In this embodiment, the mean square error is used to characterize the difference between the teacher model and the student model in terms of deep features. At the same time, the attention mask is used to characterize the importance of the candidate box area of the teacher model. Therefore, the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model is first calculated based on the mean square error. , and then calculate the inner product between the attention mask and the deep feature error loss , and finally, based on the number of deep feature layers T of the student model, the deep feature knowledge distillation loss is calculated: In the formula, represents the deep feature error loss, represents the attention mask of the deep features of the rth layer, Represents the inner product, and T represents the number of deep feature layers. It should be noted that the representation method of mean square error used for deep feature differences in this embodiment is only a preferred method rather than a specific limitation. Similar to the shallow feature differences, deep feature differences can also be represented by other error metrics such as mean absolute error, Huber loss, square absolute error, and log cosine loss.
[0065] This embodiment is based on the candidate box generated by the teacher model, combines the attention mechanism to calculate the correlation between the candidate box area of the teacher model and the deep features of the student model, and generates an attention mask to optimize the student model's learning of deep features. By making full use of the candidate box information, the student model is guided to focus on the key areas in the features, thereby effectively improving the model performance.
[0066] Furthermore, in order to solve the problem of limited performance of the student model detection head in the existing methods, the present invention designs response knowledge distillation, the purpose of which is to align the detection results of the teacher model and the student model. In the response knowledge distillation process, the detection results of the teacher model are used as soft labels, and the classification probabilities and bounding box regression results of the teacher model and the student model in the detection results are analyzed respectively to obtain the response knowledge distillation loss. The specific steps include: Through the teacher model and the student model, classification regression prediction is performed on the respective candidate boxes to obtain the detection result of the teacher model and the detection result of the student model, wherein the detection result of the teacher model includes the classification probability distribution of the teacher model and the regression result of the teacher model, and the detection result of the student model includes the classification probability distribution of the student model and the regression result of the student model; Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on KL divergence; The regression response knowledge loss between the regression results of the teacher model and the regression results of the student model is calculated based on the mean square error; The classification response knowledge loss and the regression response knowledge loss are weighted summed to obtain the response knowledge distillation loss.
[0067] In this embodiment, feature alignment is performed on the detection results of the teacher model and the student model. Based on the above model settings, the detection results of the model are based on the classification and regression prediction of the candidate boxes input by the detection head in the model, thereby obtaining the classification probability distribution and regression results. Then, knowledge distillation is performed on the detection results of the teacher model and the student model.
[0068] In the existing knowledge distillation methods, the detection head of the target detection model lacks attention to the quality of the input candidate boxes, which leads to poor quality of the candidate boxes extracted by the student model. Due to the lack of guidance on the quality of the candidate boxes, it is difficult for the student model to obtain high-quality input at the detection head stage, resulting in limited optimization capabilities of the detection head, which in turn limits the improvement of the final detection accuracy. In order to solve this problem, in a preferred embodiment, the present invention uses the candidate boxes of the teacher model as the candidate boxes of the student model, that is, the high-quality candidate boxes output by the teacher model are used as the candidate boxes of the student model. As the candidate box of the student model, it can guide the quality of the candidate box of the student model.
[0069] Then the respective candidate boxes are input into the detection heads of the teacher model and the student model for classification and regression prediction: in, Represents the detection head of the model.
[0070] Thus, the classification probability distribution of the teacher model output is obtained And the regression results , and the classification probability distribution output by the student model And the regression results .
[0071] in, , , , , N represents the N targets detected.
[0072] Then KL divergence is used to calculate the classification response knowledge loss between the classification probability distribution and the classification probability distribution : Where, L kl represents the KL divergence, represents the i-th candidate box, , represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, and N represents the total number of candidate boxes.
[0073] The regression response knowledge loss between the regression results of the teacher model and the regression results of the student model is calculated using the mean square error : In the formula, represents the regression result of the i-th candidate box in the teacher model, Represents the regression result of the i-th candidate box in the student model.
[0074] Finally, the classification response knowledge loss and regression response knowledge loss are weighted summed to obtain the response knowledge distillation loss. : That is: In the formula, represents the weight coefficient of the knowledge loss of the classification response, The weight coefficient representing the knowledge loss of the regression response.
[0075] Among them, the weights corresponding to the classification response knowledge loss and the regression response knowledge loss can be flexibly set based on the actual situation. When the weights of the classification response knowledge loss and the regression response knowledge loss are the same and both are 1, the response knowledge distillation loss can be expressed as: It should be noted that the KL divergence and mean square error used to characterize the difference in the knowledge distillation process at this stage is only a preferred method rather than a specific limitation. The KL divergence can also be replaced by error metrics such as cross entropy loss or Hinge loss, and the mean square error can be replaced by other error metrics such as mean absolute error, Huber loss, square absolute error, log cosine loss, etc., which will not be described one by one here.
[0076] This embodiment uses the candidate boxes and detection head output generated by the teacher model to guide the student model to complete the classification and bounding box regression tasks, and optimizes the detection head performance of the student model by comparing the predicted distribution and bounding box regression results of the teacher model. And by effectively guiding the student model candidate boxes, the optimization ability and final detection accuracy of the student model detection head are effectively improved.
[0077] The shallow feature knowledge distillation, deep feature knowledge distillation and response feature knowledge distillation provided by the present invention can solve the problems existing in the existing knowledge distillation technology at different stages of target detection. These three knowledge distillation methods can be used in combination, or any one or more of them can be selected for use based on actual conditions. They can also be used in combination with existing knowledge distillation technology.
[0078] When at least two knowledge distillation methods are used for model training, in a preferred embodiment, the present invention also provides a joint optimization strategy, which constructs a multi-task optimization framework by weighted summing the knowledge distillation losses corresponding to the selected multiple knowledge distillation methods. For the multi-task optimization framework, the SGD optimizer in the Pytorch framework is preferably used to solve the parameters in the model.
[0079] Taking the selection of three knowledge distillation methods as an example, the loss function of the multi-task optimization framework can be expressed as: In the formula, Represents the weight coefficient of shallow feature knowledge distillation loss, represents the weight coefficient of the deep feature knowledge distillation loss, Represents the weight coefficient of the response knowledge distillation loss. These weight coefficients can be flexibly adjusted according to the different phases of the model to balance the distillation optimization at different stages.
[0080] This embodiment achieves end-to-end optimization by integrating the losses of all feature alignment modules, and balances the importance of different alignment tasks through weight coefficients, thereby ensuring that the student model can still maintain high detection accuracy after being lightweight.
[0081] The present embodiment provides a method for detecting power line defects based on knowledge distillation. Aiming at the problems of lack of shallow feature distillation, insufficient utilization of candidate box information, and limited performance of the detection head of the student model in the traditional method, the present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model, and realizes the multi-level knowledge transfer from the teacher model to the student model by designing shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation. Among them, the shallow feature knowledge distillation optimizes the alignment of the underlying visual features by using the foreground-background regional mask technology, and improves the student model's ability to perceive details; the deep feature knowledge distillation combines the candidate box information with the attention mechanism to enhance the feature expression ability of the key area, and effectively improves the performance of the student model; the response knowledge distillation guides the optimization of the detection head through the soft labels of the classification and regression results, and improves the optimization ability and final detection accuracy of the student model detection head; in addition, the present invention also provides a joint optimization strategy, which constructs a multi-task optimization framework by integrating the loss functions of multi-level feature distillation, candidate box knowledge guidance and response knowledge distillation, and forms a full-process knowledge transfer from features to decisions, thereby improving the overall performance of the student model.
[0082] The method provided by the present invention can effectively avoid neglecting key details due to the simplification of the student model. While reducing the computational complexity of the student model, the student model retains the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on the terminal or edge device, and still has excellent detection effects even in resource-constrained scenarios. The defect detection method provided by the present invention significantly improves the overall efficiency of transmission line defect detection, provides an effective solution for large-scale, real-time, lightweight defect detection applications, and has good practicality and scalability.
[0083] See also Figure 3 Based on the same inventive concept, a second embodiment of the present invention proposes a power transmission line defect detection system based on knowledge distillation, comprising: An image acquisition module 10 is used to acquire an image to be detected of a power transmission line, wherein the image to be detected includes image data and video frame data; A defect detection module 20, used for inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; The student model in the teacher-student training method is the defect detection model, the teacher model is a pre-trained model for defect detection of power transmission lines, the teacher model and the student model have the same network architecture, but different network depths; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation; When shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated based on the regional masking technique and the error measurement method, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions; When deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; When response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
[0084] Furthermore, the defect detection module 20 is also used to perform weighted summation of the distillation losses corresponding to the adopted knowledge distillation methods when at least two of the knowledge distillation methods among shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation are used for model training, to construct a multi-task optimization framework, and to optimize the training of the student model.
[0085] Further, the defect detection module 20 includes a shallow feature knowledge distillation module; The shallow feature knowledge distillation module is used to input the image data in the data set into the teacher model and the student model respectively for shallow feature extraction, and obtain the shallow feature map of the teacher model and the shallow feature map of the student model; According to the set of annotation boxes corresponding to the input image data, the image data is divided into regions, and a regional mask is generated according to the divided regions; According to the regional mask and error measurement method, the shallow feature differences of the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions are calculated respectively; The shallow feature differences in different regions are weighted and summed to obtain the shallow feature knowledge distillation loss.
[0086] Further, the defect detection module 20 includes a deep feature knowledge distillation module; The deep feature knowledge distillation module is used to extract features from the shallow feature maps of the teacher model and the student model respectively, to obtain the deep feature maps of the teacher model and the student model; The teacher model is used to select the target frame of the deep feature map of the teacher model to obtain the candidate frame set of the teacher model; According to the candidate box set, a candidate box guidance region is generated, and the candidate box guidance region is mapped to the spatial dimension of the deep feature map of the student model through linear transformation to generate an attention mask; According to the attention mask, the deep feature map of the teacher model is aligned with the deep feature map of the student model, and the deep feature distillation loss is calculated.
[0087] Further, the defect detection module 20 includes a response knowledge distillation module; A response knowledge distillation module is used to perform classification regression prediction on respective candidate boxes through a teacher model and a student model to obtain a detection result of the teacher model and a detection result of the student model, wherein the detection result of the teacher model includes a classification probability distribution of the teacher model and a regression result of the teacher model, and the detection result of the student model includes a classification probability distribution of the student model and a regression result of the student model; Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on KL divergence; The regression response knowledge loss between the regression results of the teacher model and the regression results of the student model is calculated based on the mean square error; The classification response knowledge loss and the regression response knowledge loss are weighted summed to obtain the response knowledge distillation loss.
[0088] The technical features and technical effects of the power transmission line defect detection system based on knowledge distillation proposed in the embodiment of the present invention are the same as the method proposed in the embodiment of the present invention, and will not be described in detail here. Each module in the above-mentioned power transmission line defect detection system based on knowledge distillation can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0089] In addition, an embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0090] See also Figure 4, an internal structure diagram of a computer device in one embodiment, the computer device can specifically be a terminal or a server. The computer device includes a processor, a memory, a network interface, a display and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a power transmission line defect detection method based on knowledge distillation is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0091] It can be understood by those skilled in the art that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computing device may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.
[0092] In addition, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0093] In summary, an embodiment of the present invention proposes a method and system for power transmission line defect detection based on knowledge distillation. The method obtains an image to be detected of the transmission line, wherein the image to be detected includes image data and video frame data; the image to be detected is input into a pre-constructed defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is constructed based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; wherein the student model in the teacher-student training method is the defect detection model, and the teacher model is a pre-trained model for defect detection of transmission lines, the network architecture of the teacher model and the student model is the same, but the network depth is different; the knowledge distillation includes shallow feature knowledge distillation and deep feature knowledge distillation. and response knowledge distillation; when shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated according to the regional mask technology and the error measurement method, and the shallow feature knowledge distillation loss is calculated according to the shallow feature differences in different regions; when deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; when response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss. The present invention introduces a multi-level feature distillation and knowledge transfer mechanism between the teacher model and the student model by designing shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation, thereby realizing multi-level knowledge migration from the teacher model to the student model, effectively avoiding neglect of key details due to the simplification of the student model, and reducing the computational complexity of the student model while enabling the student model to retain the high detection accuracy of the teacher model to the greatest extent, so that the defect detection model based on the student model can be efficiently deployed on the terminal or edge device, and still has excellent detection effect even in resource-constrained scenarios. The defect detection method provided by the present invention significantly improves the overall efficiency of transmission line defect detection, provides an effective solution for large-scale, real-time, lightweight defect detection applications, and has good practicality and scalability.
[0094] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be directly referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] The above-mentioned embodiments only express several preferred implementation modes of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in the technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be based on the protection scope of the claims.
Claims
1. A power transmission line defect detection method based on knowledge distillation, characterized in that: include: Acquire an image to be detected of a power transmission line, wherein the image to be detected includes image data and video frame data; Inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and trained using a teacher-student training method based on knowledge distillation; The student model in the teacher-student training method is the defect detection model, the teacher model is a pre-trained model for defect detection of power transmission lines, the teacher model and the student model have the same network architecture, but different network depths; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation; When shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated based on the regional masking technique and the error measurement method, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions; When deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; When response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
2. The power transmission line defect detection method based on knowledge distillation according to claim 1 is characterized in that: When at least two of the knowledge distillation methods, shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation, are used for model training, the distillation losses corresponding to the adopted knowledge distillation methods are weighted summed up to construct a multi-task optimization framework and optimize the student model.
3. The power transmission line defect detection method based on knowledge distillation according to claim 1 is characterized in that: The steps of calculating the shallow feature differences between the teacher model and the student model in different regions according to the regional mask technology and the error measurement method, and calculating the shallow feature knowledge distillation loss according to the shallow feature differences in different regions include: The image data in the data set are input into the teacher model and the student model for shallow feature extraction, and the shallow feature map of the teacher model and the shallow feature map of the student model are obtained; According to the set of annotation boxes corresponding to the input image data, the image data is divided into regions, and a regional mask is generated according to the divided regions; According to the regional mask and error measurement method, the shallow feature differences of the shallow feature maps of the teacher model and the shallow feature maps of the student model in different regions are calculated respectively; The shallow feature differences in different regions are weighted and summed to obtain the shallow feature knowledge distillation loss.
4. The power transmission line defect detection method based on knowledge distillation according to claim 3 is characterized in that: The step of dividing the image data into regions according to the set of annotation boxes corresponding to the input image data, and generating a region mask according to the divided regions comprises: The area covered by each annotation box in the annotation box set corresponding to the image data is regarded as the foreground area, and the area outside each annotation box is regarded as the background area; The mask value corresponding to each pixel point in the foreground area is marked as 1, and the mask value corresponding to each pixel point in the background area is marked as 0, so as to obtain a regional mask.
5. The power transmission line defect detection method based on knowledge distillation according to claim 4 is characterized in that: The step of respectively calculating the shallow feature differences of the shallow feature map of the teacher model and the shallow feature map of the student model in different regions according to the regional mask and the error measurement method comprises: According to the regional mask and mean square error, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the foreground area is calculated; According to the inverted value and mean square error of the region mask, the error loss of the shallow feature map of the teacher model and the shallow feature map of the student model in the background area is calculated.
6. The power transmission line defect detection method based on knowledge distillation according to claim 5 is characterized in that: The following formula is used to express the shallow feature knowledge distillation loss: In the formula, represents the region mask at the coordinate point (h, w), Indicates The eigenvalue of the shallow feature map of the student model of the layer at the coordinate point (h, w) of channel c, Indicates The eigenvalue of the shallow feature map of the teacher model of the layer at the coordinate point (h, w) of channel c, represents the total number of pixels in the foreground area, represents the weight loss coefficient of the foreground area, represents the total number of pixels in the background area, Represents the background area weight loss coefficient.
7. The power transmission line defect detection method based on knowledge distillation according to claim 1 is characterized in that: The steps of associating the candidate boxes of the teacher model with the deep features of the student model based on the attention mechanism, generating an attention mask, and calculating the deep feature distillation loss include: By extracting features from the shallow feature maps of the teacher model and the student model, respectively, the deep feature maps of the teacher model and the student model are obtained; The teacher model is used to select the target frame of the deep feature map of the teacher model to obtain the candidate frame set of the teacher model; According to the candidate box set, a candidate box guidance region is generated, and the candidate box guidance region is mapped to the spatial dimension of the deep feature map of the student model through linear transformation to generate an attention mask; According to the attention mask, the deep feature map of the teacher model is aligned with the deep feature map of the student model, and the deep feature distillation loss is calculated.
8. The power transmission line defect detection method based on knowledge distillation according to claim 7 is characterized in that: The step of performing target box selection on the deep feature map of the teacher model through the teacher model to obtain a candidate box set of the teacher model includes: The teacher model is used to select the target frame of the deep feature map of the teacher model, and multiple candidate frames are obtained and formed into a candidate frame set; Calculate the intersection-and-union ratio between each candidate box in the candidate box set and each annotation box in the annotation box set corresponding to the input image data; According to the comparison relationship between the IoU ratio and the IoU ratio threshold, multiple high-quality candidate boxes are screened out from the candidate box set to form a high-quality candidate box set, and the candidate box set is updated according to the high-quality candidate box set.
9. The power transmission line defect detection method based on knowledge distillation according to claim 7, characterized in that: The step of generating a candidate box guidance region according to the candidate box set, and mapping the candidate box guidance region to the spatial dimension of the deep feature map of the student model through linear transformation, and generating an attention mask includes: Merge each candidate box in the candidate box set to generate a candidate box guidance area; Through the weight matrix, the candidate box guidance area is mapped to the query matrix, and the deep feature map of the student model is mapped to the key matrix and the value matrix; According to the attention mechanism, the query matrix and the key matrix are dot-producted and normalized to obtain the attention mask.
10. The power transmission line defect detection method based on knowledge distillation according to claim 7, characterized in that: The step of aligning the deep feature map of the teacher model and the deep feature map of the student model according to the attention mask and calculating the deep feature distillation loss includes: Calculate the deep feature error loss between the deep feature map of the teacher model and the deep feature map of the student model based on the mean squared error, and calculate the inner product between the attention mask and the deep feature error loss; According to the number of deep feature layers of the student model and the inner product, the deep feature distillation loss is calculated.
11. The power transmission line defect detection method based on knowledge distillation according to claim 10, characterized in that: The deep feature knowledge distillation loss is expressed by the following formula: In the formula, represents the deep feature error loss, represents the attention mask of the deep features of the rth layer, represents the inner product, and T represents the number of deep feature layers.
12. The power transmission line defect detection method based on knowledge distillation according to claim 7, characterized in that: Before the step of selecting a target frame on the deep feature map of the teacher model through the teacher model to obtain a candidate frame set of the teacher model, the step further includes: The deep feature map of the student model is aligned with the deep feature map of the teacher model, and the alignment operation includes size alignment and resolution alignment.
13. The power transmission line defect detection method based on knowledge distillation according to claim 1, characterized in that: The step of using the detection result of the teacher model as a soft label, calculating the classification response knowledge loss and the regression response knowledge loss between the teacher model and the student model based on the error measurement method, and obtaining the response knowledge distillation loss includes: Through the teacher model and the student model, classification regression prediction is performed on the respective candidate boxes to obtain the detection result of the teacher model and the detection result of the student model, wherein the detection result of the teacher model includes the classification probability distribution of the teacher model and the regression result of the teacher model, and the detection result of the student model includes the classification probability distribution of the student model and the regression result of the student model; Calculate the classification response knowledge loss between the classification probability distribution of the teacher model and the classification probability distribution of the student model based on KL divergence; The regression response knowledge loss between the regression results of the teacher model and the regression results of the student model is calculated based on the mean square error; The classification response knowledge loss and the regression response knowledge loss are weighted summed to obtain the response knowledge distillation loss.
14. The power transmission line defect detection method based on knowledge distillation according to claim 13, characterized in that: The step of performing classification regression prediction on respective candidate boxes through the teacher model and the student model to obtain the detection result of the teacher model and the detection result of the student model includes: The candidate boxes of the teacher model are used as the candidate boxes of the student model, and the teacher model and the student model are used to perform classification regression prediction on their respective candidate boxes to obtain the detection results of the teacher model and the detection results of the student model.
15. The power transmission line defect detection method based on knowledge distillation according to claim 13, characterized in that: The response knowledge distillation loss is expressed as follows: Where, L kl represents the KL divergence, represents the i-th candidate box, N represents the total number of candidate boxes, represents the classification probability distribution of the i-th candidate box in the teacher model, represents the classification probability distribution of the i-th candidate box in the student model, represents the regression result of the i-th candidate box in the teacher model, represents the regression result of the i-th candidate box in the student model, represents the weight coefficient of the knowledge loss of the classification response, The weight coefficient representing the knowledge loss of the regression response.
16. The power transmission line defect detection method based on knowledge distillation according to claim 1, characterized in that: Both the teacher model and the student model are constructed using fast convolutional neural networks. The teacher model uses ResNet-101 as the backbone network, and the student model uses ResNet-50 as the backbone network.
17. The power transmission line defect detection method based on knowledge distillation according to claim 16, characterized in that: The network architecture of both the teacher model and the student model includes a shallow feature extraction network, a deep feature extraction network, a region candidate network, and a detection head; Among them, the shallow feature extraction network is the first three stages of the backbone network, which is used to extract shallow features of the input image data; The deep feature extraction network is the last two stages of the backbone network, which is used to perform deep feature extraction on the shallow features output by the shallow feature extraction network; The region candidate network is used to select the target frame based on the deep features output by the deep feature extraction network and generate candidate frames; The detection head is used to perform classification regression prediction on the candidate boxes output by the region candidate network and output the detection results.
18. The power transmission line defect detection method based on knowledge distillation according to claim 2, characterized in that: The SGD optimizer is used to solve the multi-task optimization framework to obtain the optimal parameters of the student model.
19. A power transmission line defect detection system based on knowledge distillation, characterized in that: include: An image acquisition module, used to acquire an image to be detected of the power transmission line, wherein the image to be detected includes image data and video frame data; A defect detection module, used for inputting the image to be detected into a pre-built defect detection model to obtain a defect detection result of the transmission line, wherein the defect detection model is built based on a convolutional neural network and is trained using a teacher-student training method based on knowledge distillation; The student model in the teacher-student training method is the defect detection model, the teacher model is a pre-trained model for defect detection of power transmission lines, the teacher model and the student model have the same network architecture, but different network depths; The knowledge distillation includes any one or more of shallow feature knowledge distillation, deep feature knowledge distillation and response knowledge distillation; When shallow feature knowledge distillation is used for model training, the shallow feature differences between the teacher model and the student model in different regions are calculated based on the regional masking technique and the error measurement method, and the shallow feature knowledge distillation loss is calculated based on the shallow feature differences in different regions; When deep feature knowledge distillation is used for model training, the candidate boxes of the teacher model are associated with the deep features of the student model based on the attention mechanism, an attention mask is generated, and the deep feature distillation loss is calculated; When response knowledge distillation is used for model training, the detection results of the teacher model are used as soft labels, and the classification response knowledge loss and regression response knowledge loss between the teacher model and the student model are calculated based on the error measurement method to obtain the response knowledge distillation loss.
Citation Information
Patent Citations
SAR (Synthetic Aperture Radar) image target detection method combined with high-credibility knowledge distillation
CN115761511A
High-efficiency and high-precision power transmission line defect identification method and system
CN116883647A
Knowledge distillation-based lightweight power transmission line fault detection method
CN117746083A
Steel plate surface defect detection method and system based on knowledge distillation
CN118096768A
Target detection knowledge distillation method based on local attention and context relationship
CN118506149A
Cited By
Remote sensing target detection method and device based on multi-level knowledge distillation
CN120997489A
A remote sensing target detection method and device based on multi-level knowledge distillation
CN120997489B