Power operation violation identification method and system based on transfer learning and multi-teacher distillation

By adopting a lightweight violation identification method based on transfer learning and multi-teacher distillation in the power system, combined with mask-reconstruction and contrast learning pre-training models, the problem of violation operation identification in the power system is solved, and efficient and accurate violation detection is achieved.

CN119723430BActive Publication Date: 2025-05-06PING YANG XIAN CHANG TAI DIAN LI SHI YE YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510232727.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-06
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify employee violations in power systems, especially models running on edge devices that are difficult to balance global information and local details, resulting in a decrease in recognition accuracy.

Method used

The lightweight violation recognition method based on transfer learning and multi-teacher distillation is adopted, combined with mask-reconstruction and contrast learning pre-trained models, and the global and local feature recognition advantages of large models are transferred to the small models through multi-teacher distillation strategy.

Benefits of technology

It significantly improves the recognition accuracy of the model, ensures the efficient deployment of the model on the edge of the power system equipment, and meets the environmental needs of limited computing resources and storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723430B_ABST
    Figure CN119723430B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for identifying illegal power operation based on transfer learning and multi-teacher distillation, and belongs to the field of computer vision technology. It includes: by combining a mask-reconstruction pre-training model and a contrastive learning pre-training model, as well as a multi-teacher distillation strategy, the advantages of a large model in global and local feature recognition are transferred to a small model. This method not only significantly improves the recognition accuracy of the model, but also ensures that the model can be efficiently deployed on edge devices of the power system, meeting the requirements of environments with limited computing resources and storage space. Therefore, the present invention enhances the detection efficiency of illegal operations while maintaining the lightweight of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for identifying illegal power operation based on transfer learning and multi-teacher distillation, and belongs to the technical field of computer vision. Background Art

[0002] In today's era, the process of intelligentization and digitalization of power systems is advancing at an unprecedented speed, which has a profound impact on global energy management and distribution. With the advancement of technology, power systems have become more complex, involving more extensive equipment and networks, which has not only brought about efficiency improvements, but also brought new challenges. Among them, how to identify violations in power operations through automation technology has become a key issue to ensure the safety of power grid operations and improve operational efficiency. The technology needs to be able to detect and respond to various violations in a timely manner to effectively reduce the potential risks faced by the power system.

[0003] In the wave of intelligent and digital power systems, the problem of employees' illegal operations has become increasingly prominent. These illegal operations include operating errors and violations of operating procedures, such as incorrect operation of equipment switches, failure to perform maintenance procedures according to specifications, etc., and may also involve unauthorized access to key equipment, exceeding authority to enter restricted areas and other violations. These behaviors may not only cause damage to physical facilities, such as equipment failure or power outages, but may also trigger system alarms through operating errors, and even indirectly trigger a larger chain reaction. In addition, unauthorized access to equipment terminals may pose a serious threat to the stability and security of the power system. Therefore, for the power system, identifying employees' illegal operations is not only a need to improve operating efficiency and reduce the risk of equipment damage, but also a key link to ensure the safety of the power grid and maintain its stable operation.

[0004] However, in practical applications, edge devices in power systems (such as smart terminals and surveillance cameras) are generally limited in computing power and storage capacity, making it difficult to support the operation of complex large models. Therefore, it is necessary to develop lightweight and efficient models to meet the deployment requirements of edge devices. At the same time, violations in power system scenarios are usually accompanied by background interference, diversity of equipment morphology, and complexity of texture features, which places higher demands on the model's ability to extract global and local features. Existing methods often find it difficult to strike a balance between global information and local details, which affects the performance of the model in complex scenarios.

[0005] The challenges of current research on the identification of illegal operations in power systems are as follows:

[0006] First, the difficulty of knowledge fusion in multi-teacher distillation. In multi-teacher distillation technology, it is a major difficulty to effectively fuse the knowledge of multiple pre-trained models (teacher models) and transfer them to a lightweight student model. Different teacher models have significant differences in feature expression: the pre-trained model obtained by the "mask-reconstruction" method pays more attention to local relationships and detailed textures, while the pre-trained model obtained by the contrastive learning method focuses on global features and shape information. How to design a distillation strategy so that the student model can simultaneously learn and integrate these two complementary knowledge features without causing performance degradation due to conflicting information is a problem that existing technologies need to solve.

[0007] Second, the limitation of model generalization ability in transfer learning. Transfer learning relies on models pre-trained on large-scale datasets, but when the data distribution of downstream tasks differs from the pre-training data, the migration effect of the model may be significantly affected. Especially in the task of identifying power violations, the amount of data is limited and the scene is highly specific, which poses a challenge to the generalization ability of the migration model. In addition, the "mask-reconstruction" method and the contrastive learning method have different feature biases in the pre-training stage, which may cause the model to be unbalanced in the expression of local or global features when migrating to new tasks.

[0008] Third, the lightweight model lacks knowledge acquisition. In knowledge distillation, the lightweight student model has difficulty in fully learning the complex feature representation and structural information in the teacher model due to capacity limitations. In particular, when the teacher model contains multiple layers of deep representation, it is difficult for the student model to capture this multi-level, cross-scale information. In addition, current distillation methods often only focus on class tokens or global features, while ignoring the distillation of relationships between tokens, which further limits the lightweight model's ability to represent complex violation scenarios. How to enable the student model to approach the performance of the teacher model through a more efficient distillation strategy is still a major bottleneck in existing technologies. Summary of the invention

[0009] In view of the shortcomings of the existing technology, the present invention proposes a method for identifying power operation violations based on transfer learning and multi-teacher distillation.

[0010] In response to the difficulty in identifying illegal operations in the power system, the present invention innovatively proposes a lightweight violation identification method that integrates transfer learning and multi-teacher distillation technology. The present invention focuses on addressing the performance degradation problem that may be caused by the model simplification process. Its core innovation is that by combining the mask-reconstruction pre-training model and the contrastive learning pre-training model, as well as the multi-teacher distillation strategy, the advantages of large models in global and local feature recognition are transferred to small models. This method not only significantly improves the recognition accuracy of the model, but also ensures that the model can be efficiently deployed on the edge devices of the power system, meeting the environmental requirements with limited computing resources and storage space. Therefore, while maintaining the lightweight of the model, the present invention enhances the detection efficiency of illegal operations, especially when the operating environment of the power system is complex and changeable.

[0011] The present invention constructs a lightweight power violation identification model based on transfer learning and multi-teacher distillation, which is used to efficiently identify operational violations in power systems. The present invention proposes a multi-teacher distillation method that comprehensively utilizes the "mask-reconstruction" and contrastive learning pre-training models, and provides rich and diverse feature knowledge for lightweight models by integrating the advantages of the two methods. On this basis, the present invention designs a distillation strategy that adapts to power violation scenarios, focusing on different feature layers of the teacher model, especially the early layers of the "mask-reconstruction" pre-training model and the later layers of the contrastive learning pre-training model, and effectively distills the relationship between tokens, thereby enhancing the student model's ability to recognize illegal operations.

[0012] The present invention is aimed at the task of identifying illegal behaviors in power operation videos, obtains local and global information in the pre-trained model, and then distills this information into a lightweight model, significantly reducing the demand for computing resources. For example, in surveillance videos, it is possible to efficiently identify illegal operation scenarios, such as failure to operate equipment according to regulations or entering restricted areas without authorization. On the one hand, the present invention is an extension of transfer learning technology, making full use of the complementary characteristics of "mask-reconstruction" and contrastive learning; on the other hand, by combining multi-teacher distillation and lightweight models, the efficiency and accuracy of identifying illegal power behaviors are improved.

[0013] The present invention also proposes a power operation violation identification system based on transfer learning and multi-teacher distillation, which is suitable for resource-constrained scenarios and provides strong protection for power grid security.

[0014] The technical solution of the present invention is as follows:

[0015] The power operation violation identification method based on transfer learning and multi-teacher distillation includes:

[0016] Training lightweight models, including:

[0017] Resize the given input image and perform image block operations;

[0018] The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder;

[0019] Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model;

[0020] The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image;

[0021] Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping;

[0022] During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model;

[0023] Align the attention maps of the student model and the first teacher model and calculate the local alignment loss;

[0024] Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss;

[0025] Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss;

[0026] The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model;

[0027] Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; violation identification, including:

[0028] For the input image, the trained lightweight model is used to analyze the behavioral characteristics of each operation target to determine whether there is a violation.

[0029] Preferably, according to the present invention, the first teacher model is based on the MAE method, adopts the ViT-Base model, randomly masks the input image and extracts local features through the encoder; comprising:

[0030] Name the encoder of the first teacher model , name the decoder of the first teacher model ;

[0031] Input Image is randomly masked to generate a masked image ; The masked image is defined as:

[0032] ;

[0033] in, represents element-wise multiplication, is a randomly generated mask matrix;

[0034] The image that will be masked The part with 0 in the image is discarded to get the processed image. ;

[0035] The processed image Input to the MAE encoder In , extract local features to represent:

[0036] ;

[0037] in, Represents the local feature representation of the image.

[0038] Preferably, according to the present invention, the attention mechanism is used to extract attention maps of different levels, and the attention maps are mapped through MLP to align with the student model; including:

[0039] The local features are represented The masked position is inserted with a learnable mask marker, and new features are formed by random insertion. ;

[0040] Through the decoder Get the reconstructed image features :

[0041] ;

[0042] For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the first teacher model for the image ;

[0043] Extract the attention in the 2nd and 6th layer Transformer blocks in the first teacher model, named and ;

[0044] Using MLP and Mapping and , which is consistent with that in the student model, as follows:

[0045] .

[0046] According to the preferred embodiment of the present invention, the second teacher model is based on the CLIP method and adopts the clip-vit-base-patch16 model to directly extract the global features of the complete image; the encoder of the second teacher model is named ;include:

[0047] Input Image To encoder In this paper, we extract global features :

[0048] .

[0049] Preferably, according to the present invention, attention maps at different levels are obtained by combining the attention mechanism, and aligned with the student model after MLP mapping;

[0050] For global features , average pooling is performed on all image blocks to obtain the expression characteristics of the second teacher model for the image ;

[0051] Extract the attention in the 10th and 12th layer Transformer blocks in the second teacher model, named and ; Use MLP to map attention so that and Same as in the student model:

[0052] .

[0053] Preferably, according to the present invention, in the training process of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model; comprising:

[0054] The encoder of the student model adopts the ViT-S model, named , the decoder of the student model adopts the ViT-T model, named ;

[0055] The processed image Input to the encoder of the student model , and get the local feature representation:

[0056] ;

[0057] in, Represents the first 6 layers of the encoder.

[0058] Preferably, according to the present invention, the attention maps of the student model and the first teacher model are aligned, and the first local alignment loss is calculated; including:

[0059] Extract the attention in the 2nd and 6th layer Transformer blocks in the student model and name it and ;

[0060] Align the attention map of the student model with the attention map of the first teacher model, and define the first local alignment loss as:

[0061] .

[0062] Preferably, according to the present invention, the image features are reconstructed by the decoder of the student model, and the reconstruction loss is calculated; comprising:

[0063] Will The masked position is inserted with a learnable mask marker, and new features are formed by random insertion. ;

[0064] Through the decoder Get the reconstructed image features :

[0065] ;

[0066] The reconstructed image features and images Calculate the reconstruction loss :

[0067] .

[0068] Preferably, according to the present invention, extracting global features and aligning them with the global features of the second teacher model, calculating the global alignment loss; optimizing the student model by the InfoNCE loss function so that its feature expression is consistent with the teacher model; comprising:

[0069] For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image ;

[0070] For the same image in a batch , the expression characteristics of the student model for the image And the first teacher model's expression characteristics of the image Each other is a positive sample, and the other pictures are negative samples. The first Info NCE loss function is calculated as :

[0071] ;

[0072] in, is the batch size, Represents the cosine similarity between the two;

[0073] The reconstructed image features Input to the encoder of the student model to obtain the global feature representation :

[0074] ;

[0075] in, Represents the last 6 layers of the encoder;

[0076] Extract the attention maps of the 10th and 12th layers of the student model and name them as and ;

[0077] Will and Align with the attention map of the second teacher model and define the second local alignment loss as:

[0078] ;

[0079] For the global features of the image , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image .

[0080] For the same image in a batch , the expression characteristics of the student model for the image The expression characteristics of the image by the second teacher model Each other is a positive sample, and the other pictures are negative samples. Calculate the second Info NCE loss function :

[0081] ;

[0082] Designing a joint loss function , the first local alignment loss, the second local alignment loss, the first Info NCE loss function, the second Info NCE loss function and the reconstruction loss are comprehensively optimized:

[0083] ;

[0084] in, , , is a weight hyperparameter used to balance the losses of each part.

[0085] Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; complete the construction of the lightweight model and improve its performance; including:

[0086] In the power operation violation identification dataset, the encoder of the student model obtained by distillation is Process the image to get the overall expression of the image ;

[0087] The overall expression of the image After MLP, the feature size is obtained :

[0088] ;

[0089] Use YOLO v10 and name YOLO v10 Get the predicted bounding box :

[0090] ;

[0091] in, Indicates the confidence of whether the predicted bounding box contains the target, and represents the coordinates of the upper left corner of the bounding box, and The coordinates of the lower right corner of the bounding box;

[0092] Compute predicted and ground-truth bounding boxes The degree of match , calculated as:

[0093] ;

[0094] in, , , Indicates the area of ​​a region;

[0095] Calculate the confidence that the predicted bounding box contains the object There is a label with the real target Degree of match:

[0096] ;

[0097] For all images in a batch, calculate the loss function :

[0098] ;

[0099] in, is the batch size, and refers to the loss function of each sample, and is the loss weight hyperparameter.

[0100] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a method for identifying power operation violations based on transfer learning and multi-teacher distillation are implemented.

[0101] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for identifying power operation violations based on transfer learning and multi-teacher distillation.

[0102] The power operation violation identification system based on transfer learning and multi-teacher distillation includes:

[0103] The lightweight model training module is configured as follows:

[0104] Resize the given input image and perform image block operations;

[0105] The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder;

[0106] Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model;

[0107] The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image;

[0108] Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping;

[0109] During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model;

[0110] Align the attention maps of the student model and the first teacher model and calculate the local alignment loss;

[0111] Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss;

[0112] Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss;

[0113] The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model;

[0114] Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value;

[0115] The violation identification module is configured to: for the input image, use the trained lightweight model to analyze the behavioral characteristics of each operation target and determine whether it is a violation.

[0116] Compared with the prior art, the present invention has the following beneficial effects:

[0117] 0. This invention introduces multi-teacher distillation technology, fully integrating the advantages of the two pre-training models of "mask-reconstruction" and contrastive learning, capturing local detail information while taking into account global structural features, so that the accuracy of lightweight models in power violation identification tasks is effectively improved. At the same time, an adaptive distillation strategy is designed for the local attention mechanism of "mask-reconstruction" and the global feature expression of contrastive learning to enhance the model's ability to identify complex violations.

[0118] 1. Through multi-teacher distillation technology, the present invention effectively transfers the knowledge of the "mask-reconstruction" and contrastive learning teacher models with excellent performance but high resource requirements to the lightweight student model, reducing the computing resource requirements for model operation.

[0119] 2. The present invention comprehensively considers local features (equipment operation details) and global features (scene behavior associations) in monitoring videos, making the model more adaptable and robust to a variety of violations, which helps to ensure the safety of power grid operation and improve operational efficiency.

[0120] 4. The present invention can effectively identify illegal power operation violations in images and correctly locate the specific area of ​​the image where the violations are located. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 This is a flowchart of the method for identifying power operation violations based on transfer learning and multi-teacher distillation of the present invention. DETAILED DESCRIPTION

[0122] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0123] Example 1

[0124] Based on transfer learning and multi-teacher distillation, the power operation violation identification method, such as Figure 1 As shown, including:

[0125] Training lightweight models, including:

[0126] Resize the given input image and divide it into image blocks to make it meet the model input size requirements;

[0127] The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder; MAE stands for Masked Autoencoders, also known as masked self-encoders. It is a self-supervised learning method that aims to encode these image blocks by randomly masking part of the image blocks of the input image and inputting the unmasked image blocks into the encoder composed of Vision Transformer. The information output by the encoder will be sent to the decoder composed of Vision Transformer, and the decoder uses this encoded information to reconstruct the entire original image, including those masked parts. The ViT-Base model is called Vision Transformer-Small. It is a lightweight version of Vision Transformer, consisting of 12 layers of Vision Transformer blocks stacked together, with a feature dimension of 384.

[0128] The attention mechanism is used to extract attention maps at different levels, and the attention maps are mapped through MLP to align with the student model; MLP stands for Multilayer Perceptron. It is a feedforward artificial network model, including an input layer, a hidden layer, and an output layer. Each layer consists of multiple neurons, and the neurons are connected by weights.

[0129] The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image;

[0130] Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping;

[0131] During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model;

[0132] Align the attention maps of the student model and the first teacher model and calculate the local alignment loss;

[0133] Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss;

[0134] Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss;

[0135] The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model;

[0136] Combined with the YOLO detection model for prediction, the matching degree between the bounding box of the prediction result and the true value is calculated; the lightweight model is constructed and the performance is improved.

[0137] Violation identification, including:

[0138] For the input image, the trained lightweight model is used to analyze the behavioral characteristics of each operation target to determine whether there is a violation.

[0139] Example 2

[0140] The difference between the method for identifying illegal power operation based on transfer learning and multi-teacher distillation described in Example 1 is that:

[0141] Resize the given input image; including:

[0142] Get the input image After that, first, the image is resized to meet the model input size requirements, that is, , , . Represents the height of the image. Represents the width of the image, represents the number of color channels of the image. Then, the image is divided into blocks to Divide the image into blocks and obtain the image after dividing the image blocks ,in, is the number of image blocks per row and column, that is .

[0143] The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder; including:

[0144] In this embodiment, the Vision Transformer-Base (ViT-B, ViT-Base model) in Masked Autoencoders (MAE) released by facebookresearch is used as the first teacher model, and the encoder of the first teacher model is named , name the decoder of the first teacher model ;

[0145] Input Image 75% of the area is randomly masked to generate a masked image ; The masked image is defined as:

[0146] ;

[0147] in, represents element-wise multiplication, is a randomly generated mask matrix; The proportion of 0 is 75%.

[0148] The image that will be masked The part with 0 in the image is discarded to get the processed image. ;

[0149] The processed image Input to the MAE encoder In , extract local features to represent:

[0150] ;

[0151] in, Represents the local feature representation of the image.

[0152] Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model; including:

[0153] The local features are represented The masked positions are inserted with learnable mask markers, which are random features that will be used to reconstruct the masked parts in the decoder. ;

[0154] Through the decoder Get the reconstructed image features :

[0155] ;

[0156] For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the first teacher model for the image ;

[0157] Extract the attention in the 2nd and 6th layer Transformer blocks in the first teacher model, named and ;

[0158] Using MLP (Multilayer Perceptron) and Mapping and , which is consistent with that in the student model, as follows:

[0159] .

[0160] The second teacher model is based on the CLIP method and adopts the clip-vit-base-patch16 model to directly extract the global features of the complete image; the encoder of the second teacher model is named ;include:

[0161] In this embodiment, clip-vit-base-patch16 in Contrastive Language-Image Pre-Training (CLIP) released by openai is used as the second teacher model;

[0162] Input Image To encoder In this paper, we extract global features :

[0163] .

[0164] Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping;

[0165] For global features , average pooling is performed on all image blocks to obtain the expression characteristics of the second teacher model for the image ;

[0166] Extract the attention in the 10th and 12th layer Transformer blocks in the second teacher model, named and ; Use MLP to map attention so that and Same as in the student model:

[0167] .

[0168] During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model; including:

[0169] The encoder of the student model adopts the ViT-S model, named The decoder of the student model uses the ViT-T model (Vision Transformer-Tiny), named ; ViT-T model: The full name is VisionTransformer-Tiny, which is a lightweight version of Vision Transformer. It is composed of 12 layers of VisionTransformer blocks and has a feature dimension of 192.

[0170] The processed image Input to the encoder of the student model , and get the local feature representation:

[0171] ;

[0172] in, Represents the first 6 layers of the encoder.

[0173] Align the attention maps of the student model and the first teacher model, and calculate the first local alignment loss; including:

[0174] Extract the attention in the 2nd and 6th layer Transformer blocks in the student model and name it and ;

[0175] Align the attention map of the student model with the attention map of the first teacher model, and define the first local alignment loss as:

[0176] .

[0177] Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss; including:

[0178] Will The masked positions are inserted with learnable mask markers, which are random features that will be used to reconstruct the masked parts in the decoder. ;

[0179] Through the decoder Get the reconstructed image features :

[0180] ;

[0181] The reconstructed image features and images Calculate the reconstruction loss :

[0182] .

[0183] Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss; optimize the student model through the InfoNCE loss function to make its feature expression consistent with the teacher model; including:

[0184] For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image ;

[0185] For the same image in a batch , the expression characteristics of the student model for the image And the first teacher model's expression characteristics of the image Each other is a positive sample, and the other pictures are negative samples. The first Info NCE loss function is calculated as :

[0186] ;

[0187] in, is the batch size, Represents the cosine similarity between the two;

[0188] The reconstructed image features Input to the encoder of the student model to obtain the global feature representation :

[0189] ;

[0190] in, Represents the last 6 layers of the encoder;

[0191] Extract the attention maps of the 10th and 12th layers of the student model and name them as and ;

[0192] Will and Align with the attention map of the second teacher model and define the second local alignment loss as:

[0193] ;

[0194] For the global features of the image , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image .

[0195] For the same image in a batch , the expression characteristics of the student model for the image The expression characteristics of the image by the second teacher model Each other is a positive sample, and the other pictures are negative samples. Calculate the second Info NCE loss function :

[0196] ;

[0197] Designing a joint loss function , the first local alignment loss, the second local alignment loss, the first Info NCE loss function, the second Info NCE loss function and the reconstruction loss are comprehensively optimized:

[0198] ;

[0199] in, , , is a weight hyperparameter used to balance the losses of each part.

[0200] Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; complete the construction of the lightweight model and improve its performance; including:

[0201] In this embodiment, in the power operation violation identification dataset, the encoder of the student model obtained by distillation is Process the image to get the overall expression of the image ;

[0202] The overall expression of the image After MLP, we get the same feature size as YOLO v10 :

[0203] ;

[0204] Use YOLO v10 released by ultralytics and name YOLO v10 Get the predicted bounding box :

[0205] ;

[0206] in, Indicates the confidence of whether the predicted bounding box contains the target, and represents the coordinates of the upper left corner of the bounding box, and The coordinates of the lower right corner of the bounding box;

[0207] Compute predicted and ground-truth bounding boxes The degree of match , calculated as:

[0208] ;

[0209] in, , , Indicates the area of ​​a region;

[0210] Calculate the confidence that the predicted bounding box contains the object There is a label with the real target Degree of match:

[0211] ;

[0212] For all images in a batch, calculate the loss function :

[0213] ;

[0214] in, is the batch size, and refers to the loss function of each sample, and is the loss weight hyperparameter.

[0215] Joint loss function It refers to the loss function used in the process of multi-teacher distillation to effectively transfer the knowledge of multiple teacher models to the student model. Refers to the loss function used to fine-tune the model during the transfer learning process in order to improve the model's ability to identify violations of power operation.

[0216] In this embodiment, in the lightweight model, the input is: the original image; the output is: image features, attention in the 2nd and 6th layer Transformer blocks, and attention in the 10th and 12th layer Transformer blocks;

[0217] In the first teacher model, the input is: the original image; the output is: image features, attention in the 2nd and 6th layer Transformer blocks;

[0218] In the second teacher model, the input is: the original image; the output is: image features, attention in the 10th and 12th layer Transformer blocks;

[0219] In the YOLO detection model, the input is: original image; the output is: confidence, predicted bounding box.

[0220] Example 3

[0221] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for identifying power operation violations based on transfer learning and multi-teacher distillation described in Example 1 or 2 are implemented.

[0222] Example 4

[0223] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the method for identifying power operation violations based on transfer learning and multi-teacher distillation described in Example 1 or 2 are implemented.

[0224] Example 5

[0225] The power operation violation identification system based on transfer learning and multi-teacher distillation includes:

[0226] The lightweight model training module is configured as follows:

[0227] Resize the given input image and divide it into image blocks to make it meet the model input size requirements;

[0228] The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder;

[0229] Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model;

[0230] The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image;

[0231] Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping;

[0232] During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model;

[0233] Align the attention maps of the student model and the first teacher model and calculate the local alignment loss;

[0234] Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss;

[0235] Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss;

[0236] The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model;

[0237] Combined with the YOLO detection model for prediction, the matching degree between the bounding box of the prediction result and the true value is calculated; the lightweight model is constructed and the performance is improved.

[0238] The violation identification module is configured to: for the input image, use the trained lightweight model to analyze the behavioral characteristics of each operation target and determine whether it is a violation.

Claims

1. A method for identifying power operation violations based on transfer learning and multi-teacher distillation, characterized in that: include: Training lightweight models, including: Resize the given input image and perform image block operations; The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder; Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model; The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image; Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping; During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model; Align the attention maps of the student model and the first teacher model and calculate the local alignment loss; Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss; Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss; The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model; Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; violation identification, including: For the input image, the trained lightweight model is used to analyze the behavioral characteristics of each operation target to determine whether there is a violation.

2. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 1 is characterized in that: The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder; including: Name the encoder of the first teacher model , and name the decoder of the first teacher model ; Input Image is randomly masked to generate a masked image ; The masked image is defined as: ; in, represents element-wise multiplication, is a randomly generated mask matrix; The image that will be masked The part with 0 in the image is discarded to get the processed image. ; The processed image Input to the MAE encoder In , extract local features to represent: ; in, Represents the local feature representation of the image.

3. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 2 is characterized in that: Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model; include: The local features are represented The masked position is inserted with a learnable mask marker, and new features are formed by random insertion. ; Through the decoder Get the reconstructed image features : ; For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the first teacher model for the image ; Extract the attention in the 2nd and 6th layer Transformer blocks in the first teacher model, named and ; Using MLP and Mapping and , which is consistent with that in the student model, as follows: 。 4. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 3 is characterized in that: The second teacher model is based on the CLIP method and adopts the clip-vit-base-patch16 model to directly extract the global features of the complete image; the encoder of the second teacher model is named ;include: Input Image To encoder In this paper, we extract global features : 。 5. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 4 is characterized in that: Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping; For global features , average pooling is performed on all image blocks to obtain the expression characteristics of the second teacher model for the image ; Extract the attention in the 10th and 12th layer Transformer blocks in the second teacher model, named and ; Use MLP to map attention so that and Same as in the student model: 。 6. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 5 is characterized in that: During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model; including: The encoder of the student model adopts the ViT-S model, named , the decoder of the student model adopts the ViT-T model, named ; The processed image Input to the encoder of the student model , and get the local feature representation: ; in, Represents the first 6 layers of the encoder.

7. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 6 is characterized in that: Align the attention maps of the student model and the first teacher model, and calculate the first local alignment loss; including: Extract the attention in the 2nd and 6th layer Transformer blocks in the student model and name it and ; Align the attention map of the student model with the attention map of the first teacher model, and define the first local alignment loss as: 。 8. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 7 is characterized in that: Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss; include: Will The masked position is inserted with a learnable mask marker, and new features are formed by random insertion. ; Through the decoder Get the reconstructed image features : ; The reconstructed image features and images Calculate the reconstruction loss : 。 9. The method for identifying power operation violations based on transfer learning and multi-teacher distillation according to claim 8 is characterized in that: Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss; optimize the student model through the InfoNCE loss function to make its feature expression consistent with the teacher model; including: For the reconstructed image features , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image ; For the same image in a batch , the expression characteristics of the student model for the image And the first teacher model's expression characteristics of the image Each other is a positive sample, and the other pictures are negative samples. The first Info NCE loss function is calculated as : ; in, is the batch size, Represents the cosine similarity between the two; The reconstructed image features Input to the encoder of the student model to obtain the global feature representation : ; in, Represents the last 6 layers of the encoder; Extract the attention maps of the 10th and 12th layers of the student model and name them as and ; Will and Align with the attention map of the second teacher model and define the second local alignment loss as: ; For the global features of the image , average pooling is performed on all image blocks to obtain the expression characteristics of the student model for the image ; For the same image in a batch , the expression characteristics of the student model for the image And the second teacher model’s expression characteristics of the image Each other is a positive sample, and the other pictures are negative samples. Calculate the second Info NCE loss function : ; Designing a joint loss function , the first local alignment loss, the second local alignment loss, the first Info NCE loss function, the second Info NCE loss function and the reconstruction loss are comprehensively optimized: ; in, , , is a weight hyperparameter used to balance the losses of each part; Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; complete the construction of the lightweight model and improve its performance; including: In the power operation violation identification dataset, the encoder of the student model obtained by distillation is Process the image to get the overall expression of the image ; The overall expression of the image After MLP, the feature size is obtained : ; Use YOLO v10 and name YOLO v10 Get the predicted bounding box : ; in, Indicates the confidence of whether the predicted bounding box contains the target, and Represents the coordinates of the upper left corner of the bounding box, and The coordinates of the lower right corner of the bounding box; Compute predicted and ground-truth bounding boxes The degree of match , calculated as: ; in, , , Indicates the area of ​​a region; Calculate the confidence that the predicted bounding box contains the object There is a label with the real target Degree of match: ; For all images in a batch, calculate the loss function : ; in, is the batch size, and refers to the loss function of each sample, and is the loss weight hyperparameter.

10. A power operation violation identification system based on transfer learning and multi-teacher distillation, characterized by: include: The lightweight model training module is configured as follows: Resize the given input image and perform image block operations; The first teacher model is based on the MAE method and adopts the ViT-Base model to randomly mask the input image and extract local features through the encoder; Use the attention mechanism to extract attention maps at different levels, map the attention maps through MLP, and align them with the student model; The second teacher model is based on the CLIP method and uses the clip-vit-base-patch16 model to directly extract the global features of the complete image; Combine the attention mechanism to obtain attention maps at different levels, and align them with the student model after MLP mapping; During the training of the student model, the mask image of the first teacher model is used to extract local features through the encoder of the student model; Align the attention maps of the student model and the first teacher model and calculate the local alignment loss; Reconstruct the image features through the decoder of the student model and calculate the reconstruction loss; Extract global features and align them with the global features of the second teacher model, and calculate the global alignment loss; The student model is optimized through the InfoNCE loss function to make its feature expression consistent with the teacher model; Combined with the YOLO detection model for prediction, calculate the matching degree between the bounding box of the prediction result and the true value; The violation identification module is configured to: for the input image, use the trained lightweight model to analyze the behavioral characteristics of each operation target and determine whether it is a violation.

Citation Information

Patent Citations

  • Transform-based homestead remote sensing image illegal building identification method and system

    CN114359702A

  • Method and system for detecting illegal carrying articles of mine aerial passenger device personnel

    CN117152419A