Incremental target safety detection method for electric power operation scene

By adopting incremental target safety detection methods in power operation scenarios, using knowledge distillation and improved network structure, the shortcomings of traditional models in identifying old goals and adapting to new goals are solved, and efficient knowledge transfer and detection performance improvements are achieved.

CN120047662APending Publication Date: 2025-05-27HEBI POWER SUPPLY OF HENAN ELECTRIC POWERCORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411933008.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional machine learning models are difficult to effectively identify old target categories in power operation scenarios, and are limited by specific data sets and categories, and cannot adapt to the needs of a large number of unknown target categories in power operation scenarios, resulting in insufficient practicality and adaptability of the model.

Method used

The incremental target safety detection method is adopted, and the teacher network model and the student network model are constructed, and the knowledge of the teacher network model is transferred to the student network model by using knowledge distillation to realize the identification of new and old goals. Through the improved progressive feature pyramid network structure and multiple attention mechanism detection heads, the network recognition accuracy and adaptability are optimized.

Benefits of technology

The ability of students' network model to identify new goals while not forgetting old knowledge is realized, improves the practicality and adaptability of target detection in power operation scenarios, and enhances the network's knowledge transfer ability and detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047662A_ABST
    Figure CN120047662A_ABST
Patent Text Reader

Abstract

The invention discloses a safety detection method for an incremental target in a power operation scene, and the method comprises the steps: collecting an image in the power operation scene, marking the image, and constructing a data set; constructing a teacher network model and a student network model; training a teacher network model by using the first training data set to obtain an initial weight of the teacher network model; training a student network model by using the second training data set to obtain an initial weight of the student network model; knowledge of the teacher network model is migrated to the student network model by utilizing knowledge distillation, and incremental target safety regulation detection in the electric power work scene is realized by utilizing the student network model. According to the method, a knowledge distillation method is adopted for classification knowledge and positioning knowledge in response, incremental knowledge is extracted from the response of the classification knowledge and the positioning knowledge under reasoning information of a teacher network model, a powerful and efficient student network model is learned step by step, and incremental target detection in an electric power work scene is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and particularly to an incremental target safety regulation detection method for power operation scenarios. Background Art

[0002] The safety of power production is crucial for the stable operation of the power system. Safety accidents caused by non-standard operations in power production work account for a relatively large proportion, seriously threatening the safety of production personnel and the normal operation of the power system. Therefore, accurately and quickly identifying the illegal risk behaviors in on-site power operations and improving the safety level of on-site power operations are of great significance for power production safety.

[0003] With the development of deep learning, real-time target detection has been increasingly applied in the safety prevention and control field of power operations. However, most traditional machine learning models are deployed in static environments and rely on pre-collected data sets for offline training. Once the machine learning model is determined, it cannot be further updated. When the model encounters new data, it often overwrites or forgets the knowledge learned before, which is called catastrophic forgetting. In the field of target detection in power scenarios, catastrophic forgetting may cause the model to be unable to effectively identify the old target categories, thus reducing the practicality of the model.

[0004] In addition, the safety regulation target detection algorithm for on-site power scenarios is often limited to training with specific data sets, learning a fixed number of categories for specific scenarios. For example, if only the person category is trained on the target detection model and there is a new category such as whether to wear gloves (glove) that needs to be detected, traditional target detection requires collecting pictures of the person and glove categories and annotating a large amount of data, and then retraining with the annotated data of the two categories, which greatly limits the development and application of target detection technology. Moreover, in real scenarios, the on-site power operation scenarios contain a large number of unknown target categories, and the types of targets are numerous, far exceeding the number of types defined by humans. In this case, it is difficult to keep up by continuously annotating data and training models. Summary of the Invention

[0005] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide an incremental target safety regulation detection method for power operation scenarios.

[0006] Technical Solution: An incremental target safety regulation detection method for power operation scenarios of the present invention includes the following steps:

[0007] Step 1, collect images in power operation scenarios, construct a data set after annotating the images, and divide the data set into a first training data set and a second training data set, and there is no intersection between the two training data sets;

[0008] Step 2, construct a teacher network model and a student network model; the model structures of the teacher network model and the student network model are the same;

[0009] Step 3, use the first training dataset to train the teacher network model to obtain the initial weights of the teacher network model;

[0010] Step 4, use the second training dataset to train the student network model to obtain the initial weights of the student network model;

[0011] Step 5, use knowledge distillation to transfer the knowledge of the teacher network model to the student network model, and use the student network model to achieve incremental target safety code detection in the power operation scenario.

[0012] Further, the teacher network model includes a Backbone module, a Neck module, and a Head module;

[0013] Among them, the Backbone module is ResNet-50, the Neck module is an improved progressive feature pyramid network structure, and the Head module is a multi-attention mechanism detection head.

[0014] Further, after training the teacher network model with the first training dataset, it also includes:

[0015] The teacher network model outputs position information prediction and category information prediction through the multi-attention mechanism detection head, and regards the position information prediction and category information prediction as the knowledge of the teacher network model.

[0016] Further, using knowledge distillation to transfer the knowledge of the teacher network model to the student network model includes:

[0017] Calculate the first probability distribution of the teacher network model and the second probability distribution of the student network model respectively, minimize the KL divergence loss between the first probability distribution and the second probability distribution, and transfer the category knowledge recognized by the teacher network model to the student network model;

[0018] Calculate the divergence loss of each position box of the teacher network model and the learning model respectively, and transfer the position knowledge recognized by the teacher network model to the student network model by weighted summing the sum of the KL divergence losses of each regression box.

[0019] Further, the loss function of the student network model is:

[0020] L total =αL model_cls +L model_reg +βL dist_cls +L dist_reg

[0021] In the formula, L model_reg, L model_cls respectively represent the location knowledge loss and the category knowledge loss calculated by the student detector for the new category. L dist_reg represents the location knowledge distillation loss, and L dist_cls represents the category knowledge distillation loss. α and β are hyperparameters.

[0022] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are as follows:

[0023] (1) In the response, the present invention adopts a knowledge distillation method for classification knowledge and localization knowledge. Different from the feature-based method, the response-based method can provide the inference information of the teacher network model. By extracting incremental knowledge from the responses of classification knowledge and localization knowledge, a powerful and efficient student network model can be gradually learned;

[0024] (2) While taking into account the performance of the student network model in identifying new tasks, the loss function of the present invention also ensures that the old knowledge is not forgotten. At the same time, the knowledge distillation strategy is considered in the loss function, optimizing the student network model and improving the detection performance;

[0025] (3) In order to better improve the incremental learning task of the network, the present invention optimizes the object detection and recognition algorithm, incorporates an improved progressive feature pyramid network structure and a multi-attention mechanism detection head. While ensuring the speed of network inference, it focuses on the key area and can effectively improve the recognition accuracy of the student network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of an incremental target safety code detection method for a power operation scenario;

[0027] Figure 2 is a structural framework diagram of the teacher network model and the student network model;

[0028] Figure 3 is a schematic structural diagram of AFPNs;

[0029] Figure 4 is a schematic structural diagram of the multi-attention mechanism detection head;

[0030] Figure 5 is a schematic diagram of the principle of knowledge distillation;

[0031] Figure 6 is a flowchart of incremental training;

[0032] Figure 7 is an effect diagram of incremental target safety code detection. DETAILED DESCRIPTION OF THE INVENTION

[0033] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments.

[0034] An incremental target safety code detection method for power operation scenarios described in this embodiment has a flowchart as Figure 1 shown, and this method at least includes the following steps 1 to 5.

[0035] Step 1: Collect images in the power operation scenario, annotate the images and then construct a data set, and divide this data set into a first training data set and a second training data set, and there is no intersection between the two training data sets.

[0036] In one example, in the incremental target safety code detection of power operation scenario images, obtain on-site images during power operation, use the tool LabelImg for annotation strictly in accordance with the annotation specifications, calibrate badge, person, glove, etc. as old targets, and construct an old target data set, that is, the first training data set. Use other incremental targets that do not belong to the first training data set as new targets, and construct a new target data set, that is, the second training data set, to ensure that there is no intersection between the new and old target data sets and ensure the effectiveness of the incremental learning process.

[0037] Step 2: Construct a teacher network model and a student network model; where the model structures of the teacher network model and the student network model are the same.

[0038] Step 3: Use the first training data set to train the teacher network model to obtain the initial weights of the teacher network model.

[0039] Use the old target data set to train the teacher network model to obtain the initial weights of the teacher network model, and the initial weights can identify old targets.

[0040] Step 4: Use the second training data set to train the student network model to obtain the initial weights of the student network model.

[0041] Use the new target data set of incremental learning to train the student network model to obtain the initial weights of the student network model, so that it has the ability to identify new targets.

[0042] Step 5: Use knowledge distillation to transfer the knowledge of the teacher network model to the student network model, and use the student network model to realize incremental target safety code detection in the power operation scenario.

[0043] Use knowledge distillation to transfer the old targets learned by the teacher network model to the student network model, so that the student network model obtains the weights after incremental learning, and the student network model has the ability to identify both old targets and new targets.

[0044] As Figure 2As shown in the figure, the teacher network model includes a Backbone module, a Neck module, and a Head module;

[0045] Among them, the Backbone module is ResNet-50, the Neck module is an improved progressive feature pyramid network structure, and the Head module is a multi-attention mechanism detection head. After inputting the training set images into the teacher network model, the position information head and the class information head are output through the Head module, which are used for predicting position information and class information respectively. The network structures of the teacher network model and the student network model are the same, but the weights are different after training.

[0046] Since some of the targets to be detected in on-site power operations, such as gloves, are small in size, and there are a certain number of targets presented as small targets in the on-site power operation images and are relatively dense, this requires that the feature fusion stage of the network can integrate more information of small targets to avoid missing detection of small targets and their defects. The Asymptotic Feature Pyramid Network (AFPN) uses adaptive spatial features to gradually fuse the feature information in adjacent levels, avoiding a large semantic gap between non-adjacent levels. In this example, based on AFPN, the improved progressive feature pyramid network structure (AFPN Small, AFPNs) is constructed by further fusing the shallow network, as Figure 3 shown. The shallow stages C1 to C2 contain information of some small targets. Therefore, fusing the features of the C1 to C2 stages can also help the network focus on the information of the vibration damping hammer small targets. AFPNs adopts a progressive structure, gradually fusing the deep network from the shallow network to prevent the loss or degradation of feature information during transmission and interaction. AFPNs first uses ASFF and the top-down structure to fuse the features of the C1 to C2 stages to the C3 stage, and uses Conv-ABlock to integrate the features again in the top-down structure; then C3 and C4 are processed by ASFF, and after being integrated by Conv-ABlock respectively, they are processed with the C5 using the ASFF structure again; finally, three feature maps are output for prediction.

[0047] Due to the change of the shooting environment, especially in an adverse environment, it will lead to a reduction in the effective pixel information of the targets in the on-site power operation images. When the incremental learning algorithm identifies targets, it is easily affected by complex backgrounds and environmental factors, resulting in insufficient ability of the algorithm to extract effective features of the target objects. In order to concentrate algorithm resources as much as possible on the extraction of more critical information of the on-site power operation scene targets, the attention mechanism can be used to significantly improve the expression ability of the model's neck network without increasing the computational amount. In this example, a multi-attention mechanism detection head is used as the Head module, as Figure 4As shown, by applying the attention mechanism from three different perspectives of scale perception, spatial position, and multi-task, the algorithm's perception ability of scale, spatial position, and multi-task is improved. The multi-attention mechanism detection head includes a scale attention module, a spatial attention module, and a multi-task attention module. The scale attention module improves the expression ability at different levels to effectively enhance the scale perception ability of the object detector.

[0048] Furthermore, after training the teacher network model using the first training dataset, it also includes:

[0049] The teacher network model outputs position information prediction and class information prediction through the multi-attention mechanism detection head, and takes the position information prediction and class information prediction as the knowledge of the teacher network model.

[0050] Input the features output by the Neck module into the multi-attention mechanism detection head to obtain a position information head and a class information head respectively. The position information head is used for the prediction of position information, and the class information head is used for the prediction of class information.

[0051] As Figure 5 shown, furthermore, the transfer of the knowledge of the teacher network model to the student network model using knowledge distillation includes:

[0052] Calculate the first probability distribution of the teacher network model and the second probability distribution of the student network model respectively, minimize the KL divergence loss between the first probability distribution and the second probability distribution, and transfer the class knowledge recognized by the teacher network model to the student network model;

[0053] Calculate the divergence loss of each position box of the teacher network model and the learning model respectively, and transfer the position knowledge recognized by the teacher network model to the student network model by weighted summing the sum of the KL divergence losses of each regression box.

[0054] The core idea of knowledge distillation is to use a pre-trained teacher network model to guide the learning of a student network model. In this example, the base class badges "badge", "person", and "glove" are calibrated as old targets to train the teacher network model. For the class knowledge of the old targets, the teacher network model outputs n logics z, and after passing through a softmax transformation with a temperature τ, the first probability distribution p = S(z, τ) is obtained. The probability distribution p represents the probability of belonging to each category in "badge", "person", and "glove". At the same time, calculate the second probability distribution of the student network model, and minimize the second probability distribution p S and the first probability distribution p TWith the KL divergence loss, the category knowledge of the teacher network model, such as badge, person, and glove, can be transferred to the student network model, so as to achieve the effect that the student network model does not forget the old target categories. For the location knowledge of the old targets, calculate the KL divergence loss of the bounding boxes of each object detection between the teacher network model and the learning model. In the formula, e represents each side of the bounding box, and B represents the bounding box. represents calculating the KL divergence loss for each side of the bounding box. represents the regression response of the teacher detector to the j selected bounding boxes using new data. represents the corresponding regression response of the student detector to the j selected bounding boxes. Finally, by weighted summing the KL divergence losses of each bounding box, the location knowledge of the teacher network model is transmitted to the student network model.

[0055] Furthermore, the loss function of the student network model is:

[0056] L total =αL model_cls +L model_reg +βL dist_cls +L dist_reg

[0057] In the formula, L model_reg and L model_cls respectively represent the location knowledge loss and category knowledge loss calculated by the student detector for the new categories. L dist_reg represents the location knowledge distillation loss, and L dist_cls represents the category knowledge distillation loss. α and β are hyperparameters.

[0058] The loss term L model is the classification and localization loss for the detector, used to train the student network model to detect new targets. The hyperparameter α is a hyperparameter used to balance the loss weights, and the hyperparameter β is a hyperparameter used to balance the hierarchical distillation loss weights. The hyperparameters α and β improve the robustness of the detection performance for new and old categories, ensuring the incremental learning effect and final performance of the model. While taking into account the student network model's recognition of new tasks, this loss function also ensures that it does not forget old knowledge. For the divide-and-conquer distillation strategy, the network model is optimized, improving the detection performance.

[0059] In an example, such as Figure 6As shown, for the power operation scenario, training samples are collected starting from the incremental detection task. There are a total of 6 classes, and an old-target training dataset and a new-target training dataset are constructed. Among them, there are 3 classes for the old targets, including badge, person, and glove. There are 3 classes for the new targets, including wrongglove, operatingbar, and powerchecker. There is no overlap between the classes of the old and new target training datasets. Such a design ensures the effectiveness of the incremental learning process. Subsequently, the teacher network model is trained using the old-target dataset to obtain the initial weights of the teacher network model. The AFPNS module is adopted in the neck network of the teacher network model, and the multi-attention mechanism detector is adopted in the detection head network. Finally, systematic learning and training are carried out to obtain the teacher network model. Then, the student network model is trained using the new-target dataset to obtain the weights after incremental learning of the learning model. To enable the student network model to better learn the old knowledge of the teacher network model, the AFPNS module is adopted in the neck network of the student network, and the multi-attention mechanism detector is adopted in the detection head network to align with the teacher network model. At the same time, to avoid catastrophic forgetting of the student network, the divide-and-conquer strategy is used, and different distillation losses are used for the location information and category information respectively. The performance of the teacher network model in identifying old targets is transferred to the student network model, so that the student network model can identify both new and old targets after obtaining the weights after incremental learning, thus achieving the effect of incremental learning.

[0060] As Figure 7 shown, in the initial training, as shown in Figure (a), the initial weights obtained using the teacher network model can identify three old targets: badge, person, and glove. After incremental training, as shown in Figure (b), the incremental weights output by the student network model can still identify the three old targets: badge, person, and glove, and at the same time, it can also identify two new targets: wrongglove and powerchecker. The experimental results demonstrate the effectiveness of the incremental target safety code detection method in this example.

[0061] A method for incremental target safety code detection in the power operation scenario described in the present invention uses an incremental target detection algorithm and proposes a divide-and-conquer distillation strategy for the corresponding distillation method. Different distillation methods are used for location knowledge and category knowledge respectively, so as to transfer the knowledge of the old task from the teacher network to the student network. At this time, the student network detector has the ability to recognize new knowledge and does not forget the old knowledge, greatly improving the knowledge transfer ability of the network. In the present invention, AFPNs are built based on the progressive feature pyramid network structure to fuse more shallow features to detect small targets such as gloves in the image. In order to concentrate algorithm resources as much as possible on the extraction of information that is more critical to the targets in the power operation scenario, the present invention uses a multi-attention mechanism detection head method. By applying the attention mechanism from three different angles respectively, the algorithm's perception ability of scale, spatial position, and multi-task is improved, focusing on the key area and improving the incremental learning detection performance. The incremental detection algorithm proposed in the present invention is innovatively applied in the power operation scenario and has achieved good detection results.

Claims

1. A safety regulation detection method for incremental targets in power operation scenarios, characterized in that: The steps include: Step 1: collect images of power operation scenes, annotate the images and construct a data set, and divide the data set into a first training data set and a second training data set, and there is no intersection between the two training data sets; Step 2, construct the teacher network model and the student network model; The model structures of the teacher network model and the student network model are the same; Step 3, using the first training data set to train the teacher network model to obtain the initial weight of the teacher network model; Step 4, using the second training data set to train the student network model to obtain the initial weight of the student network model; Step 5: Use knowledge distillation to transfer the knowledge of the teacher network model to the student network model, and use the student network model to achieve incremental target safety detection in power operation scenarios.

2. The incremental target safety detection method for electric power operation scenarios according to claim 1 is characterized in that: The teacher network model includes Backbone module, Neck module and Head module; The Backbone module is ResNet-50, the Neck module is an improved progressive feature pyramid network structure, and the Head module is a multiple attention mechanism detection head.

3. The incremental target safety detection method in the power operation scene according to claim 2 is characterized in that: After training the teacher network model using the first training data set, it includes: The teacher network model outputs position information prediction and category information prediction through multiple attention mechanism detection heads, and uses the position information prediction and category information prediction as the knowledge of the teacher network model.

4. The incremental target safety detection method in the electric power operation scene according to claim 3 is characterized in that: Using knowledge distillation to transfer the knowledge of the teacher network model to the student network model includes: Calculate the first probability distribution of the teacher network model and the second probability distribution of the student network model respectively, minimize the KL divergence loss between the first probability distribution and the second probability distribution, and transfer the category knowledge recognized by the teacher network model to the student network model; The divergence loss of each position box of the teacher network model and the learning model is calculated respectively, and the position knowledge recognized by the teacher network model is transferred to the student network model by weighted summing up the KL divergence loss of each regression box.

5. The incremental target safety detection method for electric power operation scenarios according to claim 1 is characterized in that: The loss function of the student network model is: L total =αL model_cls +L model_reg +βL dist_cls +L dist_reg Where, L model_reg , L model_cls They represent the position knowledge loss and category knowledge loss calculated by the student detector for the new category, L dist_reg represents the position knowledge distillation loss, L dist_cls represents the category knowledge distillation loss, and α and β are hyperparameters.