Power image anomaly detection method and device, computer device and storage medium

By using a pre-trained power image anomaly detection model, abnormal regions in power images can be identified and marked, solving the problem of low efficiency in traditional methods and achieving efficient and accurate power image anomaly detection.

CN117132763BActive Publication Date: 2026-03-31MAINTENANCE & TEST CENTRE CSG EHV POWER TRANSMISSION CO
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional power image anomaly detection methods are inefficient and cannot efficiently process large numbers of power images.

Method used

A pre-trained power image anomaly detection model is used to identify abnormal regions in power images and compare them with pre-stored anomaly description texts. The description text with the highest matching degree is selected for labeling, and anomaly-labeled power images are generated.

Benefits of technology

It improves the efficiency and accuracy of power image anomaly detection, and can quickly identify and mark abnormal areas in power images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132763B_ABST
    Figure CN117132763B_ABST
Patent Text Reader

Abstract

The application relates to a power image anomaly detection method and device, computer equipment, a storage medium and a computer program product, which can be used in the technical field of electric power. The application can improve the efficiency and accuracy of power image anomaly detection. The method comprises the following steps: acquiring a power image; identifying an abnormal image area by using a power image anomaly detection model; identifying the matching degree between the abnormal image area and an abnormal description text by using the power image anomaly detection model to obtain a matching degree identification result; selecting the abnormal description text with the highest matching degree from the abnormal description text as a target abnormal description text according to the matching degree identification result; identifying the abnormal image area in the power image by using the target abnormal description text through the power image anomaly detection model to obtain an abnormal identification power image; and taking the abnormal identification power image as the abnormal detection result of the power image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for detecting anomalies in power images. Background Technology

[0002] With the development of power technology, ensuring the safe operation of power equipment is crucial for guaranteeing a stable power supply. Real-time anomaly detection of power images is typically used to promptly identify potential power problems. Therefore, how to efficiently detect anomalies in power images has become an important research direction.

[0003] Traditional techniques typically involve manually inspecting each power image to detect anomalies. However, when dealing with a large number of power images, this method requires a significant amount of manual inspection time, resulting in low efficiency in power image anomaly detection. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for detecting power image anomalies that can improve the efficiency of power image anomaly detection, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for detecting anomalies in power images. The method includes:

[0006] Acquire power images;

[0007] Anomaly detection model for power images is used to identify anomalies in the power images and identify abnormal image regions in the power images.

[0008] The matching degree between the abnormal image region and the pre-stored abnormal description text is identified using the pre-trained power image anomaly detection model to obtain the matching degree identification result.

[0009] Based on the matching degree recognition result, the anomaly description text with the highest matching degree is selected from the pre-stored anomaly description text and used as the target anomaly description text for the anomaly image region.

[0010] Using the pre-trained power image anomaly detection model, the abnormal image regions in the power image are identified using the target anomaly description text to obtain an anomaly-identified power image.

[0011] The abnormal power image is used as the anomaly detection result of the power image.

[0012] In one embodiment, the pre-trained power image anomaly detection model is trained in the following manner:

[0013] Obtain the first sample power image and the corresponding real anomaly detection result;

[0014] The first sample power image is input into the power image anomaly detection model to be trained to obtain the predicted anomaly detection result of the first sample power image.

[0015] Based on the difference between the predicted anomaly detection results and the actual anomaly detection results, the power image anomaly detection model to be trained is trained to obtain the pre-trained power image anomaly detection model.

[0016] In one embodiment, training the power image anomaly detection model to be trained based on the difference between the predicted anomaly detection result and the actual anomaly detection result to obtain the pre-trained power image anomaly detection model includes:

[0017] Based on the difference between the predicted anomaly detection results and the actual anomaly detection results, the power image anomaly detection model to be trained is trained to obtain a basic power image anomaly detection model.

[0018] Based on the model parameters and model structure of the basic power image anomaly detection model, a first power image anomaly detection model and a second power image anomaly detection model corresponding to the basic power image anomaly detection model are determined.

[0019] The unlabeled second sample power image is input into the first power image anomaly detection model to obtain the pseudo label of the second sample power image;

[0020] The second sample power image and its pseudo-label are used to train the second power image anomaly detection model, resulting in the pre-trained power image anomaly detection model.

[0021] In one embodiment, acquiring the first sample power image includes:

[0022] Acquire abnormal images of power equipment and abnormal human behavior within the substation area;

[0023] The abnormal images of the power equipment and the abnormal images of personnel behavior are combined to obtain the first sample power image.

[0024] In one embodiment, acquiring the power image includes:

[0025] Acquire raw power images within the substation area; the raw power images include images of power equipment and images of personnel behavior;

[0026] The original power image is preprocessed to obtain a power image.

[0027] In one embodiment, before identifying the matching degree between the abnormal image region and the pre-stored abnormal description text using the pre-trained power image anomaly detection model to obtain the matching degree identification result, the method further includes:

[0028] Historical anomaly information within the substation area is identified to obtain the types of historical anomalies within the substation area;

[0029] The historical anomaly types are identified to obtain historical anomaly description text within the substation area;

[0030] The historical anomaly description text is used as a pre-stored anomaly description text.

[0031] Secondly, this application also provides a power image anomaly detection device. The device includes:

[0032] The image acquisition module is used to acquire power images;

[0033] The image recognition module is used to identify anomalies in the power image using a pre-trained power image anomaly detection model, and to identify abnormal image regions in the power image.

[0034] The text recognition module is used to identify the matching degree between the abnormal image region and the pre-stored abnormal description text through the pre-trained power image anomaly detection model, and obtain the matching degree recognition result.

[0035] The text selection module is used to select the anomaly description text with the highest matching degree from the pre-stored anomaly description text based on the matching degree recognition result, and use it as the target anomaly description text for the anomaly image region.

[0036] The region identification module is used to identify the abnormal image region in the power image using the target abnormal description text through the pre-trained power image anomaly detection model, so as to obtain an abnormally identified power image.

[0037] The result determination module is used to take the abnormal power image as the abnormal detection result of the power image.

[0038] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0039] Acquire power images;

[0040] Anomaly detection model for power images is used to identify anomalies in the power images and identify abnormal image regions in the power images.

[0041] The matching degree between the abnormal image region and the pre-stored abnormal description text is identified using the pre-trained power image anomaly detection model to obtain the matching degree identification result.

[0042] Based on the matching degree recognition result, the anomaly description text with the highest matching degree is selected from the pre-stored anomaly description text and used as the target anomaly description text for the anomaly image region.

[0043] Using the pre-trained power image anomaly detection model, the abnormal image regions in the power image are identified using the target anomaly description text to obtain an anomaly-identified power image.

[0044] The abnormal power image is used as the anomaly detection result of the power image.

[0045] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0046] Acquire power images;

[0047] Anomaly detection model for power images is used to identify anomalies in the power images and identify abnormal image regions in the power images.

[0048] The matching degree between the abnormal image region and the pre-stored abnormal description text is identified using the pre-trained power image anomaly detection model to obtain the matching degree identification result.

[0049] Based on the matching degree recognition result, the anomaly description text with the highest matching degree is selected from the pre-stored anomaly description text and used as the target anomaly description text for the anomaly image region.

[0050] Using the pre-trained power image anomaly detection model, the abnormal image regions in the power image are identified using the target anomaly description text to obtain an anomaly-identified power image.

[0051] The abnormal power image is used as the anomaly detection result of the power image.

[0052] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0053] Acquire power images;

[0054] Anomaly detection model for power images is used to identify anomalies in the power images and identify abnormal image regions in the power images.

[0055] The matching degree between the abnormal image region and the pre-stored abnormal description text is identified using the pre-trained power image anomaly detection model to obtain the matching degree identification result.

[0056] Based on the matching degree recognition result, the anomaly description text with the highest matching degree is selected from the pre-stored anomaly description text and used as the target anomaly description text for the anomaly image region.

[0057] Using the pre-trained power image anomaly detection model, the abnormal image regions in the power image are identified using the target anomaly description text to obtain an anomaly-identified power image.

[0058] The abnormal power image is used as the anomaly detection result of the power image.

[0059] The aforementioned power image anomaly detection method, apparatus, computer equipment, storage medium, and computer program product acquire a power image; use a pre-trained power image anomaly detection model to identify anomalies in the power image, identifying anomalous image regions within the power image; use the pre-trained power image anomaly detection model to identify the matching degree between the anomalous image regions and pre-stored anomaly description text, obtaining a matching degree identification result; based on the matching degree identification result, select the anomaly description text with the highest matching degree from the pre-stored anomaly description text as the target anomaly description text for the anomalous image region; use the pre-trained power image anomaly detection model to mark the anomalous image regions in the power image using the target anomaly description text, obtaining an anomaly-marked power image; and use the anomaly-marked power image as the anomaly detection result for the power image. This scheme acquires power images to obtain images for anomaly detection. A pre-trained power image anomaly detection model identifies anomalous regions within the power images, thus identifying the regions with anomalies in the images to be detected. The pre-trained model then assesses the matching degree between these anomalous regions and pre-stored anomaly description text, obtaining a matching degree recognition result. Based on this result, the pre-stored anomaly description text with the highest matching degree is selected as the target anomaly description text for the anomalous region, quickly and accurately determining the corresponding anomaly description text. The pre-trained model then uses the target anomaly description text to label the anomalous regions in the power images, obtaining an anomaly-labeled power image, thus quickly and accurately identifying the corresponding anomaly-labeled power image. This anomaly-labeled power image is used as the anomaly detection result, thereby improving the efficiency and accuracy of power image anomaly detection. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] Figure 1 This is a flowchart illustrating a power image anomaly detection method in one embodiment;

[0062] Figure 2This is a flowchart illustrating the steps of determining a pre-trained power image anomaly detection model in one embodiment.

[0063] Figure 3 This is a flowchart illustrating the steps for determining a pre-trained power image anomaly detection model in another embodiment.

[0064] Figure 4 This is a structural block diagram of a power image anomaly detection device in one embodiment;

[0065] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0067] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0068] In one embodiment, such as Figure 1 As shown, a method for detecting anomalies in power images is provided. This embodiment illustrates the application of this method to a terminal; it is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc.; the server can be a standalone server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0069] Step S101: Obtain the power image.

[0070] In this step, the power image can be an image of the power area to be detected as an anomaly, such as an image of power equipment in a substation area and an image of human behavior in that substation area.

[0071] Specifically, the terminal acquires power images captured by the image-capturing device.

[0072] Step S102: Using a pre-trained power image anomaly detection model, anomalies are identified in the power image, and abnormal image regions in the power image are identified.

[0073] In this step, the pre-trained power image anomaly detection model can be a trained model for detecting anomalies in power images. The power image anomaly detection model can be a multimodal model, such as the GLIP model (language-image matching pre-trained model). The abnormal image regions in the power image can be the predicted image regions in the power image that are abnormal.

[0074] Specifically, the terminal inputs the power image into a pre-trained power image anomaly detection model, which then identifies anomalies in the power image and identifies abnormal image regions within it.

[0075] Step S103: Using a pre-trained power image anomaly detection model, the matching degree between the abnormal image region and the pre-stored anomaly description text is identified to obtain the matching degree identification result.

[0076] In this step, the pre-stored anomaly description text can be pre-stored or pre-set text used to describe the anomaly content. For example, the pre-stored anomaly description text may include: respirator oil seal damage, dial blurring, dial breakage, insulator damage, ground oil stains, silicone cylinder damage, abnormal door closure, hanging suspended objects, bird nests, cover plate damage, not wearing a safety helmet, not wearing work clothes, personnel smoking, abnormal respirator oil seal oil level and / or silicone discoloration; the matching degree recognition result can be used to represent the matching degree between the abnormal image area and each pre-stored anomaly description text.

[0077] Specifically, the terminal inputs the abnormal image region and the pre-stored abnormal description text into the pre-trained power image anomaly detection model. The pre-trained power image anomaly detection model identifies the matching degree between the abnormal image region and the pre-stored abnormal description text, and obtains the matching degree identification result.

[0078] Step S104: Based on the matching degree recognition result, select the anomaly description text with the highest matching degree from the pre-stored anomaly description text as the target anomaly description text for the anomaly image region.

[0079] In this step, the anomaly description text with the highest matching degree can be the anomaly description text with the highest matching degree between it and the anomaly image region; the target anomaly description text of the anomaly image region can be the text used to describe the anomaly content (anomaly information / anomaly type) of the anomaly image region.

[0080] Specifically, based on the matching degree recognition result, the terminal selects the anomaly description text with the highest matching degree between the corresponding anomaly image region from the pre-stored anomaly description text, and uses it as the target anomaly description text for the anomaly image region.

[0081] Step S105: Using a pre-trained power image anomaly detection model, the abnormal image regions in the power image are identified by the target anomaly description text to obtain an anomaly-identified power image.

[0082] In this step, the abnormal power image can be a power image that has undergone abnormal identification processing.

[0083] Specifically, the terminal uses a pre-trained power image anomaly detection model to mark the target anomaly description text in the corresponding mark area of ​​the power image, thus obtaining an anomaly-marked power image.

[0084] Step S106: The abnormal power image is identified as the abnormal detection result of the power image.

[0085] Specifically, the terminal will identify abnormal power images as the results of power image anomaly detection.

[0086] In the above-mentioned power image anomaly detection method, a power image is acquired; anomalies are identified in the power image using a pre-trained power image anomaly detection model, and abnormal image regions in the power image are identified; the matching degree between the abnormal image regions and pre-stored abnormal description texts is identified using the pre-trained power image anomaly detection model, and a matching degree identification result is obtained; based on the matching degree identification result, the abnormal description text with the highest matching degree is selected from the pre-stored abnormal description texts as the target abnormal description text for the abnormal image region; the abnormal image regions in the power image are labeled using the target abnormal description texts through the pre-trained power image anomaly detection model, and an anomaly-labeled power image is obtained; the anomaly-labeled power image is used as the anomaly detection result of the power image. This scheme acquires power images to obtain images for anomaly detection. A pre-trained power image anomaly detection model identifies anomalous regions within the power images, thus identifying the regions with anomalies in the images to be detected. The pre-trained model then assesses the matching degree between these anomalous regions and pre-stored anomaly description text, obtaining a matching degree recognition result. Based on this result, the pre-stored anomaly description text with the highest matching degree is selected as the target anomaly description text for the anomalous region, quickly and accurately determining the corresponding anomaly description text. The pre-trained model then uses the target anomaly description text to label the anomalous regions in the power images, obtaining an anomaly-labeled power image, thus quickly and accurately identifying the corresponding anomaly-labeled power image. This anomaly-labeled power image is used as the anomaly detection result, thereby improving the efficiency and accuracy of power image anomaly detection.

[0087] In one embodiment, such as Figure 2 As shown, the pre-trained power image anomaly detection model is trained in the following manner, specifically including the following:

[0088] Step S201: Obtain the first sample power image and the corresponding real anomaly detection result;

[0089] Step S202: Input the first sample power image into the power image anomaly detection model to be trained to obtain the predicted anomaly detection result of the first sample power image;

[0090] Step S203: Based on the difference between the predicted anomaly detection results and the actual anomaly detection results, train the power image anomaly detection model to be trained to obtain the pre-trained power image anomaly detection model.

[0091] In this embodiment, the first sample power image can be a historical power image used as the first sample, or it can be a labeled sample power image, wherein the label can refer to the actual anomaly detection result corresponding to the sample power image; the power image anomaly detection model to be trained can be an untrained power image anomaly detection model.

[0092] Specifically, the terminal acquires a first sample power image and the corresponding real anomaly detection result; inputs the first sample power image into the power image anomaly detection model to be trained, and obtains the predicted anomaly detection result of the first sample power image output by the power image anomaly detection model to be trained; determines the difference information based on the difference between the predicted anomaly detection result and the real anomaly detection result, and uses the difference information to train the power image anomaly detection model to be trained, and obtains the trained power image anomaly detection model, which is then used as the pre-trained power image anomaly detection model.

[0093] The technical solution provided in this embodiment trains the power image anomaly detection model to be trained based on the difference between the predicted anomaly detection results and the actual anomaly detection results, thereby obtaining a pre-trained power image anomaly detection model. This is beneficial for obtaining a more efficient and accurate pre-trained power image anomaly detection model, thus improving the efficiency and accuracy of power image anomaly detection.

[0094] In one embodiment, such as Figure 3 As shown, in the above steps, the power image anomaly detection model to be trained is trained based on the difference between the predicted anomaly detection results and the actual anomaly detection results, resulting in a pre-trained power image anomaly detection model. This specifically includes the following:

[0095] Step S301: Based on the difference between the predicted anomaly detection results and the actual anomaly detection results, train the power image anomaly detection model to be trained to obtain the basic power image anomaly detection model.

[0096] Step S302: Based on the model parameters and model structure of the basic power image anomaly detection model, determine the first power image anomaly detection model and the second power image anomaly detection model corresponding to the basic power image anomaly detection model.

[0097] Step S303: Input the unlabeled second sample power image into the first power image anomaly detection model to obtain the pseudo label of the second sample power image;

[0098] Step S304: Using the second sample power image and its pseudo-label, train the second power image anomaly detection model to obtain the pre-trained power image anomaly detection model.

[0099] In this embodiment, the basic power image anomaly detection model can be a trained power image anomaly detection model; the first power image anomaly detection model and the second power image anomaly detection model can be the teacher model and student model corresponding to the basic power image anomaly detection model, respectively. The model parameters and model structure of the first power image anomaly detection model can be the same as those of the basic power image anomaly detection model, and the model parameters and model structure of the second power image anomaly detection model can be the same as those of the basic power image anomaly detection model; the unlabeled second sample power image can be a historical power image without a corresponding real anomaly detection result; the pseudo-label of the second sample power image can be the predicted anomaly detection result of the second sample power image output by the first power image anomaly detection model.

[0100] Specifically, the terminal trains the power image anomaly detection model to be trained based on the difference between the predicted anomaly detection results and the actual anomaly detection results, thus obtaining a basic power image anomaly detection model. Based on the model parameters and model structure of the basic power image anomaly detection model, a first power image anomaly detection model and a second power image anomaly detection model corresponding to the basic power image anomaly detection model are determined. An unlabeled second sample power image is input into the first power image anomaly detection model to obtain the pseudo-label of the second sample power image output by the first power image anomaly detection model. Using the second sample power image and the pseudo-label of the second sample power image, the second power image anomaly detection model is trained to obtain a pre-trained power image anomaly detection model.

[0101] The technical solution provided in this embodiment trains the power image anomaly detection model using a semi-supervised learning approach, which helps to obtain a more efficient and accurate pre-trained power image anomaly detection model, thereby improving the efficiency and accuracy of power image anomaly detection.

[0102] In one embodiment, the above steps of obtaining the first sample power image specifically include the following: obtaining abnormal images of power equipment and abnormal images of personnel behavior in the substation area; and combining the abnormal images of power equipment and abnormal images of personnel behavior to obtain the first sample power image.

[0103] In this embodiment, the abnormal personnel behavior image can be an image of abnormal personnel behavior within the substation area.

[0104] Specifically, the terminal acquires abnormal images of power equipment and abnormal images of personnel behavior within the substation area; both abnormal images of power equipment and abnormal images of personnel behavior are used as the first sample power images.

[0105] The technical solution provided in this embodiment combines abnormal images of power equipment and abnormal images of human behavior to obtain a first sample power image. This is beneficial for obtaining a richer and more diverse first sample power image, which in turn is beneficial for obtaining a more efficient and accurate pre-trained power image anomaly detection model, thereby improving the efficiency and accuracy of power image anomaly detection.

[0106] In one embodiment, step S101, acquiring a power image, specifically includes the following: acquiring an original power image within the substation area; the original power image includes images of power equipment and images of personnel behavior; and performing image preprocessing on the original power image to obtain a power image.

[0107] In this embodiment, the original power image includes images of power equipment within the substation area and images of personnel behavior within the substation area; image preprocessing may include image size processing, image brightness and contrast processing, and / or image smoothing and sharpening processing.

[0108] Specifically, the terminal acquires images of power equipment and personnel behavior within the substation area as the original power images within the substation area. The original power images are then preprocessed to obtain the final power images.

[0109] The technical solution provided in this embodiment, by preprocessing the original power image, helps to obtain a clearer power image that is easier for the model to process, thereby improving the efficiency and accuracy of power image anomaly detection.

[0110] In one embodiment, before identifying the matching degree between the abnormal image region and the pre-stored abnormal description text through the pre-trained power image anomaly detection model and obtaining the matching degree identification result, step S103 above further includes a step of determining the pre-stored abnormal description text, specifically including the following: identifying historical anomaly information within the substation area to obtain historical anomaly types within the substation area; identifying historical anomaly types to obtain historical anomaly description text within the substation area; and using the historical anomaly description text as the pre-stored abnormal description text.

[0111] In this embodiment, historical anomaly information can be historical anomaly information within the substation area; historical anomaly type can be anomaly type of historical anomaly information within the substation area, such as historical anomaly type can include dial damage, cover plate damage, bird nest, dial blurring, ground oil stains, suspended debris, silicone tube damage, respirator oil seal damage and / or personnel smoking; historical anomaly description text can be anomaly description text.

[0112] Specifically, the terminal identifies historical anomaly information (historical anomaly content) within the substation area to obtain the historical anomaly type corresponding to the historical anomaly information within the substation area; identifies the historical anomaly type to obtain the historical anomaly description text within the substation area; and uses the historical anomaly description text as a pre-stored anomaly description text.

[0113] The technical solution provided in this embodiment identifies historical anomaly information within the substation area and determines the pre-stored anomaly description text, which helps to obtain more accurate pre-stored anomaly description text, thereby improving the accuracy of power image anomaly detection.

[0114] The following application example illustrates the power image anomaly detection method provided in this application. This application example demonstrates the method's application to a terminal, and the main steps include:

[0115] The first step involves the terminal collecting images of defects and anomalies in the substation, which can include 15 types of defects and 2,674 images. In addition to the power defect scene image data, training data including anomalies of substation personnel were also collected and compiled into a dataset.

[0116] The data can be randomly divided into a training set and a test set in an 8:2 ratio.

[0117] In the second step, the terminal inputs the dataset from the first step into the language-image matching pre-trained GLIP model for power defect detection.

[0118] (2.1) The GLIP detection model feeds the input image into a visual encoder, which typically uses a CNN (Convolutional Neural Network) or a Transformer as its backbone to extract region / box features. Each region / box feature is fed into two detection heads, namely a classifier C and a regressor R, which are trained with classification loss and localization loss, respectively. (2.2) In two-stage detectors, a Region Proposal Network (RPN) with loss is often used to distinguish foreground and background and refine anchor points. Since the semantic information of the object's category is not used, the GLIP model incorporates it into the localization loss. In single-stage detectors, the localization loss may also include centrality loss. The classifier C is usually a simple linear layer, and the classification loss can be represented by the object / region / box features of the input image, the weight matrix of the classifier C, the output classification logic, the target matching degree between the region and class calculated according to the classic many-to-one matching, which is usually cross-entropy loss for two-stage detectors and focus loss for single-stage detectors. (2.3) In the phrase matching model, the matching score between the image region and the word in the prompt needs to be calculated. This can be represented by the context word / tag features from the language encoder (which function similarly to the weight matrix in the classification loss in step 2.2). The matching model, composed of the image encoder and the language encoder, is trained end-to-end by minimizing the loss defined in steps 2.1) and 2.2). During training, the region word alignment score replaces the classification logic. (2.4) The GLIP model used employs a Swin-T (Swin Transformer) as the backbone network of the image encoder, paired with a DyHead (Dynamic Head) as the detection head. Deep fusion is introduced in the last few layers of the image and text encoders. When using a DyHead as the image encoder, it can be expressed by information such as the number of DyHead modules. Cross-modal communication is accomplished through a cross-modal multi-head self-attention module. The DyHead rewrites the output of the backbone network, i.e., the input of the detection head, into a three-dimensional tensor of horizontal × spatial × channel, and deploys a fully self-attention mechanism on the planes formed by each pair.

[0119] The third step involves the terminal applying semi-supervised learning to the GLIP model, thereby enabling training and recognition of a large number of unlabeled samples with only a small amount of labeled data; the trained model is then used for anomaly detection in power images.

[0120] The specific implementation method of the semi-supervised learning method is as follows: (3.1) In the fully supervised training stage, the model is trained using existing labeled data to initialize the detection model. (3.2) In the teacher-student model mutual learning stage, the model parameters are copied to both the teacher and student models. The teacher model generates pseudo-labels for the input unlabeled data, and then uses the pseudo-labeled data to train the student model. When the images input to the teacher model and the student model are different, the teacher model uses weakly augmented data while the student model uses strongly augmented data. Augmentation methods that can be used include random blurring, flipping, and grayscale adjustment. In order to prevent the teacher model from introducing too much noise when generating pseudo-labels and to ensure that the generated pseudo-labels meet the confidence threshold requirements, the teacher model can be gradually updated by applying an exponential moving average method after training the student model.

[0121] Among them, (1) fully supervised experiment, the model iteration rounds are 30. (2) Text prompt experiment. When the trained GLIP model tests the image, the category name and text prompt will be input into the model. Here we will explore the impact of text prompt on detection by adjusting the method. (3) Semi-supervised experiment. For a specific category, 10 images are taken for 5 rounds of fully supervised training, and then all images are used for 5 rounds of semi-supervised training. The confidence threshold during semi-supervised training is dynamically selected based on the results of supervised training. The lower the confidence threshold, the more it can learn unseen content, while the higher the confidence, the more it will tend to use the content that has been learned for repeated training. The first experiment is zero-sample and supervised quantitative experiment. The GLIP model can better match the fault text to the corresponding area of ​​the image. If there are two or more faults in the image, it can also mark them well. The recognition rate after training is greatly improved, and the model has a good recognition ability for power defects. The model has a certain recognition ability within a single category when zero samples are used. Then the text prompt quantitative experiment is conducted. Finally, the semi-supervised experiment is conducted. In summary, semi-supervised training methods, with appropriate rounds of supervised training and suitable confidence threshold selection, can significantly improve recognition performance with fewer samples and are well-suited for scenarios with limited available training samples.

[0122] The GLIP model unifies phrase matching and object detection tasks. The object detection model now inputs not only images but also text descriptions of all candidate object categories in the detection task. By using a high-performing teacher model combined with image-text pairs extracted from the network, a high-performing student model can be pre-trained, which also excels in recognizing rare samples. The GLIP model also possesses strong transfer learning capabilities. When transferring the GLIP model to downstream tasks, only a small amount of labeled data is required, and only partial model adjustments are needed to achieve good learning performance, reducing deployment costs in defect detection. Experiments were conducted on adjusting text prompts describing power defects, demonstrating that the proposed model has good zero-shot recognition and adaptability. A semi-supervised learning method is employed on the GLIP model, thus requiring only minimal labeled data to train and recognize a large number of unlabeled samples. The semi-supervised method utilizes a large amount of unlabeled data and a small amount of labeled data to train the object detection model, addressing the problem of limited existing labeled data. The unbiased teacher approach consists of two phases: a fully supervised training phase and a teacher-student model mutual learning phase. First, in the fully supervised training phase, the detection model is initialized by training the existing labeled data. Second, in the teacher-student model mutual learning phase, the model parameters are copied simultaneously to both the teacher and student models. The teacher model generates pseudo-labels for the input unlabeled data, and then uses the pseudo-labeled data to train the student model. When the input images to the teacher and student models differ, the teacher model uses weak augmentation data while the student model uses strong augmentation data. Suitable augmentation methods include random blurring, flipping, and grayscale adjustment.

[0123] The technical solution provided in this application example achieves defect recognition under various conditions through deep fusion training of excellent pre-trained models and power defect prompt text and defect images. It employs a semi-supervised learning approach, using a large amount of unlabeled data to input into the trained teacher model and generate pseudo-labels to train the student model, thereby improving defect recognition performance and enhancing the power defect detection capability. It demonstrates good transferability in the power field and provides a solution for cross-domain detection tasks. By introducing the GLIP model, based on cross-modal pre-trained data and the recognition model, it achieves excellent detection results on existing defect images after training. Based on the GLIP model's equipment defect and target detection method, using power defect image data as the training object, it analyzes the relationship between text prompts and image model recognition performance. Addressing the problem of scarce labeled data, a semi-supervised learning method is used to improve defect recognition performance, thereby enhancing the power defect detection capability and improving the efficiency and accuracy of power image anomaly detection.

[0124] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0125] Based on the same inventive concept, this application also provides a power image anomaly detection device for implementing the power image anomaly detection method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more power image anomaly detection device embodiments provided below can be found in the limitations of the power image anomaly detection method described above, and will not be repeated here.

[0126] In one embodiment, such as Figure 4 As shown, a power image anomaly detection device is provided, the device 400 may include:

[0127] Image acquisition module 401 is used to acquire power images;

[0128] The image recognition module 402 is used to identify anomalies in power images by using a pre-trained power image anomaly detection model, and to identify abnormal image regions in the power images.

[0129] The text recognition module 403 is used to identify the matching degree between abnormal image regions and pre-stored abnormal description text through a pre-trained power image anomaly detection model, and obtain the matching degree recognition result.

[0130] The text selection module 404 is used to select the anomaly description text with the highest matching degree from the pre-stored anomaly description text based on the matching degree recognition result, and use it as the target anomaly description text for the anomaly image region.

[0131] The region identification module 405 is used to identify abnormal image regions in the power image by using the target abnormal description text through a pre-trained power image anomaly detection model, so as to obtain an abnormally identified power image.

[0132] The result determination module 406 is used to identify the abnormal power image as the anomaly detection result of the power image.

[0133] In one embodiment, the device 400 further includes: a model training module, configured to acquire a first sample power image and the corresponding real anomaly detection result; input the first sample power image into a power image anomaly detection model to be trained to obtain a predicted anomaly detection result of the first sample power image; and train the power image anomaly detection model to be trained based on the difference between the predicted anomaly detection result and the real anomaly detection result to obtain a pre-trained power image anomaly detection model.

[0134] In one embodiment, the model training module is further configured to: train the power image anomaly detection model to be trained based on the difference between the predicted anomaly detection results and the actual anomaly detection results to obtain a basic power image anomaly detection model; determine a first power image anomaly detection model and a second power image anomaly detection model corresponding to the basic power image anomaly detection model based on the model parameters and model structure of the basic power image anomaly detection model; input an unlabeled second sample power image into the first power image anomaly detection model to obtain a pseudo-label for the second sample power image; and train the second power image anomaly detection model using the second sample power image and the pseudo-label of the second sample power image to obtain a pre-trained power image anomaly detection model.

[0135] In one embodiment, the model training module is further configured to acquire abnormal images of power equipment and abnormal images of personnel behavior within the substation area; and combine the abnormal images of power equipment and abnormal images of personnel behavior to obtain a first sample power image.

[0136] In one embodiment, the image acquisition module 401 is further configured to acquire raw power images within the substation area; the raw power images include images of power equipment and images of personnel behavior; and to perform image preprocessing on the raw power images to obtain the power images.

[0137] In one embodiment, the device 400 further includes: a text determination module, used to identify historical anomaly information within the substation area to obtain historical anomaly types within the substation area; identify historical anomaly types to obtain historical anomaly description text within the substation area; and use the historical anomaly description text as pre-stored anomaly description text.

[0138] Each module in the aforementioned power image anomaly detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0139] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for detecting power image anomalies. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0140] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0141] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0142] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0143] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0145] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A power image anomaly detection method, characterized by, The method comprises: acquiring a power image in a substation area; the power image comprises a power equipment image and a personnel behavior image; performing abnormality identification on the power image by using a pre-trained power image abnormality detection model to identify an abnormal image area in the power image; the power image abnormality detection model is a GLIP model; identifying historical abnormality information in the substation area to obtain a historical abnormality type in the substation area; identifying the historical abnormality type to obtain a historical abnormality description text in the substation area; taking the historical abnormality description text as a pre-stored abnormality description text; the pre-stored abnormality description text comprises respirator oil seal damage, dial plate blurring, dial plate damage, insulator damage, ground oil stain, silica gel cylinder damage, cabinet door closure abnormality, hanging suspended matter, bird nest, cover plate damage, not wearing a safety helmet, not wearing work clothes, personnel smoking, respirator oil seal oil level abnormality, and / or silica gel discoloration; identifying a matching degree between the abnormal image area and the pre-stored abnormality description text by using the pre-trained power image abnormality detection model to obtain a matching degree identification result; according to the matching degree identification result, selecting a corresponding abnormality description text with the highest matching degree from the pre-stored abnormality description text as a target abnormality description text of the abnormal image area; identifying the target abnormality description text in an identification area corresponding to the abnormal image area in the power image by using the pre-trained power image abnormality detection model to obtain an abnormality identification power image; taking the abnormality identification power image as an abnormality detection result of the power image.

2. The method of claim 1, wherein, The pre-trained power image abnormality detection model is obtained by the following method: acquiring a first sample power image and a true abnormality detection result corresponding to the first sample power image; inputting the first sample power image into a power image abnormality detection model to be trained to obtain a predicted abnormality detection result of the first sample power image; training the power image abnormality detection model to be trained according to a difference between the predicted abnormality detection result and the true abnormality detection result to obtain the pre-trained power image abnormality detection model.

3. The method of claim 2, wherein, The pre-trained power image abnormality detection model is obtained by the following method: training the power image abnormality detection model to be trained according to a difference between the predicted abnormality detection result and the true abnormality detection result to obtain the pre-trained power image abnormality detection model. The pre-trained power image abnormality detection model is obtained by the following method: training the power image abnormality detection model to be trained according to a difference between the predicted abnormality detection result and the true abnormality detection result to obtain the pre-trained power image abnormality detection model. The pre-trained power image abnormality detection model is obtained by the following method: training the power image abnormality detection model to be trained according to a difference between the predicted abnormality detection result and the true abnormality detection result to obtain the pre-trained power image abnormality detection model. The second sample power image and the pseudo label of the second sample power image are used to train the second power image anomaly detection model, so as to obtain the pre-trained power image anomaly detection model.

4. The method of claim 2, wherein, The first sample power image comprises: The power equipment anomaly image and the personnel behavior anomaly image in the substation area are obtained. The power equipment anomaly image and the personnel behavior anomaly image are combined to obtain the first sample power image.

5. The method of claim 1, wherein, The power image in the substation area comprises: The original power image in the substation area is obtained; the original power image comprises a power equipment image and a personnel behavior image; The original power image is preprocessed to obtain the power image.

6. An electric power image abnormality detection device characterized by comprising: The device comprises: An image acquisition module is configured to acquire a power image in a substation area; the power image comprises a power equipment image and a personnel behavior image; An image recognition module is configured to identify an anomaly in the power image by using a pre-trained power image anomaly detection model to identify an anomaly image region in the power image; the power image anomaly detection model is a GLIP model; A text determination module is configured to identify historical anomaly information in the substation area to obtain a historical anomaly type in the substation area; the historical anomaly type is identified to obtain a historical anomaly description text in the substation area; the historical anomaly description text is used as a pre-stored anomaly description text; the pre-stored anomaly description text comprises a respirator oil seal damage, a dial blur, a dial damage, an insulator damage, a ground oil stain, a silica gel cylinder damage, a box door closure anomaly, a hanging suspended matter, a bird nest, a cover plate damage, an unsafe hat, an unsafe clothing, a personnel smoking, a respirator oil seal oil level anomaly, and / or a silica gel discoloration; A text recognition module is configured to identify a matching degree between the anomaly image region and the pre-stored anomaly description text by using the pre-trained power image anomaly detection model to obtain a matching degree recognition result; A text selection module is configured to select a highest matching degree anomaly description text from the pre-stored anomaly description text as a target anomaly description text of the anomaly image region according to the matching degree recognition result; A region identification module is configured to identify the target anomaly description text in an identification region corresponding to the anomaly image region in the power image by using the pre-trained power image anomaly detection model to obtain an anomaly identification power image; A result determination module is configured to use the anomaly identification power image as an anomaly detection result of the power image.

7. The apparatus of claim 6, wherein, The device further comprises a model training module configured to: acquire a first sample power image and a true anomaly detection result corresponding to the first sample power image; input the first sample power image into a power image anomaly detection model to be trained to obtain a predicted anomaly detection result of the first sample power image; and train the power image anomaly detection model to be trained according to a difference between the predicted anomaly detection result and the true anomaly detection result to obtain the pre-trained power image anomaly detection model.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor, when executing the computer program, implements the steps of the method of any one of claims 1 to 5.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5.

10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5. The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image recognition method and device, electronic equipment and storage medium

    CN115331150A

  • Text image matching model training method, picture labeling method, device and equipment

    CN115359492A

  • Industrial defect detection method and device based on pre-training model and storage medium

    CN116468725A