An inspection method and system for power grid system faults
By designing multi-stage algorithms and end-to-end training strategies in grid fault inspection, the problems of high computing power demand and low detection accuracy in the existing technology are solved, and the refined positioning and classification of grid faults are realized, and the detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510261070.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing power grid fault inspection methods rely on deep learning, and the computing power demand is high, making it difficult to realize real-time monitoring on low-computing power platforms. The lightweight detection algorithm is difficult to meet the needs of high accuracy, and it is easy to ignore special fault conditions.
A multi-stage algorithm is designed, combined with a histogram equalization algorithm for pixel gradient perception for pre-processing, a lightweight first-stage algorithm is used for fault judgment, and a fine second-stage algorithm is used for fault position detection and type judgment, and an end-to-end training strategy of comparison learning and multi-task collaboration is used to improve detection efficiency and accuracy.
It realizes the refined positioning and classification of faults in complex power grid scenarios, improves detection efficiency and accuracy, reduces calculation overhead and promotion costs, and enhances the ability to identify unknown types of faults.
Smart Images

Figure CN119763050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image data processing and power monitoring, and particularly to an inspection method and system for grid system faults. Background Art
[0002] The intelligent inspection of power facilities can efficiently monitor the real-time operation of equipment such as cable lines, and promptly respond to various faults to ensure the stable operation of the power grid. Due to the wide range of power grid erection and complex environment, traditional manual inspections are difficult to meet the requirements of high efficiency and high accuracy, and also face huge labor costs and time expenditures. Therefore, the power grid system is gradually introducing intelligent inspection technology, which analyzes images captured by fixed cameras or inspection drones through artificial intelligence algorithms to autonomously detect faults in power facilities and notify relevant operation and maintenance personnel.
[0003] However, existing methods rely on deep learning to construct detection models, which have high computing power requirements and are difficult to achieve real-time monitoring on low-computing power platforms. Lightweight detection algorithms are difficult to meet the requirements of high accuracy. In addition, existing methods mainly detect known types of faults and are prone to ignoring some special fault situations.
[0004] Therefore, there is an urgent need to develop a solution to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide an inspection method and system for grid system faults. By designing a multi-stage algorithm and proposing an end-to-end training strategy based on contrast learning and multi-task collaboration, the efficiency and accuracy of fault detection are effectively improved, and fine-grained positioning and classification of faults in complex grid scenarios are realized.
[0006] The inspection method and system for grid system faults provided by the present invention adopt the following technical solutions:
[0007] In a first aspect, an inspection method for grid system faults specifically includes: obtaining an image to be detected of a corresponding power grid line at the current moment, and simultaneously extracting images at different previous moments at the current position and images of adjacent line positions at the current moment, and preprocessing them using a histogram equalization algorithm based on pixel gradient perception.
[0008] Based on the preprocessed image to be detected and images in different time and space domains, apply a lightweight first-stage algorithm for fault discrimination to determine whether there are abnormalities in the image to be detected.
[0009] For the to-be-detected images with anomalies and the corresponding images in different time and space domains, apply the refined second-stage algorithm for fault location detection and fault type discrimination. For the faults determined to be of unknown types, use the second-stage algorithm to extract the features of the faults, and perform clustering analysis together with the features of the unknown faults recorded in the database;
[0010] Based on the results of fault location detection, fault type discrimination, and clustering analysis, conduct comprehensive analysis to obtain the fault conditions of the to-be-detected images and output an inspection report, thus completing the intelligent inspection of power grid faults at the current location.
[0011] Optionally, perform preprocessing on it using the histogram equalization algorithm based on pixel gradient perception, including:
[0012] Calculate the gradient intensity of each pixel in the to-be-detected image and generate a gradient weight map;
[0013] Divide the to-be-detected image evenly into image blocks along the horizontal and vertical axes, and perform equalization processing on each image block separately;
[0014] Based on the gradient weight map, weighted-fuse the original and equalized images, suppress the influence of light interference while protecting visual details, and finally obtain the image pixel values to complete the preprocessing.
[0015] Optionally, apply the lightweight first-stage algorithm for fault discrimination to determine whether there are anomalies in the to-be-detected images. The lightweight first-stage algorithm adopts a training strategy of multi-branch collaboration. The specific training strategy includes:
[0016] Perform image matching and calibration on the to-be-detected image and the images in different time and space domains respectively: first extract the ORB feature points of the two images respectively, then match the feature points between the two images and filter out the wrong matching point pairs, and finally calculate the perspective transformation matrix mapped from the images in different time and space domains to the coordinate system of the to-be-detected image and perform image transformation;
[0017] For the to-be-detected image and the calibrated images in different time and space domains, extract three types of visual features respectively; among them, the three types of visual features include corner features, color change features, and global texture features;
[0018] For the three types of visual features extracted, send them into a detector composed of a convolutional layer, a fully connected layer, and an activation function layer. The detector first uses the convolutional layer to map various visual features, and then performs binary classification prediction based on a single visual feature through the fully connected layer and the activation function layer to determine whether there are anomalies in the to-be-detected image, and calculates the training loss using the cross-entropy function;
[0019] Splice together the mapped various visual features, and then predict the probability that the device in the currently to-be-predicted image has a fault through a fully connected layer and an activation function layer, completing the lightweight first-stage algorithm.
[0020] Optionally, for the to-be-detected image and the calibrated different spatio-temporal images, three types of visual features are respectively extracted, including:
[0021] When extracting corner features, calculate the feature point response scores in the to-be-detected image and sort them in descending order, select the top K ORB feature points. If the number of feature points in the to-be-detected image is less than K, fill them with a vector all of whose values are 0; where K is a natural number greater than 1.
[0022] When extracting color change features, respectively count the color histograms of the to-be-detected image and different temporal or spatial images, calculate the color histogram distance between the two images. When the to-be-detected image is an RGB image, respectively count the three color channels and calculate the color histogram distance.
[0023] When extracting global texture features, calculate the correlation between corner features pairwise, and construct global texture features based on the calculation results of the correlation.
[0024] Optionally, apply a refined second-stage algorithm for fault location detection and fault type discrimination. The refined second-stage algorithm adopts an end-to-end training strategy of contrastive learning and multi-task collaboration. The specific training includes:
[0025] For the input to-be-detected image and different spatio-temporal images, respectively extract visual features through the visual encoder of the pre-trained CLIP model.
[0026] Construct a text description statement based on keywords related to power facilities, and extract text features based on the text encoder of the pre-trained CLIP model; where the visual encoder and the text encoder maintain the state of frozen parameters.
[0027] Perform cross-spatio-temporal query modeling based on the visual features and text features to obtain multi-modal query features that simultaneously contain the visual information of the current scene and the text prior related to power facilities.
[0028] Apply matrix multiplication to the multi-modal query features and the visual features to obtain a preliminary fault location mask, and send the multi-modal query features into a classifier composed of a fully connected layer and an activation layer to obtain a fault category probability prediction distribution.
[0029] Calculate the loss based on the preliminary fault location mask and the predicted distribution of fault categories, design a skeleton-aware positive and negative sample balance loss function, and construct a contrastive learning loss function across time and space domains based on the visual features to narrow the visual representation between the image to be detected and images in different time and space domains, and distinguish the visual representation differences between images in different time domains and different space domains;
[0030] Based on the preliminary fault location mask, combine the multi-modal query features for mask iterative update until the maximum number of iterations is reached to end the training and complete the refined second-stage algorithm training.
[0031] Optionally, and / or, the formula of the positive and negative sample balance loss function is as follows:
[0032] ;
[0033] ;
[0034] In the formula, is the positive and negative sample balance loss, N is the number of pixels in the image, and are the numbers of pixels belonging to positive and negative samples in the image respectively, and are the true label and the predicted probability distribution of the th pixel respectively, is the number of pixels in the skeleton area;
[0035] The formula of the contrastive learning loss function across time and space domains is as follows:
[0036] ;
[0037] ;
[0038] ;
[0039] ;
[0040] In the formula, , and respectively represent the visual features extracted from the image to be detected, images in different time domains, and images in different space domains in sequence, is the contrastive learning loss function.
[0041] Optionally, perform cross-time and space domain query modeling based on the visual features and text features to obtain multi-modal query features that simultaneously contain the visual information of the current scene and the text prior related to power facilities, including:
[0042] Feature extraction obtains text features, concatenates the visual features of the image to be detected and images in different time and space domains, and combines the text features and the concatenated visual features based on the cross-attention mechanism to obtain multi-modal query features;
[0043] The calculation formula of the multi-modal query features is as follows:
[0044] ;
[0045] ;
[0046] In the formula, is the query vector, is the key vector, is the value vector, , and represent learnable weight matrices, represents the number of channels of the features, represents the Softmax function, is the text feature, is the visual feature obtained by concatenating the visual features of the image to be detected and images in different time and space domains, is the multi-modal query feature, is the transpose matrix of.
[0047] Optionally, based on the preliminary fault location mask, the multi-modal query features are iteratively updated by masking until the maximum number of iterations is reached to end the training, including:
[0048] Multiply the preliminary fault location mask with the features of the image to be detected using matrix multiplication to obtain the fault region features;
[0049] Based on the cross-attention mechanism, interact the multi-modal query features with the fault region features to update the multi-modal query features;
[0050] Multiply the updated multi-modal query features with the features of the image to be detected using matrix multiplication to obtain the updated fault location mask, and at the same time send the multi-modal query features into a classifier composed of a fully connected layer and an activation layer to obtain the updated fault category probability prediction distribution;
[0051] Iteratively update continuously until the maximum number of iterations to obtain the final fault location mask and fault category probability distribution.
[0052] The beneficial effects of an inspection method for power grid system faults provided by the present invention are as follows:
[0053] (1)High detection efficiency: By constructing a multi-stage detection algorithm, the present invention uses a lightweight first-stage algorithm to quickly pre-screen a large amount of image data, thereby significantly reducing the overall computational overhead. Moreover, the first-stage algorithm can be deployed on most low-computing-power platforms and achieve real-time detection, further reducing the promotion and application costs of the present invention;
[0054] (2)High detection accuracy: The present invention designs an equalization algorithm for pixel gradient perception for the characteristics of power grid images, which suppresses light interference while retaining key visual cues. In addition, the present invention simultaneously introduces spatio-temporal priors in the lightweight first-stage algorithm and designs a training strategy for multi-branch collaboration, so that it has better fault discovery capabilities. Then, the present invention constructs a refined second-stage algorithm, and through iterative updating of the fault mask, realizes precise detection of small-scale and irregularly shaped fault regions;
[0055] (3)High robustness: The present invention introduces visual language priors in the refined second-stage algorithm and constructs an end-to-end training strategy based on contrast learning and multi-task collaboration, thereby effectively improving the algorithm's feature modeling ability and the ability to discover faults in various complex environments, and improving the robustness of the algorithm. In addition, the present invention performs clustering analysis on unknown types of faults, improving the practicality of the algorithm.
[0056] In the second aspect, an inspection system for power grid system faults specifically includes:
[0057] An image acquisition and preprocessing module, which is used to obtain the image to be detected of the corresponding power grid line at the current moment, and at the same time extract the images at different previous moments at the current position and the images of adjacent line positions at the current moment, and perform preprocessing on them using the histogram equalization algorithm for pixel gradient perception;
[0058] An anomaly detection module, which based on the preprocessed image to be detected and images in different spatio-temporal domains, applies a lightweight first-stage algorithm to perform fault discrimination and determine whether there is an anomaly in the image to be detected;
[0059] A discrimination and clustering module, for the image to be detected with anomalies and the corresponding images in different spatio-temporal domains, applies a refined second-stage algorithm to perform fault location detection and fault type discrimination. For faults determined to be of unknown types, the second-stage algorithm is used to extract the features of the fault, and clustering analysis is performed together with the features of unknown faults recorded in the database;
[0060] An aggregation and analysis module, which based on the results of fault location detection, fault type discrimination and clustering analysis, performs comprehensive analysis to obtain the fault situation of the image to be detected and outputs an inspection report, completing the intelligent inspection of power grid faults at the current position.
[0061] The beneficial effects of the second aspect can be referred to the description of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a flowchart of the overall steps of a method for inspecting faults in a power grid system provided by the present invention;
[0063] Figure 2 It is a flowchart of the fault discrimination algorithm in the first stage of a method for inspecting faults in a power grid system provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meaning as understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects.
[0065] See Figure 1 , the embodiments of the present invention provide a method for inspecting faults in a power grid system, including the following steps:
[0066] S1. Obtain the image to be detected of the corresponding power grid line at the current moment, and at the same time extract the images at different previous moments at the current position and the image of the adjacent line position at the current moment, and preprocess them using the histogram equalization algorithm based on pixel gradient perception;
[0067] S2. Based on the preprocessed image to be detected and the images in different time and space domains, apply a lightweight first-stage algorithm for fault discrimination to determine whether there is an abnormality in the image to be detected;
[0068] S3. For the image to be detected with an abnormality and the corresponding images in different time and space domains, apply a refined second-stage algorithm for fault location detection and fault type discrimination. For the faults determined to be of unknown type, use the second-stage algorithm to extract the features of the fault and perform clustering analysis together with the features of the unknown faults recorded in the database;
[0069] S4. Based on the results of fault location detection, fault type discrimination, and clustering analysis, perform comprehensive analysis to obtain the fault situation of the image to be detected and output an inspection report, completing the intelligent inspection of the power grid fault at the current position.
[0070] In some embodiments, when performing step S1, it specifically includes:
[0071] S1-1. Obtain the image to be detected of the corresponding power grid line at the current moment, and at the same time extract the images of the current position at different previous moments, as well as the images of the adjacent line positions at the current moment;
[0072] S1-2. Perform preprocessing on it using the histogram equalization algorithm based on pixel gradient perception.
[0073] Specifically, when performing step S1-1, remote monitoring of the power grid line is achieved through the cloud platform or mobile application. The operation and maintenance personnel can perform operations such as capturing, video browsing, and remote configuration on the line through the background management software, obtain the image to be detected of the corresponding power grid line at the current moment through the cloud platform, and at the same time extract the images of the current position at different previous moments, as well as the images of the adjacent line positions at the current moment.
[0074] Specifically, when performing step S1-2, preprocessing is performed on it using the histogram equalization algorithm based on pixel gradient perception, including:
[0075] S1-2-1. Calculate the gradient intensity of each pixel in the image to be detected and generate a gradient weight map;
[0076] S1-2-2. Uniformly divide the image to be detected into image blocks along the horizontal and vertical axes, and perform equalization processing on each image block respectively;
[0077] S1-2-3. Based on the gradient weight map, weighted fusion of the original and equalized images is performed to suppress the influence of light interference while protecting visual details, and finally the image pixel values are obtained to complete the preprocessing.
[0078] Actually, when performing S1-2-1, denote the pixel at row column as , and its gradient The calculation formula is:
[0079]
[0080] Furthermore, denote the gradient weight of the pixel at row column as , and the specific calculation formula of the gradient weight map is:
[0081] ;
[0082] In the formula, represents the maximum value of the gradient intensity in the graph, which is used to normalize the gradient.
[0083] Actually, when executing S1-2-2, the equalized image pixels are denoted as , and its specific calculation formula is:
[0084] ;
[0085] ;
[0086] In the formula, represents the local contrast value of the pixel in the row and column, and respectively represent the mean and variance of the pixels in the current image block,
[0087]
[0088] ;
[0089] In the formula, the parameters are the same as above;
[0090] In some embodiments, referring to Figure 2 , when executing step S2, based on the preprocessed image to be detected and images in different time and space domains, a lightweight first-stage algorithm is applied for fault discrimination to determine whether there is an abnormality in the image to be detected. The lightweight first-stage algorithm adopts a multi-branch collaborative training strategy, and the specific training strategy includes:
[0091] S2-1. Perform image matching and calibration on the image to be detected and images in different time and space domains respectively: First, extract the ORB feature points of the two images respectively, then match the feature points between the two images and filter out the wrong matching point pairs, and finally calculate the perspective transformation matrix mapped from the images in different time and space domains to the coordinate system of the image to be detected and perform image transformation;
[0092] S2-2. For the image to be detected and the calibrated images in different time and space domains, three types of visual features are respectively extracted;
[0093] Further, the three types of visual features include corner features, color change features, and global texture features;
[0094] S2-3. For the three types of visual features extracted, send them into a detector composed of a convolutional layer, a fully connected layer, and an activation function layer. The detector first uses the convolutional layer to map each type of visual feature, and then based on a single visual feature, performs binary classification prediction through the fully connected layer and the activation function layer to determine whether there are abnormalities in the image to be detected, and calculates the training loss using the cross-entropy function;
[0095] S2-4. Concatenate the mapped visual features of each type together, and then predict the probability that the device in the current image to be predicted has a fault through the fully connected layer and the activation function layer, completing the lightweight first-stage algorithm.
[0096] Specifically, when performing step S2-1, use the extracted ORB features as corner features.
[0097] Specifically, when performing step S2-2, for the image to be detected and the calibrated images in different time and space domains, extract three types of visual features respectively, including:
[0098] S2-2-1. When extracting corner features, calculate the response scores of the feature points in the image to be detected and sort them in descending order, select the top K ORB feature points. If the number of feature points in the image to be detected is less than K, fill them with a vector of all zeros; where K is a natural number greater than 1;
[0099] S2-2-2. When extracting color change features, respectively calculate the color histograms of the image to be detected and the images in different time domains or different space domains, calculate the color histogram distance between the two images. When the image to be detected is an RGB image, calculate the color histogram distance by separately counting the three color channels.
[0100] S2-2-3. When extracting global texture features, calculate the correlation between corner features pairwise, and construct global texture features based on the calculation results of the correlation.
[0101] Actually, when performing step S2-2-1, considering that the ORB features extracted in step S2-1 can better represent the corner features around the feature points, and at the same time to reduce the computational cost, directly use the ORB features extracted in S2-1 as corner features.
[0102] Actually, when performing step S2-2-2, the calculation formula for the color histogram distance is as follows:
[0103] ;
[0104] In the formula, and are the means of the two histograms respectively, N is the number of intervals of the histogram, and They are the values of the two histograms in the th interval respectively.
[0105] Actually, when performing step S2-2-3, considering that the corner features in S2-2-1 can already reflect the local texture information to a certain extent, the global texture feature is further constructed based on the corner features, where is the number of corner feature points in the image, represents the th corner feature and the th corner feature, and the specific formula is as follows:
[0106] ;
[0107] ;
[0108] In the formula, and represent the th and the th corner features respectively, and represent the horizontal and vertical axis coordinates of the pixels corresponding to the th corner respectively, represents the maximum distance between the corner pixels in the current image;
[0109] Furthermore, when the distance between two corners is smaller and the orientation angle between the corner features is smaller, the correlation between them is greater, and vice versa.
[0110] Specifically, when performing step S2-3, the cross-entropy function is used to calculate the training loss, and the specific formula is as follows:
[0111] ;
[0112] In the formula, represents the sample label, , and represent the probability distributions predicted based on color, corner, and global texture features respectively.
[0113] Furthermore, when performing step S2-4, the cross-entropy function is also used to calculate the training loss, and the formula is the same as above.
[0114] In some embodiments, when performing step S3, it includes:
[0115] S3-1. For the images to be detected with abnormalities and the corresponding images in different time and space domains, the refined second-stage algorithm is applied to detect the fault location and identify the fault type;
[0116] S3-2. For a fault that is identified as an unknown type, the second-stage algorithm is used to extract the features of the fault, and cluster analysis is performed together with the features of the unknown faults recorded in the database.
[0117] Specifically, when executing step S3-1, a refined second-stage algorithm is applied to perform fault location detection and fault type discrimination. The refined second-stage algorithm adopts an end-to-end training strategy of contrastive learning and multi-task collaboration. The specific training includes:
[0118] S3-1-1, for the input image to be detected and images in different time and space domains, respectively extract visual features through the visual encoder of the pre-trained CLIP model;
[0119] S3-1-2, constructing a text description sentence based on keywords related to power facilities, and extracting text features based on a text encoder of a pre-trained CLIP model; wherein the visual encoder and the text encoder keep parameters frozen;
[0120] S3-1-3, based on the visual features and text features, cross-temporal and spatial query modeling is performed to obtain multimodal query features that simultaneously include current scene visual information and power facility-related text priors;
[0121] S3-1-4, applying matrix multiplication to the multimodal query feature and the visual feature to obtain a preliminary fault location mask, and sending the multimodal query feature to a classifier composed of a fully connected layer and an activation layer to obtain a fault category probability prediction distribution;
[0122] S3-1-5. Based on the preliminary fault location mask and the predicted distribution of fault category probability, the loss is calculated, the positive and negative sample balance loss function of skeleton perception is designed, and the contrast learning loss function across time and space domains is constructed based on the visual features to close the visual representation between the image to be detected and the images in different time and space domains, and distinguish the differences in visual representation between images in different time domains and different space domains;
[0123] S3-1-6. Based on the preliminary fault location mask and combined with the multimodal query features, the mask is iteratively updated until the maximum number of iterations is reached and the training is terminated, completing the refined second-stage algorithm training.
[0124] In fact, when performing steps S3-1-1 and S3-1-2, the pre-trained CLIP (Contrastive Language-Image Pre-Training) model is a multi-modal pre-trained neural network model composed of two parts: an image encoder and a text encoder. The image encoder is responsible for converting images into feature vectors, which can be a convolutional neural network (such as ResNet) or a Transformer model (such as ViT). The text encoder is responsible for converting text into feature vectors, usually a Transformer model. These two encoders achieve cross-modal information interaction and fusion by sharing a vector space.
[0125] In fact, when performing step S3-1-3, cross-temporal and spatio-temporal query modeling is performed based on the visual features and text features to obtain multi-modal query features that simultaneously contain the visual information of the current scene and the text prior related to power facilities, including:
[0126] Text features are extracted, the visual features of the image to be detected and images in different temporal and spatial domains are concatenated, and the text features and the concatenated visual features are combined based on the cross-attention mechanism to obtain multi-modal query features.
[0127] Furthermore, the calculation formula for the multi-modal query features is as follows:
[0128] ;
[0129] ;
[0130] In the formula, is the query vector, is the key vector, is the value vector, , and represent learnable weight matrices, represents the number of channels of the features, represents the Softmax function, is the text feature, is the visual feature obtained by concatenating the visual features of the image to be detected and images in different temporal and spatial domains, is the multi-modal query feature, is 's transposed matrix. By converting text information into feature , converting visual information into feature and feature , and then through matrix multiplication by , , calculate the query feature , to represent the possible defect information in the image to be detected.
[0131] Actually, when performing step S3-1-4, the multi-modal query features and the visual features of the image to be detected are applied with matrix multiplication to obtain a preliminary fault location mask . At the same time, the multi-modal query features are fed into a classifier composed of a fully connected layer and an activation layer to obtain a preliminary fault class probability prediction distribution. Among them, represents the height and width of the image feature map, represents the number of feature channels of the image feature map.
[0132] Actually, when performing step S3-1-5, the losses are calculated for the predicted preliminary fault location mask and fault class probability prediction distribution. Considering that the fault areas of power facilities are generally small in size or show irregular distributions, the present invention designs a skeleton-aware positive and negative sample balance loss function for mask prediction. The skeleton area is obtained by applying a skeleton extraction algorithm to the mask label. For the probability prediction distribution of the fault class, a cross-entropy loss function is used to calculate the training loss;
[0133] Furthermore, the formula of the positive and negative sample balance loss function is as follows:
[0134] ;
[0135] ;
[0136] In the formula, is the positive and negative sample balance loss, N is the number of pixels in the image, and are the numbers of pixels belonging to positive and negative samples in the image respectively, and are the true label and the predicted probability distribution of the -th pixel respectively, is the number of pixels in the skeleton area.
[0137] Actually, when performing step S3-1-5, a cross-temporal and cross-spatial contrast learning loss is constructed for the visual features extracted from the image to be detected and images in different time and space domains, so as to help the algorithm better model the visual correlation between images in different time domains and different space domains, and further better discover the fault areas in the images;
[0138] Furthermore, the contrastive learning loss function across time and space aims to narrow the visual representations between the image to be detected and images in different time and space domains, while ensuring that the algorithm can distinguish the visual representation differences between images in different time domains and different space domains. The formula for the contrastive learning loss function across time and space is as follows:
[0139] ;
[0140] ;
[0141] ;
[0142] ;
[0143] In the formula, , and respectively represent the visual features extracted from the image to be detected, images in different time domains, and images in different space domains in sequence. is the contrastive learning loss function.
[0144] Actually, when executing S3-1-6, based on the preliminary fault location mask and the multi-modal query features, the mask is iteratively updated until the maximum number of iterations is reached to end the training, including:
[0145] Multiply the matrix of the preliminary fault location mask and the features of the image to be detected to obtain the fault area features;
[0146] Based on the cross-attention mechanism, interact the multi-modal query features with the fault area features to update the multi-modal query features;
[0147] Multiply the matrix of the updated multi-modal query features and the features of the image to be detected to obtain the updated fault location mask, and at the same time send the multi-modal query features into the classifier composed of the fully connected layer and the activation layer to obtain the updated fault category probability prediction distribution;
[0148] Iteratively update continuously until the maximum number of iterations to obtain the final fault location mask and the fault category probability distribution.
[0149] Furthermore, it should be noted that in the inference and deployment stage of the second-stage algorithm, the final prediction result is directly used. For the masks and probability distributions predicted during the iteration process, the loss function described in S3-1-5 is also applied to calculate the training loss. Therefore, the complete training loss is as follows:
[0150] ;
[0151] In the formula, and represents the loss calculated based on the predicted probability distribution and the mask. represents the contrastive learning loss across time and space domains.
[0152] Specifically, when performing step S3-2, the database includes: a relational database, a NoSQL database, and a time series database. The relational database is used to store and manage various types of fault data, including known and unknown fault characteristics. The NoSQL database is used to store and manage the massive data generated by smart grids and Internet of Things devices. The time series database is used to store and manage the real-time operation data of the power grid.
[0153] In some embodiments, when performing step S4, a comprehensive analysis is performed based on the results of fault location detection, fault type discrimination, and clustering analysis to obtain the fault situation of the image to be detected and output an inspection report, thereby completing the intelligent inspection of the power grid fault at the current location.
[0154] Specifically, the input image is preprocessed and multi-stage detected according to the two-stage fault inspection method to determine whether there is a fault. If there is a fault, the fault location is located and the fault type is output, and the unknown type of fault is classified and divided to obtain the inspection result information to generate an inspection report, which is communicated with the control center and transmitted to the operation and maintenance personnel.
[0155] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein may have other embodiments and can be implemented or realized in various ways.
Claims
1. A method for inspecting a power grid system fault, characterized in that: include: Obtain the image to be detected of the corresponding power grid line at the current moment, and extract the images of the current position at different previous moments, as well as the image of the adjacent line position at the current moment, and pre-process them using a pixel gradient-aware histogram equalization algorithm; Based on the preprocessed image to be detected and images in different time and space domains, a lightweight first-stage algorithm is applied to perform fault discrimination to determine whether the image to be detected has an abnormality. The lightweight first-stage algorithm adopts a multi-branch collaborative training strategy. The specific training strategy includes: Perform image matching and calibration on the image to be detected and the images in different time and space domains respectively: first extract the ORB feature points of the two images respectively, then match the feature points between the two images and filter out the wrong matching point pairs, and finally calculate the perspective transformation matrix mapped from the images in different time and space domains to the coordinates of the image to be detected and perform image transformation; For the image to be detected and the calibrated images in different time and space domains, three types of visual features are extracted respectively; wherein the three types of visual features include corner point features, color change features and global texture features; The three types of extracted visual features are respectively sent to a detector consisting of a convolutional layer, a fully connected layer, and an activation function layer. The detector first uses the convolutional layer to map various visual features, and then uses the fully connected layer and the activation function layer to perform a binary classification prediction based on a single visual feature to determine whether the image to be detected has an abnormality, and uses the cross entropy function to calculate the training loss; The mapped visual features are stitched together, and then the probability of a device failure in the current image to be predicted is predicted through the fully connected layer and the activation function layer, completing the lightweight first-stage algorithm. For the images to be detected with abnormalities and the corresponding images in different time and space domains, the visual features are extracted by the visual encoder of the pre-trained CLIP model for the input images to be detected and the images in different time and space domains. Constructing a text description sentence based on keywords related to power facilities, and extracting text features based on a text encoder of a pre-trained CLIP model; wherein the visual encoder and the text encoder keep parameters frozen; Based on the visual features and text features, query modeling across time and space is performed to obtain multimodal query features that simultaneously include current scene visual information and text priors related to power facilities; The multimodal query feature is multiplied by the visual feature using matrix multiplication to obtain a preliminary fault location mask, and the multimodal query feature is sent to a classifier composed of a fully connected layer and an activation layer to obtain a fault category probability prediction distribution; Based on the preliminary fault location mask and the predicted distribution of fault category probability, the loss is calculated, a skeleton-aware positive and negative sample balance loss function is designed, and a cross-temporal and spatial contrast learning loss function is constructed based on the visual features to narrow the visual representation between the image to be detected and the images in different temporal and spatial domains, and distinguish the differences in visual representation between images in different temporal domains and different spatial domains; Based on the preliminary fault location mask and the features of the image to be detected, matrix multiplication is used to obtain the fault area features; Based on the cross-attention mechanism, the multimodal query features are interacted with the fault area features to update the multimodal query features; The updated multimodal query features are multiplied by the features of the image to be detected using matrix multiplication to obtain an updated fault location mask. At the same time, the multimodal query features are sent to a classifier consisting of a fully connected layer and an activation layer to obtain an updated fault category probability prediction distribution. The algorithm is continuously updated until the maximum number of iterations is reached to obtain the final fault location mask and fault category probability distribution, completing the refined second-stage algorithm training. For faults that are identified as unknown types, the second-stage algorithm is used to extract the features of the fault and perform cluster analysis together with the features of unknown faults recorded in the database. A comprehensive analysis is performed based on the results of fault location detection, fault type discrimination and cluster analysis to obtain the fault condition of the image to be detected and output an inspection report to complete the intelligent inspection of the power grid fault at the current location.
2. A method for inspecting a power grid system fault according to claim 1, characterized in that: The pixel gradient-aware histogram equalization algorithm is used for preprocessing, including: Calculate the gradient strength of each pixel in the image to be detected and generate a gradient weight map; The image to be detected is evenly divided into image blocks along the horizontal axis and the vertical axis, and each image block is subjected to equalization processing respectively; The original and equalized images are weightedly fused based on the gradient weight map to suppress the influence of light interference while protecting the visual details, and finally the image pixel value is obtained to complete the preprocessing.
3. A method for inspecting a power grid system fault according to claim 1, characterized in that: For the image to be detected and the calibrated images in different time and space domains, three types of visual features are extracted respectively, including: When extracting corner features, calculate the response scores of the feature points in the image to be detected and sort them in descending order, select the first K ORB feature points, and if the number of feature points in the image to be detected is less than K, fill them with vectors with all values zero; K is a natural number greater than 1; When extracting color change features, the color histograms of the image to be detected and the images in different time domains or different spatial domains are counted respectively, and the color histogram distance between the two images is calculated. When the image to be detected is an RGB image, the three color channels are counted respectively and the color histogram distance is calculated; When extracting global texture features, the correlation between corner point features is calculated pairwise, and the global texture features are constructed based on the calculation results of the correlation.
4. A method for inspecting a power grid system fault according to claim 1, characterized in that: And / or, the positive and negative sample balance loss function formula is as follows: ; ; In the formula, is the loss of positive and negative samples balance, is the loss weight coefficient, N is the number of pixels in the image, and are the number of pixels in the image that belong to positive and negative samples, respectively. and They are The true label and predicted probability distribution of pixels, is the number of pixels in the skeleton area; The formula for the contrastive learning loss function across time and space is as follows: ; ; ; ; In the formula, , and They represent the visual features extracted from the image to be detected, images in different time domains, and images in different spatial domains, respectively. is the contrastive learning loss function.
5. A method for inspecting a power grid system fault according to claim 1, characterized in that: Based on the visual features and text features, query modeling across time and space is performed to obtain multimodal query features that contain both current scene visual information and text priors related to power facilities, including: Feature extraction obtains text features, splices the visual features of the image to be detected and the images in different time and space domains, and combines the text features and the spliced visual features based on the cross-attention mechanism to obtain multimodal query features; The formula for calculating multimodal query features is as follows: ; ; In the formula, is the query vector, is the key vector, is a value vector, , and represents the learnable weight matrix, The number of channels representing features, represents the Softmax function, Text features, It is the visual feature obtained by concatenating the visual features of the image to be detected and the images in different time and space domains. is the multimodal query feature, for The transposed matrix of .
6. A power grid system fault inspection system, characterized in that: include: The image acquisition preprocessing module is used to obtain the image to be detected of the corresponding power grid line at the current moment, and at the same time extract the images of the current position at different previous moments, as well as the images of the adjacent line position at the current moment, and preprocess them using a pixel gradient-aware histogram equalization algorithm; The anomaly detection module uses a lightweight first-stage algorithm to perform fault discrimination based on the preprocessed image to be detected and the images in different time and space domains to determine whether the image to be detected has an anomaly. The lightweight first-stage algorithm adopts a multi-branch collaborative training strategy. The specific training strategy includes: performing image matching and calibration on the image to be detected and the images in different time and space domains respectively: first extracting the ORB feature points of the two images respectively, then matching the feature points between the two images and filtering out the wrong matching point pairs, and finally calculating the perspective transformation matrix mapped from the images in different time and space domains to the coordinates of the image to be detected and performing image transformation; for the image to be detected and the calibrated images in different time and space domains, respectively Three types of visual features are obtained; wherein the three types of visual features include corner features, color change features and global texture features; the three types of extracted visual features are respectively sent to a detector consisting of a convolutional layer, a fully connected layer and an activation function layer, the detector first uses the convolutional layer to map various types of visual features, and then uses the fully connected layer and the activation function layer to perform a binary classification prediction based on a single visual feature to determine whether the image to be detected has an abnormality, and uses a cross entropy function to calculate the training loss; the mapped various types of visual features are spliced together, and then the probability of a device failure in the current image to be predicted is predicted through the fully connected layer and the activation function layer, completing the lightweight first-stage algorithm; The discriminant clustering module extracts visual features of the input images to be detected and the corresponding images in different time and space domains through the visual encoder of the pre-trained CLIP model respectively; constructs text description sentences based on keywords related to power facilities, and extracts text features based on the text encoder of the pre-trained CLIP model; wherein the visual encoder and the text encoder keep the parameters frozen; performs query modeling across time and space domains based on the visual features and text features, and obtains multimodal query features that simultaneously contain the current scene visual information and the text priors related to power facilities; multiplies the multimodal query features with the visual features using matrix multiplication to obtain a preliminary fault location mask, and sends the multimodal query features to a classifier composed of a fully connected layer and an activation layer to obtain a fault category probability prediction distribution; calculates the loss based on the preliminary fault location mask and the fault category probability prediction distribution, designs a skeleton-aware positive and negative sample balance loss function, and uses the visual features to calculate the fault location mask and the fault category probability prediction distribution. The feature constructs a cross-temporal and spatial contrast learning loss function to bring the visual representation between the image to be detected and the images in different temporal and spatial domains closer, and distinguishes the visual representation differences between images in different temporal and spatial domains; the preliminary fault location mask is multiplied with the features of the image to be detected by matrix multiplication to obtain the fault area features; based on the cross-attention mechanism, the multimodal query features are interacted with the fault area features to update the multimodal query features; the updated multimodal query features are multiplied with the features of the image to be detected by matrix multiplication to obtain an updated fault location mask, and the multimodal query features are sent to a classifier composed of a fully connected layer and an activation layer to obtain an updated fault category probability prediction distribution; the iteration update is continuously performed until the maximum number of iterations is reached to obtain the final fault location mask and fault category probability distribution, and the refined second-stage algorithm training is completed; for faults that are identified as unknown types, the second-stage algorithm is used to extract the features of the fault, and cluster analysis is performed together with the features of unknown faults recorded in the database; The cluster analysis module performs a comprehensive analysis based on the results of fault location detection, fault type identification and cluster analysis, obtains the fault condition of the image to be detected and outputs an inspection report, completing the intelligent inspection of the power grid fault at the current location.
Citation Information
Patent Citations
Online grid fault detection method based on relative protection entropy and nominal transition resistance
CN104316836A
Power grid inspection method and device, computer equipment and storage medium
CN116188370A
Power grid unmanned aerial vehicle inspection system and method for long-distance communication
CN116935245A
Point cloud quality evaluation method based on projection and multi-scale features and related device
CN117372325A
Fault studying and judging method and system based on multi-mode power grid operation characteristics
CN117648443A