Power fault identification method and system based on small sample target detection

By preprocessing and enhancing the feature extraction of power fault data, combining interactive matching to generate aggregated features, and using the YOLOv1n detector, the problem of insufficient real-time performance in power fault identification is solved, and the detection accuracy and speed are improved.

CN121053441APending Publication Date: 2025-12-02CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +3
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511142551.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing small-sample target detection technologies cannot meet the real-time requirements in power fault identification. Traditional methods have time delay bottlenecks and cannot be effectively applied to power inspection.

Method used

A power fault identification method based on small sample target detection is adopted. By preprocessing power equipment fault data and image label data, query data and support data are generated. Feature extraction and feature enhancement modules are used to generate aggregated features. Interactive matching is combined to characterize fault co-occurrence patterns. YOLOv1n is used as the basic detector, and the detection speed and accuracy are improved through anchorless design.

Benefits of technology

While maintaining model speed, the detection performance of small sample fault categories has been improved, meeting the real-time requirements of power line inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053441A_ABST
    Figure CN121053441A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and discloses a power fault identification method and system based on small sample target detection, and the method comprises the steps: carrying out the preprocessing of collected power equipment fault data and locally pre-stored image label data, and obtaining query data and support data; respectively and sequentially performing feature extraction and feature enhancement on the query data and the support data to respectively generate query features and support features; the query features and the support features are fused together through interactive matching, and aggregation features representing a fault co-occurrence mode are generated; and performing target detection according to the aggregation features, and outputting a power fault positioning and classification result. According to the method, feature extraction and bounding box regression are decoupled through anchor-point-free design, and the detection speed of the model is increased; in the aspect of small sample adaptation, a meta-learning module based on feature re-calibration is developed, fault feature expression is enhanced through a channel attention mechanism, and the detection precision of a small sample target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, and specifically relates to a power fault identification method and system based on small sample target detection. Background Technology

[0002] In the process of intelligent transformation of power systems, equipment operation and maintenance models are undergoing fundamental changes. Traditional maintenance methods relying on regular manual inspections are no longer adequate to meet the demands of modern power grid expansion. Their labor-intensive nature and reliance on subjective experience lead to bottlenecks in operation and maintenance efficiency. With breakthroughs in computer vision and deep learning technologies, the industry has begun to explore intelligent operation and maintenance solutions based on automated inspections, building equipment status awareness networks by deploying intelligent terminals such as visible light / infrared imaging devices and vibration sensors. However, power equipment faults exhibit a typical long-tail distribution, and many defects in actual operation are rare anomalies, causing traditional supervised learning models to fall into a "data hunger" dilemma due to sample scarcity. This small-sample learning challenge directly restricts the large-scale application of AI technology in the field of power inspection.

[0003] Existing few-shot object detection technologies mainly rely on two-stage detection frameworks, such as the Faster R-CNN series. These frameworks use a serial processing mechanism—generating candidate boxes and then classifying them through Region Proposal Networks (RPNs)—which, while ensuring detection accuracy to some extent, has fundamental flaws in real-time power system inspection scenarios and cannot meet the real-time requirements of power inspections. The inherent latency bottleneck of this algorithm architecture directly restricts the engineering application of few-shot learning techniques in the field of dynamic power fault identification. Summary of the Invention

[0004] The purpose of this invention is to provide a power fault identification method and system based on small sample target detection, so as to solve the problem that the existing technology cannot meet the real-time requirements of power inspection.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a power fault identification method based on small sample target detection, comprising: The collected power equipment fault data and locally stored image tag data are preprocessed to obtain query data and support data. The query data and supporting data are subjected to feature extraction and feature enhancement respectively, generating query features and supporting features respectively; The query features and supporting features are fused together through interactive matching to generate aggregate features that represent the co-occurrence pattern of faults. Target detection is performed based on aggregated features, and the results of power fault location and classification are output.

[0006] Furthermore, the collected power equipment fault data and locally pre-stored image tag data include: A power fault dataset is collected. The faults in the dataset are labeled with their location and category through annotation and correction, resulting in a training dataset Data, containing a category set C. The frequency of each category of faults in the training dataset Data is counted, and a base class is selected based on a threshold th. and novel In the basic class, the number of times each type of fault appears in the training data set Data is greater than th; in the novel class, the number of times each type of fault appears in the training data set Data is less than or equal to th. Select samples corresponding to the basic classes from the training dataset Data to form the basic training set. , The number of images in the middle is Then, a support set dataset is constructed based on the basic training set; Select samples corresponding to novel classes from the training dataset Data to form a novel dataset. Approximately th samples are randomly selected from the training dataset Data for each base class and added to the fine-tuning set, which serves as the final fine-tuning set. ; The number of images in the middle is The class set is the complete set of classes. A support set dataset is constructed based on the fine-tuned training set; Locally pre-stored data is stored in a fixed location. Each category has a separate folder in this location, containing image files and label files for that category. Each label file contains only one fault marker belonging to that category, and categories in different folders are not duplicated. When the data acquisition module collects data, it notifies the local loading module to randomly load the support set data for that category from each folder under the fixed address.

[0007] Furthermore, the preprocessing process is as follows: The training dataset Data is processed as follows: Read each image one by one and generate a multidimensional array representation; adjust the image size while maintaining the image aspect ratio, and record the image scaling ratio. The image is processed by channel normalization, the color space is converted from RGB to BGR format, the mean of each channel is normalized, and the standard deviation is scaled to keep the original standard deviation of each channel. Perform boundary padding on the image, specifying the number of columns to pad each of the four sides of the image (top, bottom, left, and right). The data format is standardized, the input image is converted into a tensor of the deep learning framework, the image dimensions are rearranged, and finally, CPU / GPU memory is allocated according to the device GPU status. The image scaling ratio and edge padding are constructed into tensors as additional information for the image. The operation is repeated until all images are processed. For the locally stored image label data, perform the following operations: Select one data point from the locally stored image label data, perform channel normalization on the image, convert the color space from RGB to BGR format, normalize the mean of each channel, and scale the standard deviation to keep the original standard deviation of each channel. The image size is adjusted and a mask is generated. The original image is scaled and a mask data matching the image is generated. The image and the mask are concatenated on the channel to obtain the processed image. The data format is standardized by converting the processed image data into tensor format for the deep learning framework, including tensor dimension conversion and data type adaptation. The labels are converted into tensor format and normalized according to the original image size. This process is repeated until all data has been processed.

[0008] Furthermore, the step of extracting features from the query data and supporting data separately includes: Feature extraction is performed using a feature extraction module that includes 5 convolutional modules, 2 bottleneck modules, and 2 attention modules. After the query data and support data are input into the feature extraction module, the query feature FA and support feature FB are obtained. The attention module comprises 6 convolutional modules, 2 attention sub-modules, 4 addition modules, and 1 connection module. The attention sub-modules use the region attention mechanism proposed in YOLOv12, including: (1) Let the shape of the input x be B×C×H×W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions, and a is the number of regions. By performing a convolution operation on x, we obtain Q, K, V, and Pe: Q,K=Conv1(X), V=Conv2(X), Pe=Conv3(V) (2) Convert Q, K, and V into a multi-head form, with the number of heads being head, and divide the spatial dimension into a regions: Q=Reshape(Q,[B×a,head,C / head,N / a]) K=Reshape(K,[B×a,head,C / head,N / a]) V=Reshape(V,[B×a,head,C / head,N / a]) (3) Scale the dot product attention and normalize it using Softmax:

[0009] (4) Apply attention weights to Value and perform dimensional transformation on the output:

[0010] Z = Reshape(Z,[B,C,H,W]) (5) Weight the sum of the value Z and the position code Pe, and reconstruct the output using a convolution operation: Y = Conv3(Z + Pe) The obtained Y is used as the output of the attention submodule.

[0011] Furthermore, feature enhancement is performed on the query feature FA and supporting feature FB: The feature fusion module, a path aggregation network, employs a dual-path feature fusion approach: top-down to enhance semantic information and bottom-up to enhance detail information. Feature enhancement is achieved using a feature fusion module, which comprises two upsampling modules, four connection modules, three attention modules, and one bottleneck module. First, the feature F3 obtained from the backbone network is upsampled by 2 times using the upsampling module. Then, the upsampling result is concatenated with the feature F2 obtained from the backbone network using the connection module to perform channel-dimensional concatenation. Finally, the feature is refined through the attention module to generate the enhanced feature F1. The upsampling module is used to perform a 2x nearest neighbor upsampling on the enhanced feature F1. Then, the upsampling result is concatenated with the feature F1 obtained from the backbone network using the connection module to achieve multi-scale feature fusion. Finally, the attention module is used for processing to generate the enhanced feature F2. The enhanced feature F2 is downsampled using convolution, and then the downsampled result is concatenated with the enhanced feature F1 using a connection module to construct a cross-scale feature connection. Finally, it is processed by an attention module to generate the enhanced feature F3. The enhanced feature F3 is downsampled using convolution. Then, the downsampled result is concatenated with the feature F3 obtained from the backbone network using a connection module to form the final multi-scale fusion. Finally, the bottleneck module is used to extract deep features and generate the enhanced feature F4.

[0012] Furthermore, the process of fusing query features and supporting features through interactive matching to generate aggregated features representing fault co-occurrence patterns includes: For the query features, obtain enhanced features F2, F3, and F4; for the supporting features, obtain enhanced features F2', F3', and F4'. First, the enhanced features F2, F3, and F4 are input into two consecutive convolutional modules for processing to obtain features A1, A2, and A3. Then, the strong feature F2', enhanced feature F3', and enhanced feature F4' are input into two consecutive convolutional modules for processing, and the output features are processed through an adaptive max pooling layer. The output size of the adaptive max pooling layer is set to 1×1 to obtain features A1', A2', and A3'. Finally, feature A1 and feature A1' are multiplied by channel-wise scalar to obtain aggregated features, feature A2 and feature A2' are multiplied by channel-wise scalar to obtain aggregated features, and feature A3 and feature A3' are multiplied by channel-wise scalar to obtain aggregated features.

[0013] Furthermore, the step of performing target detection based on aggregated features and outputting the location and classification results of power faults includes: The object detection module uses the previously obtained enhanced features and aggregated features as input features, with feature sizes of respectively. , , , , , ; Each enhanced feature is processed using convolutional modules and convolutional layers to obtain class and location predictions. These predictions are then further processed by concatenating them along the channels, performing dimensionality reduction adjustments, and finally concatenating them along the final dimension. The category prediction and location prediction are separated on the channel and their dimensions are swapped to obtain category prediction D4 and location prediction D4. The size of category prediction D4 is [missing value]. The predicted location D4 size is ; Finally, the location prediction D4 is decoded using the generated anchor points to obtain the actual location prediction. This step first converts the network's output distribution prediction into specific bounding box offsets, and then converts the offsets into absolute coordinates. The sigmoid activation function is then applied to the class prediction to obtain the final class prediction result.

[0014] Secondly, the present invention provides a power fault identification system based on small sample target detection, comprising: The data acquisition module is used to preprocess the acquired power equipment fault data and locally stored image tag data to obtain query data and support data. The feature extraction module is used to extract and enhance features from the query data and support data respectively, generating query features and support features respectively. The aggregation module is used to fuse query features and supporting features through interactive matching to generate aggregated features that represent the co-occurrence pattern of faults. The output module is used to perform target detection based on aggregated features and output the location and classification results of power faults.

[0015] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the power fault identification method based on small sample target detection.

[0016] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the power fault identification method based on small sample target detection.

[0017] Compared with the prior art, the present invention has the following technical effects: The proposed power fault identification method creatively integrates a few-shot learning paradigm with a single-stage detection architecture. At the detection framework level, YOLOv12n is used as the basic detector, and feature extraction and bounding box regression are decoupled through an anchorless design to improve the detection speed of the model. At the few-shot adaptation level, a meta-learning module based on feature recalibration is developed, which strengthens the representation of fault features through a channel attention mechanism to improve the detection accuracy of targets in few samples.

[0018] Introducing few-sample techniques into single-stage target detection methods can improve the detection performance of few-sample fault categories while maintaining model speed. Attached Figure Description

[0019] Figure 1 The reasoning process diagram of the power fault identification method based on small sample target detection.

[0020] Figure 2 Network structure of the feature extraction module.

[0021] Figure 3 Bottleneck module structure.

[0022] Figure 4 Attention module structure.

[0023] Figure 5Feature fusion module.

[0024] Figure 6 Feature aggregation module.

[0025] Figure 7 Target detection module.

[0026] Figure 8 This is a flowchart of the present invention. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings: Example 1, please refer to Figure 8 This invention provides a power fault identification method based on small sample target detection, comprising: The collected power equipment fault data and locally stored image tag data are preprocessed to obtain query data and support data. The query data and supporting data are subjected to feature extraction and feature enhancement respectively, generating query features and supporting features respectively; The query features and supporting features are fused together through interactive matching to generate aggregate features that represent the co-occurrence pattern of faults. Target detection is performed based on aggregated features, and the results of power fault location and classification are output.

[0028] Example 2: This invention provides a power fault identification method based on small sample target detection, specifically including: The module includes a data processing module, a feature extraction module, a feature fusion module, a feature aggregation module, and a target detection module. Figure 1 This paper demonstrates the inference process of the proposed power fault identification method based on few-sample target detection. In actual deployment, our model remains running in memory as a service. After acquiring data from the data acquisition device, the data is stored as Data 1, and the local loader is notified to load Data 2 from the local machine. Data 1 is used as retrieval data and sequentially input into Data Processing Module 1, Feature Extraction Module 1, and Feature Fusion Module 1 to obtain query features. Feature Extraction Module 1 and Feature Extraction Module 2 share the same structure and parameters except for the first layer. Feature Fusion Module 1 and Feature Fusion Module 2 also share the same structure and parameters. Data 2 is used as support data and sequentially input into Data Processing Module 2, Feature Extraction Module 2, and Feature Fusion Module 2 to obtain support features. The query features and support features are input into the Feature Aggregation Module to obtain aggregated features, which are then input into the target detection module to obtain the power fault identification result. The training and inference processes maintain a similar structure. After obtaining the detection results, a loss function calculation module is added to the training process, and the parameters of each module are updated through backpropagation based on the calculated loss. Each module is described in detail below: 1. Data processing module: The data processing module processes the data to be input into the model. The data processing module is different in the training and inference stages. The training stage is further divided into the basic training stage and the fine-tuning stage, and their data processing methods are also different. We will introduce them separately.

[0029] (1) Data processing in the training stage Before model training, first collect the power failure dataset according to business needs. Mark the positions and categories of the faults in the dataset through the method of manual annotation + expert correction to obtain the training data set Data, including the category set C. Count the number of occurrences of each type of fault in the training data set Data, and select the basic classes and novel classes through the threshold th (th = 100). The number of occurrences of each type of fault in the basic class in the training data set Data is greater than th, and the number of occurrences of each type of fault in the novel class in the training data set Data is less than or equal to th.

[0030] 1) Data processing in the basic training stage First, select the samples corresponding to the basic classes from the training data set Data to form the basic training set , The number of images in is . Then, construct the support set data set based on the basic training set. The process of constructing the support set data set is as follows: Select th images with the label of the category c from the training set without replacement for the category set in the basic class set, and select a label belonging to the category c for each image to form the support set data of this category. Each image can appear in at most one category set.

[0031] Construct an index for the support set data of this category:

[0032] represents the index of the i-th data of the c-th category in the support set, is the index of the selected image in the basic training set, is the index of the selected label among all the labels of the corresponding image. If the number of available labeled data n < th for a certain category, then the selected data for this category is repeated (th / / n + 1) times.

[0033] Repeat the above steps until the support set data set is formed for each category.

[0034] Generate matching base training set length Supports indexing of data loading order in sets. Each time, it loads from the base class collection. A random selection num1 without replacement (num1= Given (number of categories), select num2 (num2=1) data points without replacement from the support set of each category to form a support set data loading order index set. Repeat this step until creation. This is a set of indexes that support the data loading order.

[0035] During the basic training phase, we acquire a batch of data each time, which can be divided into two parts: query data and support data.

[0036] Use data processing module 1 to process the queried data as follows: Implement a multi-scale training strategy, randomly selecting a number divisible by 32 from [640, 1280] as the size of the longest side of the image.

[0037] Image files are loaded from the basic training dataset in the default order, raw digital images are read from the specified path, the initial decoding of the pixel matrix is ​​completed, and a multidimensional array representation is generated.

[0038] Load the labeled data, simultaneously load the labeled files that match the image, and parse out the set of bounding box locations of the fault targets and their corresponding set of category labels.

[0039] Perform multi-scale dynamic adjustments on the image, resizing it and adjusting the longest side to W1 while maintaining the aspect ratio. Scale the bounding box position consistently.

[0040] The image is randomly augmented by performing a horizontal flip with a 50% probability, simultaneously flipping the bounding box position. Then, random color transformations and random noise additions are performed to improve the model's spatial invariance.

[0041] Channel normalization is performed, the color space is converted from RGB to BGR format, and the mean of each channel is normalized by subtracting [B:103.53, G:116.28, R:123.675], and the standard deviation is scaled to keep the original standard deviation of each channel (std=[1.0,1.0,1.0]).

[0042] Perform boundary padding to meet network input requirements, padding the shorter side of the image with zeros to multiples of 32. Perform a consistent translation operation on the bounding box position.

[0043] The data format is standardized, the bounding box positions are normalized according to the length and width of the image, the input image, the set of bounding box positions, and the set of category labels are converted into tensors of the deep learning framework PyTorch, the image is rearranged in dimensions, and finally GPU memory is allocated according to the configuration.

[0044] Repeat steps (b)-(h) above until a batch of data has been loaded and processed. Concatenate the images along the channels to form a four-dimensional tensor, Imgs. Create a num3x6 two-dimensional tensor, where num3 is the number of all labels corresponding to the batch of data. Each label can be represented as [Imgidx,cls,x,y,w,h]. Imgidx is the index of the image corresponding to that label in Imgs, cls is the fault category corresponding to that label, and [x,y,w,h] are the coordinates of the center point and the width and height of the fault location corresponding to that label.

[0045] The supporting data is processed using data processing module 2 as follows: Based on the index of the current batch, read the pre-built set of support set data loading order indexes from the support set data loading index.

[0046] Read one piece of data from the support set data loading order index set. The image is retrieved based on the image index, and the original image file is read from the specified path to complete the initial loading operation of pixel data.

[0047] Synchronously load the annotation data corresponding to the image, and retain only the annotation data. Specifies the bounding box coordinates and category label for the index.

[0048] The image is processed by channel normalization, the color space is converted from RGB to BGR format, and the mean is normalized by subtracting [B:103.53, G:116.28, R:123.675] from each channel. The standard deviation is scaled to keep the original standard deviation of each channel (std=[1.0,1.0,1.0]).

[0049] The image size is adjusted and a mask is generated. The original image is scaled to 224×224 pixels, and mask data matching the image is generated. The image and the mask are concatenated on the channel to obtain the processed image.

[0050] Data format standardization involves converting the processed image data into tensor format for the PyTorch deep learning framework, including tensor dimension conversion and data type adaptation. Labels are also converted to tensor format and normalized according to the original image size.

[0051] Repeat the process of (b)-(f) until all the corresponding data in the support set data loading order index set are processed. Concatenate a batch of images on the channels to obtain a four-dimensional tensor Imgs'. Create a two-dimensional tensor of num4x6, where num4 is the number of support set pictures loaded, and each annotation can be represented as [Imgidx’,cls’,x’,y’,w’,h’]. Imgidx’ is the index of the picture corresponding to this annotation in Imgs’, cls’ is the fault category corresponding to this annotation, and [x’,y’,w’,h’] are the center coordinates and length and width of the fault position corresponding to this annotation.

[0052] After the query data and support data are processed, the processed query data Query and support data Support are obtained, which are used for the subsequent process.

[0053] 2) Data preparation for the fine-tuning stage Select the samples corresponding to the novel classes from the training data set Data to form a novel data set , and randomly select about th samples for each base class from the training data set Data and add them to the fine-tuning set as the final fine-tuning set . [[ID=,13]] The number of images in , and the class set is the full set of classes . Then, based on the fine-tuning training set, construct a support set data set. The construction of the support set data set: From the fine-tuning set For the classes c in the full set of classes Without replacement, select th / / 2 images with the label of this class, and select one label belonging to class c for each image to form the support set data of this class. Each image can appear in at most one class set.

[0054] Construct an index for the support set data of this class:

[0055] represents the index of the i-th data of the c-th class in the support set, is the index of the selected image in the fine-tuning training set, is the index of the selected label among all the labels of the corresponding image. If the number of available labeled data n < th / / 2 for a certain class, then the selected data for this class is repeated (th / / 2n + 1) times.

[0056] Repeat the above steps until a support set data set is formed for each class.

[0057] Generate a length matching the fine-tuning training set Supports indexing of the data loading order. Each time, it loads from the full class collection. A random selection num5 without replacement (num5= Given (number of categories), select num6 (num6=1) data points without replacement from the support set of each category to form a support set data loading order index set. Repeat this step until creation. This is a set of indexes that support the data loading order.

[0058] Similar to the basic training phase, the fine-tuning phase acquires a batch of data each time. The acquired data can be divided into two parts: query data and support data. Their specific processing procedures are the same as those in the basic phase.

[0059] After the query data and support data have been processed, the processed query data and support data are obtained. These are used in subsequent processes.

[0060] (2) Data preparation in the reasoning stage During the inference phase, after the data acquisition module collects the data, it saves it to data 1. The data processing module 1 then processes the data as follows: Read the images one by one from Data 1 and generate a multidimensional array representation.

[0061] Adjust the image size by setting the longest side to 1280 while maintaining the aspect ratio, and record the image scaling ratio.

[0062] The image is processed by channel normalization, the color space is converted from RGB to BGR format, and the mean is normalized by subtracting [B:103.53, G:116.28, R:123.675] from each channel. The standard deviation is scaled to keep the original standard deviation of each channel (std=[1.0,1.0,1.0]).

[0063] Perform boundary padding to meet network input requirements, padding the shorter sides of the image with zeros to multiples of 32. Record the padding details [pt1, pb2, pl1, pr2], which represent the column number of each of the four sides of the image (top, bottom, left, right) to be padded.

[0064] The data format is standardized by converting the input image into a tensor using the PyTorch deep learning framework. The image dimensions are rearranged, and CPU / GPU memory is allocated based on the device's GPU status. The image scaling ratio and edge padding are constructed into a 1x4 tensor as additional image information.

[0065] Repeat steps (a)-(e) above until all images in Data 1 have been processed. Concatenate the images along the channels to form a four-dimensional tensor, ImgsInfer. Create a 7x6 two-dimensional tensor, where 7 is the number of images provided by Data 1. Each piece of additional information can be represented as [Imgidxinfer, ratio, pt1, pb2, pl1, pr2]. Imgidxinfer is the index of the image corresponding to that piece of additional information in ImgsInfer, ratio is the scaling ratio of the image, and [pt1, pb2, pl1, pr2] represents the edge configuration of the image.

[0066] Support data for the inference phase is stored in a fixed local location. Each category has a separate folder within this location, containing the image and label files for that category. Each label file contains only one fault marker belonging to that category, and categories in different folders are unique. After the data acquisition module collects data, it instructs the local loading module to randomly load one support set data for that category from each folder at the fixed address, resulting in Data 2. Data 2 contains both images and labels. Data 2 is then processed as follows: Select one data point from Data 2, perform channel normalization on the image, convert the color space from RGB to BGR format, subtract [B:103.53, G:116.28, R:123.675] from each channel to normalize the mean, and scale the standard deviation to keep the original standard deviation of each channel (std=[1.0,1.0,1.0]).

[0067] The image size is adjusted and a mask is generated. The original image is scaled to 224×224 pixels, and mask data matching the image is generated. The image and the mask are concatenated on the channel to obtain the processed image.

[0068] Data format standardization involves converting the processed image data into tensor format for the PyTorch deep learning framework, including tensor dimension conversion and data type adaptation. Labels are also converted to tensor format and normalized according to the original image size.

[0069] Repeat steps (a)-(c) until all data in Data 2 has been processed. Concatenate a batch of images along the channels to form a four-dimensional tensor `ImgsInfer'`. Create a num8x6 two-dimensional tensor, where num8 is the number of images in Data 2. Each label can be represented as `[ImgidxInfer',clsInfer',xInfer',yInfer',wInfer',hInfer']`. `ImgidxInfer'` is the index of the image corresponding to this label within `ImgsInfer'`, `clsInfer'` is the fault category corresponding to this label, and `[xInfer',yInfer',wInfer',hInfer']` are the center coordinates and dimensions of the fault location corresponding to this label.

[0070] After the query data and support data have been processed, the processed query data and support data are obtained. These are used in subsequent processes.

[0071] 2. Feature Extraction Module: The feature extraction module remains consistent between the training and inference phases; the only difference is that gradient updates are performed on the parameters of the feature extraction module during the training phase. Below, we will describe the feature extraction module in detail: The network structure of the feature extraction module we used is as follows: Figure 2 As shown, it contains 5 convolutional modules, 2 bottleneck modules, and 2 attention modules. After data is input into the feature extraction module, it sequentially enters convolutional module B1, convolutional module B2, bottleneck module B1, convolutional module B3, bottleneck module B2, convolutional module B4, attention module B1, convolutional module B5, and attention module B2 to obtain the extracted features. Convolutional module B1 has 3 or 4 input channels and c output channels (c=16). Convolutional module B2 has c input channels and 2c output channels. Bottleneck module B1 has 2c input channels and 4c output channels. Convolutional module B3 has 4c input channels and 4c output channels. Bottleneck module B2 has 4c input channels and 8c output channels. Convolutional module B4 has 8c input channels and 8c output channels. Attention module B1 has 8c input channels and 8c output channels. The convolution module B5 has 8c input channels and 16c output channels. The attention module B2 has 16c input channels and 16c output channels.

[0072] The structure of the bottleneck module is as follows Figure 3As shown, it contains 4 convolutional modules, 1 segmentation module, and 1 addition module. Features input from the module preceding the bottleneck module i (i=1, 2) enter the bottleneck module, first passing through convolutional module i1, and then the output is input to segmentation module i1. Segmentation module i1 segments the input features into two identical feature maps, i1 and i2. Feature i2 is then sequentially input into convolutional modules i2 and i3, and the output is added to feature i2 using addition module i1 to obtain feature i3. Features i1, i2, and i3 are then concatenated along the channels using connection module i1. Finally, convolutional module i4 processes the output of the connection module to obtain the bottleneck module's output.

[0073] The structure of the attention module is as follows Figure 4 As shown, it contains 6 convolutional modules, 2 attention sub-modules, 4 addition modules, and 1 connection module. Features input from the module preceding the attention module (p=1, 2) enter the attention module and first pass through convolutional module p1 to obtain feature p1. Feature p1 is then input into attention sub-module p1, and the output is added to feature p1 using addition module p1 to obtain feature p2. Feature p2 is then sequentially input into convolutional modules p2 and p3, and the result is added to feature p2 using addition module p2 to obtain feature p3. Feature p3 is then input into attention sub-module p2, and the result is added to addition module p3 to obtain feature p4. Finally, feature p4 is sequentially input into convolutional modules p4 and p5, and the result is added to addition module p4 to obtain feature p5. Finally, features p1, p3, and p5 are input into the connection module p1 to be connected on the channels, and the result is input into the convolution module p6 to obtain the output of the attention module.

[0074] The attention submodule uses the region attention mechanism proposed in YOLOv12, which includes the following 5 steps: (1) First, input processing is performed. Assume the shape of input x is B×C×H×W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions, and the number of regions is a. Q, K, V, and Pe are obtained by performing a convolution operation on x: Q,K=Conv1(X), V=Conv2(X), Pe=Conv3(V) (2) Transform the Q, K, and V feature dimensions. Convert Q, K, and V into a multi-head form, with the number of heads being 'head', and divide the spatial dimension into 'a' regions: Q=Reshape(Q,[B×a,head,C / head,N / a]) K=Reshape(K,[B×a,head,C / head,N / a]) V=Reshape(V,[B×a,head,C / head,N / a]) (3) Scale the dot product attention and normalize it using Softmax:

[0075] (4) Apply attention weights to Value and perform dimensional transformation on the output:

[0076] Z = Reshape(Z,[B,C,H,W]) (5) Weight the sum of the value Z and the position code Pe, and reconstruct the output using a convolution operation: Y = Conv3(Z + Pe) The obtained Y is used as the output of the attention submodule.

[0077] Query data A is input into feature extraction module 1 for feature extraction, resulting in query feature FA. Supporting data B is input into feature extraction module 2, resulting in supporting feature FB. Since query data A and supporting data B have different dimensions, the first layer of feature extraction modules 1 and 2 differs, while subsequent layers remain the same. The first layer of feature extraction modules 1 and 2 does not share any parameters, while subsequent layers share parameters. The first layer of extraction module 1 is set to support 3 channels, and the first layer of feature extraction module 2 is set to support 4 channels.

[0078] 3. Feature Fusion Module: The feature fusion module remains consistent between the training and inference phases, with the only difference being that its parameters are updated using gradients during training. Below, we describe the feature fusion module in detail: The feature fusion module path aggregation network adopts a dual-path feature fusion approach, using a top-down path to enhance semantic information and a bottom-up path to enhance detail information.

[0079] The structure of the feature fusion module is as follows: Figure 5 As shown, features F1, F2, and F3 obtained from the backbone network are used as input to the feature fusion module, which contains two upsampling modules, four connection modules, three attention modules, and one bottleneck module. This can be divided into four steps: The first step is to perform a 2x nearest neighbor upsampling on feature F3 using the upsampling module F1. Then, the upsampling result is concatenated with feature F2 obtained from the backbone network using the connection module F1 along the channel dimension to achieve the fusion of shallow detailed features and deep semantic features. Finally, the attention module F1 is used to refine the features and generate the enhanced feature F1.

[0080] The second step is to first use the upsampling module F2 to perform a 2x nearest neighbor upsampling on the enhanced feature F1, and then use the connection module F2 to concatenate the upsampling result with the feature F1 obtained from the backbone network to achieve multi-scale feature fusion. Finally, the attention module F2 is used for processing to generate the enhanced feature F2.

[0081] The third step first applies a 3×3 convolution with a stride of 2 to downsample the enhanced feature F2. Then, the downsampled result is concatenated with the enhanced feature F1 using the connection module F3 to construct a cross-scale feature connection. Finally, it is processed by the attention module F3 to generate the enhanced feature F3.

[0082] The fourth step first involves downsampling the enhanced feature F3 using a 3×3 convolution with a stride of 2. Then, the downsampling result is concatenated with the feature F3 obtained from the backbone network using the connection module F4 to form the final multi-scale fusion. Finally, the bottleneck module F1 is used for deep feature extraction to generate the enhanced feature F4.

[0083] Feature fusion module 1 and feature fusion module 2 share the same structure and parameters, so they will not be described separately here.

[0084] 4. Feature aggregation module: The feature aggregation module remains consistent between the training and inference phases; the only difference is that gradient updates are performed on the parameters of the feature aggregation module during the training phase. The structure of the feature aggregation module is as follows: Figure 6 As shown, the feature aggregation module obtains enhanced features F2, F3, and F4 from the feature fusion module 1, and the feature aggregation module obtains enhanced features F2', F3', and F4' from the feature fusion module 2.

[0085] First, the enhanced features F2, F3, and F4 are input into two consecutive convolutional modules for processing to obtain features A1, A2, and A3.

[0086] Then, the strong feature F2', enhanced feature F3', and enhanced feature F4' are input into two consecutive convolutional modules for processing. The output features are then processed through an adaptive max pooling layer with an output size of 1×1, resulting in features A1', A2', and A3'.

[0087] Finally, feature A1 and feature A1' are multiplied by channel-wise scalar to obtain aggregated feature A1, feature A2 and feature A2' are multiplied by channel-wise scalar to obtain aggregated feature A2, and feature A3 and feature A3' are multiplied by channel-wise scalar to obtain aggregated feature A3.

[0088] 5. Target Detection Module: The object detection module takes the previously obtained enhanced features F2, F3, F4, aggregated features A1, A2, and A3 as input features, with feature sizes of respectively. , , , , , The object detection module differs between the training and inference phases; the only difference between the basic training and fine-tuning phases lies in the feature size of the category prediction output. The training and inference phases are described below.

[0089] (1) Training phase The processing structure of the object detection module during the training phase is as follows: Figure 7 As shown, each enhancement feature is first processed using convolutional modules and convolutional layers. Aggregated feature A1 is processed by convolutional layer D1 to obtain category prediction D1, and enhancement feature F2 is processed by convolutional modules D1, D2, and D2 to obtain location prediction D1. Aggregated feature A2 is processed by convolutional layer D3 to obtain category prediction D2, and enhancement feature F3 is processed by convolutional modules D3, D4, and D4 to obtain location prediction D2. Aggregated feature A3 is processed by convolutional layer D5 to obtain category prediction D3, and enhancement feature F4 is processed by convolutional modules D5, D6, and D6 to obtain location prediction D3.

[0090] Then, the obtained category prediction results and location prediction results are processed. Specifically, the category prediction D1 and location prediction D1 are concatenated on the channels, and their shape is changed from 4D. Adjust to three-dimensional Concatenate the category prediction D2 and the location prediction D2 along the channels to change their shape from 4D. Adjust to three-dimensional Concatenate the category prediction D3 and the location prediction D3 along the channels to change their shape from 4D. Adjust to three-dimensional The results are then concatenated along the final dimension to obtain... The category prediction and location prediction are separated by channel and their dimensions are swapped to obtain category prediction D4 and location prediction D4. The size of category prediction D4 is [missing value]. The predicted location D4 size is .

[0091] Finally, the location prediction D4 is decoded using the generated anchor points to obtain the actual location prediction. This step first converts the network output distribution prediction into specific bounding box offsets, and then converts the offsets into absolute coordinate format. The sigmoid activation function is applied to the class prediction to obtain the final class prediction result.

[0092] (2) Reasoning stage The process preceding the inference phase is the same as the training phase. After obtaining the position prediction results in absolute coordinate format and the final category prediction results, the following processing is performed: 1) Take the maximum value of the class probability for each anchor box to obtain the confidence matrix of shape (batch_size, num_anchors). Use topk to select the min(max_det, num_anchors) anchor boxes with the highest confidence in each image and record their indices.

[0093] 2) Extract the following from the index using the gather operation: the corresponding bounding box coordinates and the corresponding complete category probability distribution.

[0094] 3) Secondary screening of the highest-scoring detection results: flatten the filtered category prediction results into a two-dimensional tensor, and execute topk again to select the top max_det highest-scoring detection results globally to obtain the final retained score value.

[0095] 4) Integrate the final output. The final output shape is (batch_size, max_det, 6), which includes: [x, y, w, h, max_class_prob, class_index]. [x, y, w, h] are the center position and width and height of the detection box, max_class_prob is the probability of the predicted class, and class_index is the index of the predicted class.

[0096] 6. Loss Function Calculation Module: The loss function calculation module is only used during model training. The loss function used in basic training and fine-tuning is basically the same, differing only in the number of categories processed. The loss function calculation module is described in detail below.

[0097] After obtaining the location and category prediction results during the training phase, the loss must be calculated between the prediction and the label.

[0098] First, the labels are processed. This preprocessing method converts the label data in the object detection task into a normalized tensor format that matches the input batch size. The main process is as follows: (1) Input parsing stage: receive raw label data and extract data dimension information: nl represents the number of samples and ne represents the parameter dimension of each label.

[0099] (2) Special handling for empty data: When there is no target data (nl==0), return an all-zero tensor with dimensions (B,0, ne-1) to maintain compatibility with subsequent processing.

[0100] (3) In the target allocation stage, extract the image index column (column 0) corresponding to all targets from the original labels, count the number of labels contained in each image, and obtain the unique image index and its corresponding target count. Create an output tensor out with dimensions (B_size, maximum target count, ne-1), and fill unused positions with zeros.

[0101] (4) Data filling loop: Iterate through each image index in the batch and generate a Boolean mask `matches` to mark the targets belonging to the current image. When a valid target exists, fill the corresponding target parameters (skipping column 0) into the corresponding position of the output tensor.

[0102] (5) During the coordinate system transformation stage, the coordinate columns (columns 1-4, assumed to be in xywh format) of the output tensor are transformed. First, a scaling factor is applied for normalization and restoration, and then the center coordinates + width and height format is converted to the bounding box coordinate format.

[0103] (6) Return the normalized result, and finally output a tensor that is strictly aligned with the input batch.

[0104] Then the prediction results are matched with the labels.

[0105] Finally, the loss is calculated based on the loss function. The loss consists of three parts: distribution focus loss, location prediction loss, and category prediction loss.

[0106] α, β, and λ are three coefficients used to balance the proportions of the distribution focus loss, location prediction loss, and category prediction loss in the total loss. In implementation, we set α=1.5, β=7.5, and λ=0.5.

[0107] For location prediction loss, we use the following loss function to calculate:

[0108] K is the set of foreground elements in the image, and C is the number of categories. It is the score of the k-th target box in the c-th category. It is the coordinate vector of the k-th prediction box, in the format x1, y1, x2, y2. It is the coordinate vector of the k-th label box, in the format x1, y1, x2, y2. Calculate the CIoU loss between the predicted bounding box and the label bounding box.

[0109] For category prediction loss, we use the cross-entropy loss function.

[0110] in It is the sigmoid activation function, where C is the number of classes. It's a category label. It is the predicted probability that the target belongs to the c-th class.

[0111] In another embodiment of the present invention, a power fault identification system based on small sample target detection is provided, which can be used to implement the above-mentioned power fault identification method based on small sample target detection. Specifically, the system includes: The data acquisition module is used to preprocess the acquired power equipment fault data and locally stored image tag data to obtain query data and support data. The feature extraction module is used to extract and enhance features from the query data and support data respectively, generating query features and support features respectively. The aggregation module is used to fuse query features and supporting features through interactive matching to generate aggregated features that represent the co-occurrence pattern of faults. The output module is used to perform target detection based on aggregated features and output the location and classification results of power faults.

[0112] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0113] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used to operate a power fault identification method based on small sample target detection.

[0114] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the power fault identification method based on small sample target detection in the above embodiments.

[0115] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0116] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A power fault identification method based on small sample target detection, characterized in that, include: The collected power equipment fault data and locally stored image tag data are preprocessed to obtain query data and support data. The query data and supporting data are subjected to feature extraction and feature enhancement respectively, generating query features and supporting features respectively; The query features and supporting features are fused together through interactive matching to generate aggregate features that represent the co-occurrence pattern of faults. Target detection is performed based on aggregated features, and the results of power fault location and classification are output.

2. The power fault identification method based on small sample target detection according to claim 1, characterized in that, The collected power equipment fault data and locally pre-stored image tag data include: A power fault dataset is collected. The faults in the dataset are labeled with their location and category through annotation and correction, resulting in a training dataset Data, containing a category set C. The frequency of each category of faults in the training dataset Data is counted, and a base class is selected based on a threshold th. and novel In the basic class, the number of times each type of fault appears in the training data set Data is greater than th; in the novel class, the number of times each type of fault appears in the training data set Data is less than or equal to th. Select samples corresponding to the basic classes from the training dataset Data to form the basic training set. , The number of images in the middle is Then, a support set dataset is constructed based on the basic training set; Select samples corresponding to novel classes from the training dataset Data to form a novel dataset. Approximately th samples are randomly selected from the training dataset Data for each base class and added to the fine-tuning set, which serves as the final fine-tuning set. ; The number of images in the middle is The class set is the complete set of classes. A support set dataset is constructed based on the fine-tuned training set; Locally pre-stored data is stored in a fixed location. Each category has a separate folder in this location, containing image files and label files for that category. Each label file contains only one fault marker belonging to that category, and categories in different folders are not duplicated. When the data acquisition module collects data, it notifies the local loading module to randomly load the support set data for that category from each folder under the fixed address.

3. The power fault identification method based on small sample target detection according to claim 2, characterized in that, The preprocessing process is as follows: The training dataset Data is processed as follows: Read each image one by one and generate a multidimensional array representation; adjust the image size while maintaining the image aspect ratio, and record the image scaling ratio. The image is processed by channel normalization, the color space is converted from RGB to BGR format, the mean of each channel is normalized, and the standard deviation is scaled to keep the original standard deviation of each channel. Perform boundary padding on the image, specifying the number of columns to pad each of the four sides of the image (top, bottom, left, and right). The data format is standardized, the input image is converted into a tensor of the deep learning framework, the image dimensions are rearranged, and finally, CPU / GPU memory is allocated according to the device GPU status. The image scaling ratio and edge padding are constructed into tensors as additional information for the image. The operation is repeated until all images are processed. For the locally stored image label data, perform the following operations: Select one data point from the locally stored image label data, perform channel normalization on the image, convert the color space from RGB to BGR format, normalize the mean of each channel, and scale the standard deviation to keep the original standard deviation of each channel. Image resizing and mask generation: The original image is scaled and mask data matching the image is generated. The image and the mask are concatenated on the channels to obtain the processed image. Data format standardization: Convert the processed image data into tensor format for deep learning frameworks, including tensor dimension conversion and data type adaptation, convert labels into tensor format, and normalize them according to the original image size; repeat until all data has been processed.

4. The power fault identification method based on small sample target detection according to claim 1, characterized in that, The step of extracting features from the query data and supporting data separately includes: Feature extraction is performed using a feature extraction module that includes 5 convolutional modules, 2 bottleneck modules, and 2 attention modules. After the query data and support data are input into the feature extraction module, the query feature FA and support feature FB are obtained. The attention module comprises 6 convolutional modules, 2 attention sub-modules, 4 addition modules, and 1 connection module. The attention sub-modules use the region attention mechanism proposed in YOLOv12, including: (1) Let the shape of the input x be B×C×H×W, where B is the batch size, C is the number of channels, H and W are the spatial dimensions, and a is the number of regions. By performing a convolution operation on x, we obtain Q, K, V, and Pe: Q,K=Conv1(X), V=Conv2(X), Pe=Conv3(V) (2) Convert Q, K, and V into a multi-head form, with the number of heads being head, and divide the spatial dimension into a regions: Q=Reshape(Q,[B×a,head,C / head,N / a]) K=Reshape(K,[B×a,head,C / head,N / a]) V=Reshape(V,[B×a,head,C / head,N / a]) (3) Scale the dot product attention and normalize it using Softmax: (4) Apply attention weights to Value and perform dimensional transformation on the output: Z = Reshape(Z,[B,C,H,W]) (5) Weight the sum of the value Z and the position code Pe, and reconstruct the output using a convolution operation: Y = Conv3(Z + Pe) The obtained Y is used as the output of the attention submodule.

5. The power fault identification method based on small sample target detection according to claim 4, characterized in that, Feature enhancement is performed on query feature FA and supporting feature FB: The feature fusion module, a path aggregation network, employs a dual-path feature fusion approach: top-down to enhance semantic information and bottom-up to enhance detail information. Feature enhancement is achieved using a feature fusion module, which comprises two upsampling modules, four connection modules, three attention modules, and one bottleneck module. First, the feature F3 obtained from the backbone network is upsampled by 2 times using the upsampling module. Then, the upsampling result is concatenated with the feature F2 obtained from the backbone network using the connection module to perform channel-dimensional concatenation. Finally, the feature is refined through the attention module to generate the enhanced feature F1. The upsampling module is used to perform a 2x nearest neighbor upsampling on the enhanced feature F1. Then, the upsampling result is concatenated with the feature F1 obtained from the backbone network using the connection module to achieve multi-scale feature fusion. Finally, the attention module is used for processing to generate the enhanced feature F2. The enhanced feature F2 is downsampled using convolution, and then the downsampled result is concatenated with the enhanced feature F1 using a connection module to construct a cross-scale feature connection. Finally, it is processed by an attention module to generate the enhanced feature F3. The enhanced feature F3 is downsampled using convolution. Then, the downsampled result is concatenated with the feature F3 obtained from the backbone network using a connection module to form the final multi-scale fusion. Finally, the bottleneck module is used to extract deep features and generate the enhanced feature F4.

6. The power fault identification method based on small sample target detection according to claim 5, characterized in that, The process of fusing query features and supporting features through interactive matching to generate aggregated features representing fault co-occurrence patterns includes: For the query features, obtain enhanced features F2, F3, and F4; for the supporting features, obtain enhanced features F2', F3', and F4'. First, the enhanced features F2, F3, and F4 are input into two consecutive convolutional modules for processing to obtain features A1, A2, and A3. Then, the strong feature F2', enhanced feature F3', and enhanced feature F4' are input into two consecutive convolutional modules for processing, and the output features are processed through an adaptive max pooling layer. The output size of the adaptive max pooling layer is set to 1×1 to obtain features A1', A2', and A3'. Finally, feature A1 and feature A1' are multiplied by channel-wise scalar to obtain aggregated features, feature A2 and feature A2' are multiplied by channel-wise scalar to obtain aggregated features, and feature A3 and feature A3' are multiplied by channel-wise scalar to obtain aggregated features.

7. The power fault identification method based on small sample target detection according to claim 1, characterized in that, The step of performing target detection based on aggregated features and outputting the location and classification results of power faults includes: The object detection module uses the previously obtained enhanced features and aggregated features as input features, with feature sizes of respectively. , , , , , ; Each enhanced feature is processed using convolutional modules and convolutional layers to obtain class and location predictions. These predictions are then further processed by concatenating them along the channels, performing dimensionality reduction adjustments, and finally concatenating them along the final dimension. The category prediction and location prediction are separated on the channel and their dimensions are swapped to obtain category prediction D4 and location prediction D4. The size of category prediction D4 is [missing value]. The predicted location D4 size is ; Finally, the generated anchor points are used to decode the location prediction D4 to obtain the actual location prediction. This step first converts the distribution prediction output by the network into specific bounding box offsets, and then converts the offsets into absolute coordinate format; the sigmoid activation function is applied to the category prediction to obtain the final category prediction result.

8. A power fault identification system based on small sample target detection, characterized in that, include: The data acquisition module is used to preprocess the acquired power equipment fault data and locally stored image tag data to obtain query data and support data. The feature extraction module is used to extract and enhance features from the query data and support data respectively, generating query features and support features respectively. The aggregation module is used to fuse query features and supporting features through interactive matching to generate aggregated features that represent the co-occurrence pattern of faults. The output module is used to perform target detection based on aggregated features and output the location and classification results of power faults.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the power fault identification method based on small sample target detection as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the power fault identification method based on small sample target detection as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cross-equipment rolling bearing small sample fault diagnosis method based on feature fusion

    CN116186641A

  • Electric power small sample defect detection method and device, computer equipment and storage medium

    CN116205916A

  • Small sample regulating valve fault diagnosis method using attention enhancement mechanism

    CN116975601A

  • Bearing fault diagnosis method and device based on small sample learning and medium

    CN118606815A

  • Equipment fault diagnosis method based on deep Brown distance and attention mechanism

    CN120408295A