Detection Method and Storage Medium for AC Filter Defects

Through the trained defect detection model, feature extraction, enhancement and fusion of AC filter images is solved, and the problems of low real-time detection of AC filter defects in the prior art are solved, and efficient defect detection and identification are achieved.

CN119722670BActive Publication Date: 2025-05-30STATE GRID ANHUI ULTRA HIGH VOLTAGE CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220975.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing defect detection algorithm based on deep learning has problems such as low real-time, slow speed and insufficient detection capabilities for small targets in AC filter defect detection, which cannot meet the detection needs of large targets and small target defects in AC filters.

Method used

An AC filter defect detection method is proposed. Through the trained defect detection model, feature extraction, enhancement and fusion of the AC filter image to be detected, and a combination of feature extraction module, feature fusion module and defect detection module is used to achieve rapid detection of defects of large and small targets.

Benefits of technology

The efficiency of AC filter defect detection is improved, and the detection needs of AC filters with large and small target defects can be met, and the rapid identification and positioning of AC filter defects can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722670B_ABST
    Figure CN119722670B_ABST
Patent Text Reader

Abstract

The present invention discloses a detection method and a storage medium for defects of an AC filter, relating to the technical field of image processing. The detection method for defects of an AC filter includes: acquiring an image of the AC filter to be detected; inputting the image of the AC filter to be detected into a trained defect detection model for defect detection to obtain a detection result; the defect detection model includes a feature extraction module, a feature fusion module, and a defect detection module, the feature extraction module is used for extracting features from the image of the AC filter to be detected to obtain multiple-scale features; the feature fusion module is used for enhancing features of at least one scale, and performing fusion processing on the enhanced features and the unenhanced features by means of affine transformation; the defect detection module is used for performing defect detection according to the features after fusion processing. This detection method can improve the detection efficiency of defects of the AC filter and meet the detection requirements for large-target and small-target defects existing in the AC filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method for detecting defects of an AC filter and a storage medium. Background Art

[0002] With the development of smart grids, the condition monitoring and defect detection of AC filters have become important guarantees for the safe operation of the power grid. As a key device in the DC transmission system, the AC filter is mainly used to suppress harmonic interference and improve power quality. Since AC filters mostly operate outdoors and are exposed to long-term wind, sun, and rain, they are prone to component failures, mechanical damages, connector aging and other defects, and the defect appearances are as shown in the positions circled by the black boxes in Figure 1 , Figure 2 , Figure 3 . These defects may not only lead to increased harmonics and unbalanced three-phase currents, but also affect the stability of the transmission system and even cause serious problems such as equipment tripping.

[0003] Currently, the detection methods for AC filter defects mainly include manual inspection, infrared temperature measurement, and protection alarm. Among them, manual inspection is limited by low frequency and strong randomness, making it difficult to detect problems in a timely manner; although infrared temperature measurement can effectively detect hot spots on the surface of the AC filter, it cannot detect external defects of the AC filter; while the protection alarm is often triggered only after the operation of the AC filter is affected. Therefore, it is particularly important to research and apply technical means that can identify AC filter defects in the early stage.

[0004] In recent years, object detection technology has been widely applied in the field of industrial inspection, providing a new solution for equipment status monitoring in complex scenarios. Traditional object detection methods are based on manually designed features and classifiers, such as HOG (Histogram of Oriented Gradients) + SVM (Support Vector Machine) and Cascade models. Although this method performs well in simple scenarios, its detection ability for complex equipment defects is limited. With the development of deep learning technology, object detection methods based on CNN (Convolutional Neural Networks), such as Faster R-CNN, YOLO (You Only Look Once), and SSD (Single Shot MultiBox Detector), have rapidly emerged, demonstrating powerful feature extraction and object localization capabilities. Therefore, in equipment defect detection, object detection technology can fully utilize its processing ability for complex scenarios and subtle features. Combining a large amount of visible light image defect data, deep learning models can automatically extract defect features, such as foreign objects, damages, and rust, and quickly locate and classify them. By introducing object detection technology, intelligent monitoring of equipment operating status can be achieved, significantly improving the detection efficiency and accuracy, and providing technical support for the stable operation of the power system.

[0005] However, directly applying existing deep learning-based defect detection algorithms to AC filter defect detection has the following drawbacks: (1) The Faster R-CNN method needs to generate candidate regions first and then classify these candidate regions and perform fine-grained bounding box regression. This cumbersome working mode results in low real-time performance. (2) The YOLO algorithm needs to use NMS (Non-Maximum Suppression) to sort a large number of detection boxes and calculate overlapping regions, which will affect the speed of inferring defects in actual applications and is not conducive to the real-time requirements of the model, especially in the case of AC filters with a large number of targets to be detected. (3) SSD (Single Shot MultiBox Detector) relies on high-level features in the network for small target detection, and high-level features have a large receptive field, resulting in insufficient ability to capture detailed information of small targets. Small targets may project onto very small regions or even be lost on the feature map. Therefore, for the detection task of AC filters where both large and small target defects may appear simultaneously, it cannot meet the detection requirements. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent. To this end, the object of the present invention is to provide a method and a storage medium for detecting defects in AC filters, so as to improve the detection efficiency of AC filter defects and meet the detection requirements for large and small target defects in AC filters.

[0007] In a first aspect, an embodiment of the present invention provides a method for detecting defects in an AC filter. The method includes: obtaining an image of the AC filter to be detected; inputting the image of the AC filter to be detected into a trained defect detection model for defect detection to obtain a detection result. Wherein, the defect detection model includes a feature extraction module, a feature fusion module, and a defect detection module. The feature extraction module is configured to extract features from the image of the AC filter to be detected to obtain N features of different scales, where N is an integer greater than or equal to 2. The feature fusion module is configured to perform enhancement processing on features of at least one scale, and perform fusion processing on the enhanced features and the unenhanced features by means of affine transformation. The defect detection module is configured to perform defect detection based on the features after fusion processing.

[0008] Exemplarily, the training process of the defect detection model includes: obtaining a number of positive sample images and a number of negative sample images of various types of defects in the AC filter; randomly dividing the positive sample images and the negative sample images into a training set and a test set according to a first preset ratio, where the first preset ratio is greater than 1; constructing the defect detection model, and training and testing the defect detection model by using the training set and the test set to obtain the trained defect detection model.

[0009] Exemplarily, the feature fusion module includes a spatial / channel interaction self-attention sub-module and a cross-scale fusion sub-module. The feature fusion module performs enhancement processing on features of at least one scale, and performs fusion processing on the enhanced features and the unenhanced features by means of affine transformation, including: the spatial / channel interaction self-attention sub-module respectively performs enhancement processing on the features of the smallest scale among the N-scale features through spatial self-attention and channel self-attention to obtain enhanced features, where the number of self-attention heads used by the spatial self-attention and the channel self-attention is different; the cross-scale fusion sub-module performs fusion processing on the enhanced features and the unenhanced features by means of affine transformation.

[0010] Exemplarily, the spatial / channel interaction self-attention sub-module includes a feature mapping unit, a multi-head division unit, a spatial self-attention unit, a channel self-attention unit, and a first splicing unit; the spatial / channel interaction self-attention sub-module enhances the features of the smallest scale among the N-scale features through spatial self-attention and channel self-attention respectively, including: the feature mapping unit maps the features of the smallest scale into query features, key features, and value features using a preset convolution kernel; the multi-head division unit groups the query features, key features, and value features according to the number of attention heads in a second preset ratio to obtain a first group and a second group, where the second preset ratio is not equal to 1; the spatial self-attention unit obtains a first enhanced sub-feature based on the query features, key features, and value features of the first group according to spatial self-attention; the channel self-attention unit obtains a second enhanced sub-feature based on the query features, key features, and value features of the second group according to channel self-attention; the first splicing unit splices the first enhanced sub-feature and the second enhanced sub-feature to obtain the enhanced feature.

[0011] Exemplarily, the spatial self-attention unit obtains the first enhanced sub-feature through the following formula:

[0012]

[0013]

[0014] Where represents the enhanced sub-feature of the i-th attention head in the first group, represents the normalization function, represents the cosine similarity function, T represents the transpose, 、 、 respectively represent the features of the query feature, key feature, and value feature of the i-th attention head in the first group after a preset-size window shift operation in space, represents the constant coefficient, B represents the relative position bias, represents the first enhanced sub-feature, represents the 1×1 convolution, inverse window shift, and shape transformation operations, represents the channel splicing function, represents the number of attention heads corresponding to the first group.

[0015] Exemplarily, the channel self-attention unit obtains the second enhanced sub-feature through the following formula:

[0016]

[0017]

[0018] Among them, represents the enhancer feature of the i-th attention head in the second group, represents the normalization function, represents the cosine similarity function, T represents the transpose, , , respectively represent the query feature, key feature, and value feature of the i-th attention head in the second group, represents a constant coefficient, represents the second enhancer feature, represents the 1×1 convolution and shape transformation operation, represents the channel concatenation function, represents the number of attention heads corresponding to the first group, and L represents the total number of attention heads.

[0019] Exemplarily, the first splicing unit obtains the enhanced feature through the following formula:

[0020]

[0021] Among them, represents the enhanced feature, represents the 3×3 convolution, represents the channel concatenation function, represents the first enhancer feature, represents the second enhancer feature.

[0022] Exemplarily, the cross-scale fusion sub-module includes a second splicing unit, N - 1 first affine fusion units, a second affine fusion unit, and a third affine fusion unit. The N - 1 first affine fusion units are connected in sequence and correspond one by one to N - 1 features except the smallest-scale feature. The first first affine fusion unit is used to fuse the enhanced feature and the feature of the corresponding scale. The first affine fusion units except the first one are used to fuse the feature output by the previous first affine fusion unit and the feature of the corresponding scale. The second affine fusion unit is used to fuse the features output by the last two first affine fusion units. The third affine fusion unit is used to fuse the enhanced feature and the feature output by the second affine fusion unit; the first affine fusion unit, the second affine fusion unit, and the third affine fusion unit all perform fusion processing on the two features to be fused , through the following formula:

[0023]

[0024]

[0025]

[0026]

[0027] Among them, represents a 3×3 convolution, represents a 1×1 convolution, represents the learned reflection transformation coefficient, represents the learned shift coefficient, represents a convolution with a 3×3 convolution kernel for learning of represents a convolution with a 3×3 convolution kernel for learning of The corresponding scale is smaller than the corresponding scale, represents the preliminary fusion feature, represents the non-parametric instance normalization function, represents the feature and the feature of the fused feature;

[0028] The second splicing unit is used to splice the fused features output by the last first affine fusion unit, the second affine fusion unit, and the third affine fusion unit to obtain the final fused feature.

[0029] Exemplarily, the feature extraction module adopts the EfficientViT model, and the defect detection module adopts the RT-DETR network.

[0030] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the method for detecting AC filter defects described in the first aspect above.

[0031] The method for detecting AC filter defects and the storage medium according to the embodiments of the present invention divide the detection process of AC filter defects into three stages through a trained defect detection model: the first stage is the feature extraction stage, in which multiple different-scale features of the AC filter image to be detected can be extracted; the second stage is feature enhancement and fusion. In terms of feature enhancement, at least one scale of feature enhancement processing is realized to improve the discrimination intensity of the features; in terms of feature fusion, the enhanced features and other unenhanced features are fused, which can make full use of feature semantics and details to mine large-target and small-target features; the third stage is the prediction stage, in which defect recognition and localization are completed according to the fused features. Thus, the detection efficiency of AC filter defects can be improved, and the detection requirements for large-target and small-target defects in AC filters can be met. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a schematic diagram of a defect of an exemplary AC filter;

[0033] Figure 2 is a schematic diagram of a defect of another exemplary AC filter;

[0034] Figure 3 is a schematic diagram of a defect of yet another exemplary AC filter;

[0035] Figure 4 is a schematic structural diagram of a defect detection model according to an embodiment of the present invention;

[0036] Figure 5 is a schematic structural diagram of a spatial / channel interaction self-attention sub-module according to an embodiment of the present invention;

[0037] Figure 6 is a schematic structural diagram of a spatial self-attention unit according to an embodiment of the present invention;

[0038] Figure 7 is a schematic structural diagram of a channel self-attention unit according to an embodiment of the present invention;

[0039] Figure 8 is a schematic diagram of a structure for realizing cross-scale feature fusion according to an embodiment of the present invention;

[0040] Figure 9 is a schematic diagram of an mAP curve of an example of the present invention;

[0041] Figure 10 is a schematic diagram of an mAP curve of another example of the present invention. Detailed implementation manners

[0042] Aiming at the low detection efficiency of the existing recognition model and the inability to meet the detection requirements for large and small target defects in AC filters, the present invention provides a detection method and a storage medium for AC filter defects.

[0043] The following describes the detection method and the storage medium for AC filter defects according to the embodiments of the present invention with reference to the accompanying drawings.

[0044] In the embodiments of the present invention, the detection method for AC filter defects includes the following steps S21-S22:

[0045] S21, obtaining an AC filter image to be detected.

[0046] Exemplarily, the AC filter image to be detected can be obtained by a monitoring camera at a corresponding position, or can be obtained by a drone carrying a camera.

[0047] S22. Input the AC filter image to be detected into the trained defect detection model for defect detection to obtain the detection result.

[0048] In this embodiment, as Figure 4 shown, the defect detection model includes a feature extraction module, a feature fusion module, and a defect detection module. The feature extraction module is used to extract features from the AC filter image to be detected, obtaining N different-scale features, where N is an integer greater than or equal to 2, such as 3, 4, etc.; the feature fusion module is used to enhance at least one scale of features and perform fusion processing on the enhanced features and the unenhanced features by means of affine transformation; the defect detection module is used to perform defect detection based on the fused features.

[0049] Specifically, the defect detection model can be configured in electronic devices such as intelligent terminals and host computers. This electronic device can communicate with the aforementioned monitoring camera to obtain the AC filter image for detection. Then, the AC filter image to be detected can be input into the trained defect detection model for defect detection to obtain the detection result. Refer to Figure 4 , when using the defect detection model for defect detection, first, the feature extraction module extracts features from the AC filter image to be detected, obtaining N-scale features, such as obtaining three-scale features , , . By extracting multi-scale features, the accuracy and robustness of the defect detection model for defect detection can be improved. Then, the feature fusion module enhances at least one scale of features, such as high-level features (corresponding to the smallest scale) to improve the recognition of various defect features for subsequent detection; and performs cross-scale fusion on the enhanced and unenhanced different-scale features, fully mining the defect features of large targets and small targets by integrating defect features of different scales. Finally, the defect detection module realizes the rapid decoupling of the fused features, and finally realizes defect detection, which can include detecting whether there are defects, defect types, defect positions, etc.

[0050] Thus, this detection method can improve the detection efficiency of AC filter defects and meet the detection requirements for AC filters with large-target and small-target defects.

[0051] In some examples, the feature extraction module can use the EfficientViT model.

[0052] In this example, EfficientViT, as the feature extraction backbone network in the object detection task, has significant advantages. First of all, it has excellent computational and memory efficiency. By reducing the redundant operations of MHSA (Multi-Head Self-Attention) and adopting CGA (Cascaded Group Attention), it effectively improves the computational speed and feature extraction ability. Secondly, its combination of multi-scale feature fusion and efficient lightweight encoders enables it to perform excellently in processing multi-scale object detection, especially in resource-constrained scenarios (such as drone and satellite platforms) where it can ensure detection accuracy and speed.

[0053] Specifically, the EfficientViT model combines a lightweight multi-scale attention mechanism and MBConv (Mixed Convolution) blocks to achieve efficient performance in image processing. The lightweight multi-scale attention module is used to extract context information to assist the model in understanding the global information in the image; the MBConv block is mainly used to extract and utilize local information. The global receptive field and multi-scale learning are very important for the object detection task and can significantly improve the performance of the model. Therefore, using EfficientViT as the backbone network for AC filter defect detection can achieve an effective balance among detection speed, accuracy, and model complexity.

[0054] Exemplarily, the feature extraction module can also adopt but is not limited to methods such as image pyramid, multi-scale convolutional layer, SSD (Single Shot MultiBox Detector), FPN (Feature Pyramid Network), etc. to extract N different scale features.

[0055] In some examples of the present invention, the defect detection module can adopt the RT-DETR (Real-Time DEtectionTRansformer) network as the head network of the defect detection model.

[0056] Specifically, RT-DETR is a real-time end-to-end object detection model based on the Transformer architecture, which not only has efficient and fast real-time object detection performance but also can maintain high accuracy.

[0057] Exemplarily, the defect detection module can also adopt but is not limited to the YOLO network, DETR network, etc.

[0058] In some embodiments of the present invention, the training process of the defect detection model includes the following steps S41 - S43:

[0059] S41. Obtain a number of positive sample images and a number of negative sample images of various types of defects of the AC filter.

[0060] Among them, the positive sample image refers to an image containing defects of the AC filter (such as breakage, rust, foreign objects, etc.), and the negative sample image refers to an image without defects of the AC filter. The positive and negative sample images can be obtained from the network or by taking pictures with a camera.

[0061] S42. Randomly divide the positive sample images and negative sample images into a training set and a test set according to a first preset ratio.

[0062] Among them, the first preset ratio is greater than 1, such as 7:3, 8:2, etc., and can be specifically set according to needs. Exemplarily, the labels involved in the training set and the test set may include whether there are defects in the AC filter, defect types, defect locations, etc.

[0063] S43. Construct a defect detection model, and use the training set and the test set to train and test the defect detection model to obtain a trained defect detection model.

[0064] Optionally, the test set may not be set, and the defect detection model may be directly trained by the training set.

[0065] In some embodiments of the present invention, as Figure 4 shown, the feature fusion module includes a spatial / channel interaction self-attention sub-module and a cross-scale fusion sub-module.

[0066] In this embodiment, the feature fusion module enhances the features of at least one scale, and uses an affine transformation method to fuse the enhanced features and the unenhanced features, including: the spatial / channel interaction self-attention sub-module respectively enhances the features of the smallest scale (such as Figure 4 in ) among the N-scale features through spatial self-attention and channel self-attention to obtain enhanced features, where the number of self-attention heads used in spatial self-attention and channel self-attention is different; the cross-scale fusion sub-module uses an affine transformation method to fuse the enhanced features and the unenhanced features.

[0067] Specifically, to improve high-level vision tasks such as defect detection, it is crucial to refer to the feature distribution in the entire image (global or non-global context). Enhancing the semantic information of features can effectively improve the distinctiveness of AC filter defects, enhance the separation between defect categories, and improve the classification performance of the model for different defects. To this end, in order to obtain high-level semantic features with local and global context, the present invention designs a spatial / channel interaction self-attention sub-module. This sub-module enables the high-level features (i.e., the features at the smallest scale) output by the backbone network to interact within the scale through the combined interaction of spatial self-attention and channel self-attention, enhancing its ability to express local and global context information, so as to enhance the semantic information of the features, thereby effectively improving the distinctiveness of AC filter defect features (i.e., enhancing the discrimination intensity of the features), enhancing the separation between defect categories, and improving the classification performance of the model for different defects.

[0068] Since high-level features have more abstract semantic information while low-level features have more detailed information, simply splicing and fusing high-level and low-level features after upsampling and downsampling will not only affect the detailed information but also cause repetition and confusion between high-level and low-level features. Therefore, the present invention designs a cross-scale fusion sub-module based on affine transformation, which provides adaptively modified image features for the entire fusion process and reduces unnecessary impacts. Through cross-scale feature fusion based on affine transformation, features of large and small targets can be fully exploited, providing valuable feature guidance for defect recognition of different scales of the final AC filter.

[0069] Exemplarily, referring to Figure 4 , before enhancing and fusing the features, the N-scale features can also be respectively subjected to convolutional processing, such as using a 3×3 convolutional kernel, which can accelerate the speed of model training and detection.

[0070] In some examples, as Figure 5 shown, the spatial / channel interaction self-attention sub-module includes a feature mapping unit, a multi-head division unit, a spatial self-attention unit, a channel self-attention unit, and a first splicing unit. This spatial / channel interaction self-attention sub-module focuses on local and global context respectively by operating spatial self-attention (S-SA) and channel self-attention (C-SA) in parallel.

[0071] In this example, the spatial / channel interaction self-attention sub-module enhances the features of the smallest scale among the N-scale features through spatial self-attention and channel self-attention respectively, including: the feature mapping unit maps the features of the smallest scale into query features, key features, and value features using a preset convolution kernel (such as 3×3); the multi-head division unit groups the query features, key features, and value features according to the number of attention heads in the second preset ratio, obtaining a first group and a second group, where the second preset ratio is not equal to 1, such as 6:4, so that the number of self-attention heads used by spatial self-attention and channel self-attention is different; the spatial self-attention unit obtains a first enhanced sub-feature based on spatial self-attention according to the query features, key features, and value features of the first group; the channel self-attention unit obtains a second enhanced sub-feature based on channel self-attention according to the query features, key features, and value features of the second group; the first splicing unit splices the first enhanced sub-feature and the second enhanced sub-feature to obtain an enhanced feature.

[0072] As an implementation, the spatial self-attention unit obtains the first enhanced sub-feature through the following formulas (1) and (2):

[0073] (1)

[0074] (2)

[0075] Where represents the enhanced sub-feature of the i-th attention head in the first group, represents the normalization function, represents the cosine similarity function, T represents the transpose, 、 、 respectively represent the features of the query feature, key feature, and value feature of the i-th attention head in the first group after the preset size (which can be set as needed, such as 8×8) window shift operation in space, represents the constant coefficient (the value can be calibrated as needed, such as 0.01), B represents the relative position bias, represents the first enhanced sub-feature, represents the 1×1 convolution, inverse window shift, and shape transformation operation (i.e., the reshape operation, used to transform a specified matrix into a matrix of a specific dimension), represents the channel splicing function, represents the number of attention heads corresponding to the first group.

[0076] Specifically, taking N = 3 as an example, referring to Figure 4 、 Figure 5 , the high-level features output by the feature extraction module are input to the feature mapping unit. The feature mapping unit uses a 3×3 convolution to map the features It is mapped into query feature (Q), key feature (K), and value feature (V), and then the multi-head division unit divides the attention heads (the total number is ) in a ratio of 6:4.

[0077] The implementation structure of S-SA can be as Figure 6 shown. The calculation process of S-SA can be described as follows: First, , and are subjected to an 8×8 window shifting operation in space. After that, the features are rearranged and split into vectors , and . The cosine similarity is calculated using and and added to the relative position bias . Then, after normalization by , the weights are obtained. The obtained weights are multiplied by for reweighting. For the reweighted feature vectors, they are first passed through a 1×1 convolution, then an inverse window shifting operation, and then a reshape operation to obtain the feature . The above process can be represented by the above formulas (1) and (2).

[0078] As an implementation, the channel self-attention unit obtains the second enhancer feature through the following formulas (3) and (4):

[0079] (3)

[0080] (4)

[0081] Among them, represents the enhancer feature of the i-th attention head in the second group, represents the normalization function, represents the cosine similarity function, T represents the transpose, , , respectively represent the query feature, key feature, and value feature of the i-th attention head in the second group, represents the constant coefficient, represents the second enhancer feature, represents the 1×1 convolution and reshape operations, represents the channel concatenation function, represents the number of attention heads corresponding to the first group, and L represents the total number of attention heads.

[0082] Since an 8×8 window shifting operation is used in the calculation process of S-SA, each , and is restricted within an 8×8 window, so the receptive field of S-SA is limited and only contains local context information. While C-SA directly uses the obtained 、 and to participate in attention calculation, so C-SA contains global context information. The implementation structure of C-SA can be as shown in Figure 7 and can be represented by the above formulas (3) and (4).

[0083] As an implementation manner, the first splicing unit obtains the enhanced feature through the following formula (5):

[0084] (5)

[0085] wherein, represents the enhanced feature, represents a 3×3 convolution, represents the channel splicing function, represents the first enhancer feature, represents the second enhancer feature.

[0086] Specifically, referring to Figure 5 , after obtaining the first enhancer feature and the second enhancer feature , and are spliced on the channel, and the local features and global features are mixed through a 3×3 convolution, and finally the spatial / channel attention is obtained, and the implementation process can be represented by the above formula (5). has high-level semantic features with local and global contexts, so it can effectively improve the distinguishability of AC filter defects and enhance the separation degree between defect categories.

[0087] In some embodiments of the present invention, referring to Figure 4 (illustrated with N = 3 as an example), the cross-scale fusion sub-module includes a second splicing unit, N - 1 first affine fusion units, a second affine fusion unit, and a third affine fusion unit. The N - 1 first affine fusion units are connected in sequence and correspond one by one to N - 1 features except the smallest-scale feature. The first first affine fusion unit is used to fuse the enhanced feature and the feature of the corresponding scale. The first affine fusion units except the first first affine fusion unit are used to fuse the feature output by the previous first affine fusion unit and the feature of the corresponding scale. The second affine fusion unit is used to fuse the features output by the last two first affine fusion units. The third affine fusion unit is used to fuse the enhanced feature and the feature output by the second affine fusion unit.

[0088] Among them, the first affine fusion unit, the second affine fusion unit, and the third affine fusion unit all perform fusion processing on the two features to be fused through the following formula , as follows:

[0089] (6)

[0090] (7)

[0091] (8)

[0092] (9)

[0093] Among them, represents a 3×3 convolution, represents a 1×1 convolution, represents the learned reflection transformation coefficient, represents the learned shift coefficient, represents a convolution with a convolution kernel size of 3×3 for learning , represents a convolution with a convolution kernel size of 3×3 for learning . The corresponding scale is smaller than the scale corresponding to , represents the preliminary fusion feature, represents the non-parametric instance normalization function, represents the feature and the feature of the fusion feature.

[0094] Specifically, see Figure 4, taking N = 3 as an example, two first affine fusion units are connected successively from bottom to top, the second affine fusion unit and the third affine fusion unit are connected successively from top to bottom, and the second affine fusion unit is connected to the two topmost first affine fusion units. First, the size and number of channels of the enhanced feature p5 are changed using Conv1 and the upsampling operation. Then, the first first affine fusion unit is used to fuse p4 and p5. At this time, the output of the first first affine fusion unit contains the information of both p4 and p5 features, and the output is denoted as (p4 + p5). Then, the above operation is repeated to fuse p3 using the second first affine fusion unit, and the output is denoted as (p3 + p4 + p5). Considering that the object detection type framework has a greater demand for high-level semantic features, that is, features with smaller scales, the second affine fusion unit is used to fuse (p4 + p5) with (p3 + p4 + p5) (first use the Conv3 operation, that is, use a convolution with a stride of 2 and a kernel size of 3×3 to adjust the size), and the output is denoted as [(p4 + p5)+(p3 + p4 + p5)]. And so on, the third affine fusion unit fuses [(p4 + p5)+(p3 + p4 + p5)] with the un-upsampled p5 after the Conv1 operation. It should be noted that Figure 4 Conv1 and Conv3 in the shown cross-scale fusion sub-module are convolutions with a kernel size of 1×1 and 3×3 using SiLU as the activation function respectively. Among them, Conv1 is used to modify the number of channels, and Conv3 with a stride of 2 is used to modify the feature map size.

[0095] During the process of feature fusion in each affine fusion unit, the affine transformation parameters , are learned from the semantic prior, rather than using the image or semantic map as the conditional input, so as to reduce the risk of repetition and confusion between high-level features and low-level features. The structure of the affine fusion unit is as Figure 8 shown. First, non-parametric instance normalization ) is used to normalize the input low-level feature (the low-level feature here is relative, such as relative to is the low-level feature). Then two different parameter sets (the semantic prior is the high-level feature, and the high-level feature is also relative, such as relative to is the high-level feature) are learned from the semantic prior , (which can be represented by the above formulas (6) and (7)) in order to perform a spatial pixel affine transformation on the low-level feature (which can be represented by the above formula (8)).

[0096] Feature Although the detailed and semantic information is available, in order to further extract local and global information and considering the efficiency of the integration and fusion process, the present invention uses simple convolutions to implement this process. Considering that smaller convolution kernels are suitable for extracting detailed features and local textures, while larger convolution kernels are more suitable for global image information. Therefore, are respectively input into two paths. The first path uses a single 1×1 convolution operation to extract local features; the second path uses a single 1×1 convolution operation and three 3×3 convolution operations respectively to achieve deeper information extraction. Finally, the local features and global features are added together to obtain the final fusion features at two scales . The implementation process can be represented by the above formula (9).

[0097] The second splicing unit (represented by a circle containing C in Figure 4 ) is used to splice the fusion features output by the last first affine fusion unit, the second affine fusion unit, and the third affine fusion unit to obtain the final fusion features. Refer to Figure 4 , finally, three different (that is, , , , all of which may contain high-level semantic features) are spliced and input into the defect detection module based on RT-DETR to achieve the final defect detection. Therefore, the final cross-scale fusion sub-module based on affine transformation adaptively fuses high / low-level features at three different scales, fully extracts the features of large and small targets, and provides valuable feature guidance for the final defect detection.

[0098] In some other embodiments of the present invention, the cross-scale fusion sub-module includes a second splicing unit, N-1 first affine fusion units, and N-1 fourth affine fusion units. The connection structure (connected from bottom to top) and functions of the N-1 first affine fusion units are the same as those of the N-1 first affine fusion units in the above embodiments. The N-1 fourth affine fusion units are connected in sequence (connected from top to bottom) and correspond to the N-1 first affine fusion units one by one. The first fourth affine fusion unit is used to fuse the features output by the corresponding first affine fusion unit (after Conv3 operation) and its input features (the features that have not been upsampled after Conv1 operation), and the other fourth affine fusion units are used to fuse the features that have not been upsampled and input below the corresponding first affine fusion unit and the features output by the previous fourth affine fusion unit (after Conv3 operation). The second splicing unit is used to splice the fusion features output by the last first affine fusion unit, the first fourth affine fusion unit, and the last fourth affine fusion unit.

[0099] Exemplarily, to reduce the processing complexity, the feature The second splicing unit is directly used for splicing, that is, without going through the process of the above formula (9).

[0100] The following combines Figure 9 and Figure 10 , and through the tracking results of mAP (mean Average Precision) during the training process, the detection performance of the defect detection model designed by the present invention is described;

[0101] Figure 9 and Figure 10 respectively show the change curves of mAP50 and mAP50:90 during the training process. The abscissa is the epoch, representing the training rounds of the defect detection model, and the ordinate is the mAP value. It can be seen that the change curves of mAP50 and mAP50:90 can quickly reach a stable state. Among them, mAP50 can be maintained above 0.75, and mAP50:90 can be maintained above 0.50, indicating that the average precision of the designed defect detection model for different defect categories of AC filters is very impressive.

[0102] It should be noted that in the object detection task, mAP is used to measure the average precision of the model under different categories and different IoU (Intersection over Union) thresholds. Among them, mAP50 is to measure the average precision of the model when the IoU threshold is 0.5; mAP50:95 is to measure the average precision of the model in the range of IoU thresholds from 0.5 to 0.95.

[0103] Based on the above-described method for detecting AC filter defects in the embodiments, the present invention also proposes a computer-readable storage medium.

[0104] In this embodiment, a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method for detecting AC filter defects in the above-described embodiments is implemented.

[0105] In summary, for the method and storage medium for detecting AC filter defects according to the embodiments of the present invention, when detecting AC filter defects, the defect detection model adopted divides the detection process into three stages: the first stage is the feature extraction stage, in which image feature extraction is implemented based on the EfficientViT model. This network containing the Transfomer structure can achieve efficient feature extraction. The second stage is the feature enhancement and fusion stage. In terms of feature enhancement, a Transformer-style spatial / channel interaction self-attention is designed. By using different numbers of spatial and channel self-attentions, interaction within the high-level feature scale is realized, and high-level semantic features with local and global contexts are extracted to enhance the discrimination intensity of the features. In terms of feature fusion, cross-scale fusion based on affine transformation is designed, and the affine transformation method is used to achieve efficient fusion of high-level and low-level features, making full use of feature semantics and details to mine large-target and small-target features, providing valuable feature guidance for final defect detection and being conducive to improving the detection accuracy of the model. The third stage is the prediction stage, in which the head network of RT-DETR is used to complete defect recognition and localization. This network containing the Transfomer structure can quickly process multi-scale defect features by decoupling spatial / channel interaction self-attention and cross-scale fusion based on affine transformation, improving the detection speed and accuracy and being adaptable to AC filter defects with different scales.

[0106] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0107] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0108] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0109] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0110] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0111] In the present invention, unless otherwise clearly specified or limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed broadly. For example, it may be a fixed connection, a detachable connection, or an integral body; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0112] In the present invention, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0113] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for detecting defects in an AC filter, characterized in that: The method comprises: Acquire an image of the AC filter to be tested; Inputting the AC filter image to be inspected into a trained defect detection model to perform defect detection and obtain a detection result; The defect detection model includes a feature extraction module, a feature fusion module and a defect detection module. The feature extraction module is used to extract features of the AC filter image to be detected to obtain features of N different scales, where N is an integer greater than or equal to 2; the feature fusion module is used to enhance features of at least one scale, and fuse the enhanced features with the unenhanced features by affine transformation; the defect detection module is used to perform defect detection based on the fused features; The feature fusion module includes a spatial / channel interactive self-attention submodule and a cross-scale fusion submodule; the feature fusion module enhances features of at least one scale, and fuses the enhanced features with the unenhanced features by using an affine transformation, including: The spatial / channel interactive self-attention submodule enhances the smallest scale feature among N scale features by spatial self-attention and channel self-attention respectively to obtain enhanced features, wherein the number of self-attention heads used by the spatial self-attention and the channel self-attention is different; the cross-scale fusion submodule fuses the enhanced features with the unenhanced features by affine transformation; The spatial / channel interactive self-attention submodule includes a feature mapping unit, a multi-head division unit, a spatial self-attention unit, a channel self-attention unit and a first splicing unit; the spatial / channel interactive self-attention submodule enhances the smallest scale feature among N scale features through spatial self-attention and channel self-attention, respectively, including: The feature mapping unit maps the minimum scale feature into a query feature, a key feature and a value feature using a preset convolution kernel; The multi-head division unit groups the query features, key features, and value features according to a second preset ratio of attention heads to obtain a first group and a second group, wherein the second preset ratio is not equal to 1; The spatial self-attention unit obtains the first enhancer feature by the following formula: in, represents the enhancer feature of the i-th attention head in the first group, represents the normalization function, represents the cosine similarity function, T represents the rank transformation, , , They respectively represent the query features, key features, and value features of the i-th attention head in the first group after the preset size window shift operation is performed in space, represents the constant coefficient, B represents the relative position bias, represents the first enhancer feature, represents 1×1 convolution, inverse window shift and shape transformation operations, represents the channel splicing function, represents the number of attention heads corresponding to the first group; The channel self-attention unit obtains the second enhancer feature by the following formula: in, represents the enhancer feature of the i-th attention head in the second group, represents the normalization function, represents the cosine similarity function, T represents the rank transformation, , , denote the query feature, key feature, and value feature of the i-th attention head in the second group, respectively. represents the constant coefficient, represents the second enhancer feature, represents a 1×1 convolution and shape transformation operation, represents the channel splicing function, represents the number of attention heads corresponding to the first group, and L represents the total number of attention heads; The first splicing unit obtains the enhanced feature by the following formula: in, represents the enhanced feature, represents 3×3 convolution, represents the channel splicing function, represents the first enhancer feature, represents the second enhancer feature; The cross-scale fusion submodule includes a second splicing unit, N-1 first affine fusion units, a second affine fusion unit and a third affine fusion unit, the N-1 first affine fusion units are connected in sequence and correspond one-to-one to the N-1 features except the minimum scale feature, the first first affine fusion unit is used to fuse the enhanced feature and the feature of the corresponding scale, the first affine fusion unit except the first first affine fusion unit is used to fuse the feature output by the previous first affine fusion unit and the feature of the corresponding scale, the second affine fusion unit is used to fuse the features output by the last two first affine fusion units, and the third affine fusion unit is used to fuse the enhanced feature and the feature output by the second affine fusion unit; The first affine fusion unit, the second affine fusion unit and the third affine fusion unit all use the following formula to combine the two features to be combined , Perform fusion processing: in, represents 3×3 convolution, represents 1×1 convolution, represents the learned reflection transformation coefficient, represents the learned shift coefficient, Indicates that the convolution kernel size is 3×3 for learning The convolution of Indicates that the convolution kernel size is 3×3 for learning The convolution of The corresponding scale is smaller than The corresponding scale, represents the initial fusion feature, represents the non-parametric instance normalization function, Representation characteristics and Features The fusion characteristics of The second splicing unit is used to splice the fusion features output by the last first affine fusion unit, the second affine fusion unit and the third affine fusion unit to obtain the final fusion features.

2. The method for detecting defects of an AC filter according to claim 1, characterized in that: The training process of the defect detection model includes: Acquire a number of positive sample images and a number of negative sample images of multiple types of defects of AC filters; Randomly dividing the positive sample images and the negative sample images into a training set and a test set according to a first preset ratio, wherein the first preset ratio is greater than 1; The defect detection model is constructed, and the defect detection model is trained and tested using the training set and the test set to obtain the trained defect detection model.

3. The method for detecting defects of an AC filter according to claim 1 or 2, characterized in that: The feature extraction module adopts the EfficientViT model, and the defect detection module adopts the RT-DETR network.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting defects of an AC filter according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Multi-modal medical image fusion method based on significant information enhancement

    CN117333411A

  • Mini LED defect detection method, electronic equipment and medium

    CN119090851A