Photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention
By introducing feature pyramids and cascade group attention mechanisms into the YOLOv11 algorithm, the problem of high-resolution image processing in photovoltaic panel defect detection is solved, and the precise positioning and classification of photovoltaic panel defects is achieved, and the detection effect and recognition are improved.
Patent Information
- Application Number
- CN202510158109.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing PV panel defect detection technologies are difficult to accurately classify and locate PV panel defects in high-resolution infrared images, especially when processing complex features and high-dimensional data, computing resources are consumed and difficult to scale to larger data sets.
The improved YOLOv11 algorithm based on feature pyramids and cascade group attention is adopted, and multi-scale features are fused through feature pyramid structures, and a cascade group attention mechanism is introduced into the detection head to reduce redundant feature calculations and improve the refinement feature extraction ability.
It realizes the precise positioning and classification of photovoltaic panel defects, improves the recognition and effect of detection, can more effectively process high-resolution infrared images, and is suitable for defect detection of large-scale photovoltaic panels.
Smart Images

Figure CN119992213A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention. Background Art
[0002] Photovoltaic power generation converts solar energy into electrical energy through the photovoltaic effect, and is a renewable energy solution to energy crises and environmental pollution. Large-scale deployment of photovoltaic panels can effectively meet energy needs and promote sustainable development, but equipment aging and failure problems need to be solved urgently. Therefore, timely detection of photovoltaic panel failures is the key to ensuring their stable operation;
[0003] Defect detection of photovoltaic panels can be carried out by collecting data with handheld cameras. In recent years, with the popularity of drones, more industries have adopted drones for real-time shooting and infrared collection. The collected information is more, more complete, more stable, and of higher quality, which is conducive to providing data support for subsequent defect detection.
[0004] At present, the technology of photovoltaic panel defect detection is widely used. However, there is little research on how to accurately classify the detected faults. In actual situations, it is necessary not only to detect which photovoltaic panels have defects, but also to determine the type of defects, so as to know the severity of the specific defects and whether they can be repaired, so as to facilitate the subsequent repair or replacement of photovoltaic panels.
[0005] In the research of defect classification and detection of photovoltaic panels, the traditional method is based on machine vision to detect defects, using random forest (RF), support vector machine (SVM) and K-means clustering for classification processing, but the ability to process complex features and high-dimensional data is limited, not only it consumes a lot of computing resources, but also it is difficult to expand to a larger data set, limiting the feasibility in actual industrial applications. For example, in high-resolution infrared images of photovoltaic panels, the size, texture and other features of different categories of defects are obviously different, and traditional methods are easily interfered by redundant features, resulting in misjudgment.
[0006] Deep learning has shown powerful effects in the field of target detection in recent years due to its powerful computing and learning capabilities, among which the YOLO algorithm is the most representative. Since most of the defects in photovoltaic panels are small (such as debris cover and small area hot spots), the feature information is not rich enough, and the performance of the YOLO algorithm in small target detection is poor. Therefore, based on the YOLOv11 algorithm, the present invention makes adaptive improvements to some modules therein, so that the model pays more attention to the extraction of refined features and reduces redundant feature calculations on the basis of the fusion of multi-scale features, and at the same time uses the idea of task alignment to achieve better synergy in photovoltaic panel defect location and classification detection. Summary of the invention
[0007] 1. Technical issues to be resolved
[0008] In view of the deficiencies in the prior art, the present invention provides a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention to solve the above problems.
[0009] (II) Technical solution
[0010] To achieve the above object, the present invention provides the following technical solution: a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention, comprising the following steps:
[0011] S1. Data collection, annotation and preprocessing;
[0012] S2. Use the improved YOLOv11 algorithm to train the model;
[0013] S3. Predict the input infrared image, accurately locate the defects and mark the categories.
[0014] Preferably, the specific content of S1 is:
[0015] S1.1: Obtain infrared imaging images taken by the drone;
[0016] S1.2: Analyze the internal circuit of the photovoltaic panel and deal with the one-to-one correspondence between the fault type and the manifestation;
[0017] S1.3: Use annotation tools to carefully annotate defects and faults, including location information and classification groups;
[0018] S1.4: Use the opencv library to perform data enhancement operations such as rotation, brightness and contrast adjustment, and scaling on the photovoltaic panel images to generate more samples for subsequent model training.
[0019] Preferably, the S2 adaptively improves the YOLOv11 network, passes the image into the pre-trained network, and trains the model, and the specific content is:
[0020] S2.1: The detection module is divided into the backbone network, the feature pyramid neck, and the final detection head;
[0021] The downsampling layer in the backbone network uses multi-layer convolution to extract multi-dimensional features from the infrared-photographed photovoltaic panel images that are passed in, and uses the SPPF module to enhance the features of important information to obtain valuable defect information.
[0022] In the Neck part, a feature pyramid structure is constructed by introducing upsampling layers and multiple convolution modules, and connected with the multi-layer features of the backbone network to achieve multi-scale feature fusion;
[0023] The detection head contains several Detect modules, in which anchor boxes are drawn and redundant duplicate boxes are removed through non-maximum suppression. The maximum probability category is selected in the output probability distribution, and the defect conditions of small samples and large samples are output in multiple channels.
[0024] The specific structure of the backbone network is as follows:
[0025] S2.2: In the backbone network, multiple Conv convolutional layers C3k2 layers and SCC3k2 layers are alternately connected to build a deep network to continuously sample and extract information. For specific connection methods, see the attached Figure 2 The gray part in the figure is the improved module. The SCC3k2 module uses the spatial and channel reconstruction convolution SCConv to replace the convolution part in the traditional C3k2 module to enhance the deep feature extraction effect;
[0026] The SCConv module consists of two core parts, the spatial reconstruction unit (SRU) and the channel reconstruction unit (CRU), which are sequentially composed. The SRU is used to reduce the redundancy in the spatial dimension, and then the CRU is used to reduce the redundancy in the channel dimension. SCConv can be seamlessly integrated into the existing CNN to replace the standard convolution operation.
[0027] After several Conv convolutional layers, C3k2 layers and SCC3k2 layers are alternately connected, they are processed by SPPF parallel maximum pooling to further enhance the effective defect information;
[0028] At the end of the backbone network, C2WCGA is used to replace the original C2PSA module. The cascade group attention based on local window parallel processing is used to focus more on learning detailed features with rich information, thus improving the diversity of attention.
[0029] Since the input image has a high resolution, a local window attention mechanism is introduced before it is input into the module. It is used to split the width and height of the input into multiple sub-regions according to a certain resolution, so that the width and height are reduced, the channels remain unchanged, and the batch size is increased. After dividing multiple local windows, the cascade group attention mechanism is applied independently to each window. This method can reduce the amount of calculation, especially when processing high-resolution input, and can increase the model's attention to local features, thereby improving performance;
[0030] The CGA cascade group attention part under this module contains more attention heads, which perform attention calculations on different channels and perform cascade operations, that is, the output of the previous head is spliced to the input of the next head. This not only enhances the information flow within each window, but also promotes information fusion across windows, which is conducive to the model gradually generating a global understanding of the entire image;
[0031] S2.3: The core idea of feature pyramid FPN is to generate feature maps of multiple scales in different convolutional layers and enrich the representation by gradually fusing high-level and low-level features. In the neck layer, the UnSample upsampling layer refines and expands the features and concats them with the C3k2 and SCC3k2 modules in the backbone network to fuse high-level and low-level information to enrich the features. In addition, connecting different SCC3k2 modules at different feature output dimensions can further enhance the effect of multi-scale fusion.
[0032] S2.4: In the final Head layer, the Anchor-Free concept is adopted. Compared with Anchor-Based, Anchor-Free does not need to generate predefined anchor boxes, making the model more flexible and able to more effectively identify various types of defects in photovoltaic panels;
[0033] At the same time, this module adopts a decoupled approach, which can maximize the role of multiple Detect modules and output small sample and large sample defects on multiple channels respectively; in addition, the original Detect module is changed to the DTADetect module, and the corresponding dynamic task alignment (DynamicTask Align) structure design is performed on the detection head with reference to the TOOD (task-aligned single-stage target detection) idea, so as to enhance the collaborative performance of the two tasks of photovoltaic panel defect location and category prediction. The extractor can extract different refined features from different convolutional layers and interact with each other, thereby improving the performance of defect classification detection;
[0034] In this part, multiple anchor boxes are first generated, and then analyzed based on IOU (intersection over union), a key indicator for measuring the degree of overlap of prediction boxes. The larger the IOU value, the more serious the overlap of prediction boxes. Next, the NMS (non-maximum suppression) technology is used to remove redundant anchor boxes with high overlap to ensure the reliability of the final prediction results. On this basis, the corresponding defect categories are assigned to the selected labeled defect anchor boxes according to the generated category probabilities.
[0035] Preferably, the specific content of S3 is:
[0036] S3.1: specify a specified file directory and use the batch of photovoltaic panel infrared images in it as input;
[0037] S3.2: Specify the configuration file and load the model trained using the improved YOLOv11 algorithm;
[0038] S3.3: Traverse each image in the folder, use the model to classify and detect photovoltaic panel defects, and generate anchor boxes in the image to locate the position and size of the defects. Each anchor box will be marked with the category name and confidence level.
[0039] Preferably, the SCConv convolution in the SCC3k2 module is used to efficiently capture information in deep space and channels, reduce the extraction of redundant information and improve training efficiency;
[0040] The SRU module of SCConv includes separation and reconstruction operations to reduce the redundancy of spatial information. Although the traditional C3k2 reduces the computational complexity to a certain extent through the residual connection block, the data processing part still includes redundant feature information, which increases the computational complexity and brings unnecessary noise interference;
[0041] In the SRU separation operation, Group Normalization is used to extract the scaling factor of the feature map. The larger the variance, the more effective information about the photovoltaic panel defects. By comparing the values in the variance, the spatial distribution of the effective information can be known, and it can be divided into two parts according to a specific threshold. The information above the threshold is effective information, and the information below the threshold is redundant information. In order to further integrate these two parts of information, it is necessary to use the reconstruction operation to rearrange and integrate these features by cross reconstruction to enhance the expressiveness of the features;
[0042] The formula for the separation operation is as follows:
[0043] W=T(Sigmoid(W γ(GN(X)) ));
[0044] This formula indicates that in the separation operation, after the input X is processed by GroupNormalization, combined with the scaling factor W γ Feature extraction is performed, and then the Sigmoid activation function is applied to make the weight value smoother. Finally, the features are divided into rich features and redundant features through T threshold gating, and the final feature map weight W is obtained.
[0045] In the SRU module of SCC3k2, the threshold T is changed from the default 0.5 to 0.4 to extract more and richer photovoltaic panel defect features, because the defects caused by categories such as debris covering in photovoltaic panels are usually small targets on photovoltaic panels and there are more of them. If the threshold is not changed, there is a greater probability that they will be identified as redundant features.
[0046] The relevant formula for threshold allocation is as follows:
[0047]
[0048] Among them, W1 is the upper layer rich feature map, and W2 is the shallow layer redundant feature map.
[0049] Preferably, in the backbone network construction and feature extraction, the channel reconstruction unit CRU is divided into three steps: segmentation, transformation, and fusion;
[0050] In the segmentation step, CRU first obtains the spatially refined feature map obtained in the previous step and divides its channels into two parts, including αC channels and (1-α)C channels respectively. Then, the feature map is compressed through 1×1 convolution to improve computational efficiency.
[0051] In the transformation step, the separated upper half feature map is transformed by group-wise convolution (GWC) and point-wise convolution (PWC) to extract rich high-level features; the lower half is residually concatenated with the original shallow input through a cheap 1×1 convolution to supplement the shallow feature information;
[0052] In the fusion part, the upper and lower parts of the information are weighted and fused through global average pooling and the β parameter related to the soft attention mechanism to extract the final channel refined feature map. Compared with direct connection, fusion is more efficient in extracting information;
[0053] In the Bottleneck bottleneck structure originally included in C3k2, the second layer Conv is replaced with SCConv. In order to reduce the computational burden, the present invention chooses to replace only the convolutional layer in the second half of the model, thereby achieving performance improvement while effectively controlling the computational cost. Finally, the residual connection is used to form SCBottleneck to replace Bottleneck, forming an improved SCC3k2 module.
[0054] Preferably, a cascade group attention mechanism (CGA) is applied to enhance the feature representation step. The CGA module is used in the final C2WCGA of the backbone network. By performing different segmentations of the complete features on each head in the multi-head attention, each subtask can extract more accurate and refined features, and finally connect them to obtain more effective feature outputs to provide valuable context information.
[0055] Specifically, the input feature map is first spatially segmented into multiple local windows, and then multiple sub-regions are formed. In this process, the height and width dimensions are reduced while the number of channels remains unchanged, and the batch size is increased, which can effectively improve the efficiency of parallel computing.
[0056] Next, a multi-head attention mechanism is applied in parallel on each window. First, the input feature map passes through the QKV layer to calculate the query, key, and value, and then it is divided into multiple sub-features, and each sub-task is processed separately. Each head performs feature modeling through the self-attention mechanism to capture local context information. An attention bias is added to the attention module to guide the attention mechanism to focus more finely on different areas and features, improving the performance of the model. The specific formula is as follows:
[0057]
[0058] Among them, d k is the dimension of the key vector, which is used to scale the dot product to avoid excessive values. B is a bias term, which is used to adjust the similarity between the query Q and the key K. The softmax operation converts the similarity into weights, and weights the value vector V by these weights, and finally outputs the weighted value vector V. In the code implementation, a feat feature map array is introduced to record the multi-head output results obtained by parallel calculation. The output results of each head will be cascaded to the input of the next head in turn, and multiple feature fusions and information transfers will be performed to ensure effective information exchange between different heads. The specific cascade formula is as follows:
[0059]
[0060] Among them, the actual input features in the first layer are the same as the original input features, and the actual input features f of the i-th layer are i It is necessary to superimpose the current original input x i And the output of the previous layer Get it.
[0061] Beneficial Effects
[0062] Compared with the prior art, the present invention provides a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention, which has the following beneficial effects:
[0063] This is a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention. According to the internal circuit structure of the photovoltaic panel and the infrared imaging image, the defects on the photovoltaic panel are accurately located, labeled and classified.
[0064] The latest YOLOv11 target detection algorithm is used and improved, and some C3k2 modules are replaced with SCC3k2 to make it more accurate in capturing spatial and channel information and reduce the adverse effects of redundant information extraction;
[0065] C2WCGA is used to replace the original basic C2PSA to enhance the fusion performance of refined features, making the recognition of photovoltaic panel defect classification detection higher and the effect better;
[0066] In the detection head, the Detect module is improved to the DTADetect module, which learns the interactive features between multiple tasks to extract more refined information.
[0067] After the infrared image is processed by the trained model, it outputs an image with defect annotations and classification information, as well as processed table data, which makes it easier for photovoltaic panel staff to repair or replace the photovoltaic panels at the corresponding locations. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is an algorithm processing flow chart of a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention of the present invention;
[0069] Figure 2 This is a structural diagram of an improved YOLOv11 in a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention of the present invention;
[0070] Figure 3 This is a SCConv structure diagram in a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention in the present invention;
[0071] Figure 4 It is a schematic diagram of the photovoltaic panel defect classification detection effect in a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention in the present invention. DETAILED DESCRIPTION
[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0074] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0075] See also Figure 1-4 , a photovoltaic panel defect category detection algorithm based on feature pyramid and cascaded group attention, including the following steps:
[0076] S1. Data collection, annotation and preprocessing;
[0077] S2. Use the improved YOLOv11 algorithm to train the model;
[0078] S3. Predict the input infrared image, accurately locate the defects and mark the categories.
[0079] It should be noted that the specific content of S1 is:
[0080] S1.1: Obtain infrared imaging images taken by the drone;
[0081] S1.2: Analyze the internal circuit of the photovoltaic panel and deal with the one-to-one correspondence between the fault type and the manifestation;
[0082] S1.3: Use annotation tools to carefully annotate defects and faults, including location information and classification groups;
[0083] S1.4: Use the opencv library to perform data enhancement operations such as rotation, brightness and contrast adjustment, and scaling on the photovoltaic panel images to generate more samples for subsequent model training.
[0084] It should be noted that S2 adaptively improves the YOLOv11 network, passes the image into the pre-trained network, and trains the model. The specific content is:
[0085] S2.1: The detection module is divided into the backbone network, the feature pyramid neck, and the final detection head;
[0086] The downsampling layer in the backbone network uses multi-layer convolution to extract multi-dimensional features from the infrared-photographed photovoltaic panel images that are passed in, and uses the SPPF module to enhance the features of important information to obtain valuable defect information.
[0087] In the Neck part, a feature pyramid structure is constructed by introducing upsampling layers and multiple convolution modules, and connected with the multi-layer features of the backbone network to achieve multi-scale feature fusion;
[0088] The detection head contains several Detect modules, in which anchor boxes are drawn and redundant duplicate boxes are removed through non-maximum suppression. The maximum probability category is selected in the output probability distribution, and the defect conditions of small samples and large samples are output in multiple channels.
[0089] The specific structure of the backbone network is as follows:
[0090] S2.2: In the backbone network, multiple Conv convolutional layers C3k2 layers and SCC3k2 layers are alternately connected to build a deep network to continuously sample and extract information. For specific connection methods, see the attached Figure 2 The gray part in the figure is the improved module. The SCC3k2 module uses the spatial and channel reconstruction convolution SCConv to replace the convolution part in the traditional C3k2 module to enhance the deep feature extraction effect;
[0091] The SCConv module consists of two core parts, the spatial reconstruction unit (SRU) and the channel reconstruction unit (CRU), which are sequentially composed. The SRU is used to reduce the redundancy in the spatial dimension, and then the CRU is used to reduce the redundancy in the channel dimension. SCConv can be seamlessly integrated into the existing CNN to replace the standard convolution operation.
[0092] After several Conv convolutional layers, C3k2 layers and SCC3k2 layers are alternately connected, they are processed by SPPF parallel maximum pooling to further enhance the effective defect information;
[0093] At the end of the backbone network, C2WCGA is used to replace the original C2PSA module. The cascade group attention based on local window parallel processing is used to focus more on learning detailed features with rich information, thus improving the diversity of attention.
[0094] Since the input image has a high resolution, a local window attention mechanism is introduced before it is input into the module. It is used to split the width and height of the input into multiple sub-regions according to a certain resolution, so that the width and height are reduced, the channels remain unchanged, and the batch size is increased. After dividing multiple local windows, the cascade group attention mechanism is applied independently to each window. This method can reduce the amount of calculation, especially when processing high-resolution input, and can increase the model's attention to local features, thereby improving performance;
[0095] The CGA cascade group attention part under this module contains more attention heads, which perform attention calculations on different channels and perform cascade operations, that is, the output of the previous head is spliced to the input of the next head. This not only enhances the information flow within each window, but also promotes information fusion across windows, which is conducive to the model gradually generating a global understanding of the entire image;
[0096] S2.3: The core idea of feature pyramid FPN is to generate feature maps of multiple scales in different convolutional layers and enrich the representation by gradually fusing high-level and low-level features. In the neck layer, the UnSample upsampling layer refines and expands the features and concats them with the C3k2 and SCC3k2 modules in the backbone network to fuse high-level and low-level information to enrich the features. In addition, connecting different SCC3k2 modules at different feature output dimensions can further enhance the effect of multi-scale fusion.
[0097] S2.4: In the final Head layer, the Anchor-Free concept is adopted. Compared with Anchor-Based, Anchor-Free does not need to generate predefined anchor boxes, making the model more flexible and able to more effectively identify various types of defects in photovoltaic panels;
[0098] At the same time, this module adopts a decoupled approach, which can maximize the role of multiple Detect modules and output small sample and large sample defects on multiple channels respectively; in addition, the original Detect module is changed to the DTADetect module, and the corresponding dynamic task alignment (DynamicTask Align) structure design is performed on the detection head with reference to the TOOD (task-aligned single-stage target detection) idea, so as to enhance the collaborative performance of the two tasks of photovoltaic panel defect location and category prediction. The extractor can extract different refined features from different convolutional layers and interact with each other, thereby improving the performance of defect classification detection;
[0099] In this part, multiple anchor boxes are first generated, and then analyzed based on IOU (intersection over union), a key indicator for measuring the degree of overlap of prediction boxes. The larger the IOU value, the more serious the overlap of prediction boxes. Next, the NMS (non-maximum suppression) technology is used to remove redundant anchor boxes with high overlap to ensure the reliability of the final prediction results. On this basis, the corresponding defect categories are assigned to the selected labeled defect anchor boxes according to the generated category probabilities.
[0100] It should be noted that the use of drones for data collection and infrared imaging technology to obtain high-resolution infrared TIR photovoltaic photography images;
[0101] Use the opencv library to perform data enhancement by rotating the image and adjusting the brightness contrast map to obtain more samples as model input;
[0102] Since infrared images contain most of the information, such as local temperature rise caused by hot spots and local brightening of image color, the labelimg annotation tool is used to select detailed defect classification and annotation of infrared images according to the characteristics corresponding to different defects of photovoltaic panels, including broken panels, damaged batteries, shadow masks caused by trees, etc.
[0103] It should be noted that the SCConv convolution in the SCC3k2 module is used in the backbone network to efficiently capture information in deep space and channels, reduce the extraction of redundant information and improve training efficiency;
[0104] The SRU module of SCConv includes separation and reconstruction operations to reduce the redundancy of spatial information. Although the traditional C3k2 reduces the computational complexity to a certain extent through the residual connection block, the data processing part still includes redundant feature information, which increases the computational complexity and brings unnecessary noise interference;
[0105] In the SRU separation operation, Group Normalization is used to extract the scaling factor of the feature map. The larger the variance, the more effective information about the photovoltaic panel defects. By comparing the values in the variance, the spatial distribution of the effective information can be known, and it can be divided into two parts according to a specific threshold. The information above the threshold is effective information, and the information below the threshold is redundant information. In order to further integrate these two parts of information, it is necessary to use the reconstruction operation to rearrange and integrate these features by cross reconstruction to enhance the expressiveness of the features;
[0106] The formula for the separation operation is as follows:
[0107] W=T(Sigmoid(W γ(GN(X)) ));
[0108] This formula indicates that in the separation operation, after the input X is processed by Group Normalization, combined with the scaling factor W r Feature extraction is performed, and then the Sigmoid activation function is applied to make the weight value smoother. Finally, the features are divided into rich features and redundant features through T threshold gating, and the final feature map weight W is obtained.
[0109] In the SRU module of SCC3k2, the threshold T is changed from the default 0.5 to 0.4 to extract more and richer photovoltaic panel defect features, because the defects caused by categories such as debris covering in photovoltaic panels are usually small targets on photovoltaic panels and there are more of them. If the threshold is not changed, there is a greater probability that they will be identified as redundant features.
[0110] The relevant formula for threshold allocation is as follows:
[0111]
[0112] Among them, W1 is the upper layer rich feature map, and W2 is the shallow layer redundant feature map.
[0113] It should be noted that in the backbone network construction and feature extraction, the channel reconstruction unit CRU in SCConv is divided into three steps: segmentation, transformation, and fusion;
[0114] In the segmentation step, CRU first obtains the spatially refined feature map obtained in the previous step and divides its channels into two parts, including αC channels and (1-α)C channels respectively. Then, the feature map is compressed through 1×1 convolution to improve computational efficiency.
[0115] In the transformation step, the separated upper half feature map is transformed by group-wise convolution (GWC) and point-wise convolution (PWC) to extract rich high-level features; the lower half is residually concatenated with the original shallow input through a cheap 1×1 convolution to supplement the shallow feature information;
[0116] In the fusion part, the upper and lower parts of the information are weighted and fused through global average pooling and the β parameter related to the soft attention mechanism to extract the final channel refined feature map. Compared with direct connection, fusion is more efficient in extracting information;
[0117] In the Bottleneck bottleneck structure originally included in C3k2, the second layer Conv is replaced with SCConv. In order to reduce the computational burden, the present invention chooses to replace only the convolutional layer in the second half of the model, thereby achieving performance improvement while effectively controlling the computational cost. Finally, the residual connection is used to form SCBottleneck to replace Bottleneck, forming an improved SCC3k2 module.
[0118] It should be noted that the cascade group attention mechanism (CGA) is used to enhance the feature representation step. The CGA module is used in the final C2WCGA of the backbone network. By performing different segmentations of the complete features on each head in the multi-head attention, each subtask can extract more accurate and refined features, and finally connect them to obtain more effective feature outputs to provide valuable context information.
[0119] Specifically, the input feature map is first spatially segmented into multiple local windows, and then multiple sub-regions are formed. In this process, the height and width dimensions are reduced while the number of channels remains unchanged, and the batch size is increased, which can effectively improve the efficiency of parallel computing.
[0120] Next, a multi-head attention mechanism is applied in parallel on each window. First, the input feature map passes through the QKV layer to calculate the query, key, and value, and then it is divided into multiple sub-features, and each sub-task is processed separately. Each head performs feature modeling through the self-attention mechanism to capture local context information. An attention bias is added to the attention module to guide the attention mechanism to focus more finely on different areas and features, improving the performance of the model. The specific formula is as follows:
[0121]
[0122] Among them, d k is the dimension of the key vector, which is used to scale the dot product to avoid excessive values. B is a bias term, which is used to adjust the similarity between the query Q and the key K. The softmax operation converts the similarity into weights, and weights the value vector V by these weights, and finally outputs the weighted value vector V. In the code implementation, a feat feature map array is introduced to record the multi-head output results obtained by parallel calculation. The output results of each head will be cascaded to the input of the next head in turn, and multiple feature fusions and information transfers will be performed to ensure effective information exchange between different heads. The specific cascade formula is as follows:
[0123]
[0124] Among them, the actual input features in the first layer are the same as the original input features, and the actual input features f of the i-th layer are i It is necessary to superimpose the current original input x i And the output of the previous layer Get it.
[0125] Finally, the outputs of all heads will be reconnected to form the final cascaded group attention result, which will be passed as input to the next layer of the network for subsequent processing after shape transformation;
[0126] The neck layer also introduces the SCC3k2 module to replace the traditional C3k2 structure. Through more sophisticated and rich feature extraction and multi-scale information fusion, the model's perception of defects of different scales is significantly improved, especially the detection accuracy under complex backgrounds. Specifically, the SCC3k2 module uses spatial channel reconstruction convolution to reduce the calculation of redundant information, so that the model can pay more attention to the fusion of rich and detailed features. It not only enhances the expression ability of features at all levels, but also effectively solves the problems caused by scale changes, so that the network can more accurately capture defects of different sizes and categories on the surface of photovoltaic panels, including large-area circuit board failures and tiny bird droppings debris coverings. In complex scenarios with different factors such as lighting and angles, the model can also show high accuracy, further enhancing the robustness of the model.
[0127] In the final Head layer, multiple DTADetect modules are used for output. This model uses the idea of dynamic task alignment, task decomposition and dynamic convolution mechanism to improve the accuracy and efficiency of target detection. The specific contents are as follows:
[0128] First, multi-scale feature processing is performed, for each input scale x i , first extract features through the shared convolution layer ShareConv, and then concatenate these features to obtain a fused feature map;
[0129] For the concatenated feature map, firstly, adaptive average pooling is used to reduce its dimension to an average feature to obtain global context information, and then the task decomposition module is used to perform regression and classification feature extraction on the feature map respectively;
[0130] In the regression task, we first generate an offset and a mask through spatial convolution offset. These offsets will help adjust the regression features to make the regression of the target box more accurate. Then, we use the DyDCNV2 dynamic convolution module to input the regression features together with the offset and the mask to further adjust and align the regression features.
[0131] For classification tasks, a convolutional layer is used to obtain classification features. After ReLU activation, the classification probability of each region is generated through a category convolutional layer. The classification features and category probabilities are multiplied to obtain the final category prediction.
[0132] For each scale output, the regression prediction and category prediction are concatenated to merge the results to obtain the final prediction output of each anchor box. The predicted target box position is decoded and converted into actual coordinates to complete the drawing of the final bounding box.
[0133] By sharing convolutional layers at multiple scales, this module not only reduces the amount of computation, but also effectively copes with the detection requirements of targets of different scales. By using adaptive feature processing and task alignment strategies, the model can efficiently process input features during training and inference, output accurate target boxes and category probabilities, and further improve the detection head's ability to adapt to multi-scale and multi-category targets.
[0134] The training process continues to converge. Since its convergence rate is not completely monotonic, it is necessary to record not only the last model, but also the model with the smallest loss in all processes.
[0135] Use the best trained model file to classify and detect defects in photovoltaic panel images. After obtaining the results, integrate the defect information to achieve accurate defect positioning.
[0136] Anchor frames of different colors are marked on the image, and the image results are output; at the same time, the processed defect fault coordinates, categories and other information are sorted into an Excel table for storage;
[0137] Relevant photovoltaic staff can locate and analyze defective photovoltaic panels based on the visual defect classification detection diagram and Excel data table, which is convenient for subsequent photovoltaic panel repair and replacement work.
[0138] It should be noted that the specific content of S3 is:
[0139] S3.1: specify a specified file directory and use the batch of photovoltaic panel infrared images in it as input;
[0140] S3.2: Specify the configuration file and load the model trained using the improved YOLOv11 algorithm;
[0141] S3.3: Traverse each image in the folder, use the model to classify and detect photovoltaic panel defects, and generate anchor boxes in the image to locate the position and size of the defects. Each anchor box will be marked with the category name and confidence level.
[0142] In summary, the present invention provides a photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention accurately locates, marks and classifies defects on photovoltaic panels according to the internal circuit structure of photovoltaic panels and infrared imaging images; the latest YOLOv11 target detection algorithm is adopted and improved, and some C3k2 modules are replaced with SCC3k2, so that the information capture on space and channels is more accurate, and the adverse effects of redundant information extraction are reduced; C2WCGA is used to replace the original basic C2PSA to enhance the fusion performance of refined features, so that the recognition of photovoltaic panel defect classification detection is higher and the effect is better; the detection head part improves the Detect module into a DTADetect module, learns the interactive features between multiple tasks, and extracts more refined information; after the infrared picture is processed by the trained model, the picture with defect annotation and classification information and the processed table data are output, which is convenient for the photovoltaic panel staff to repair or replace the photovoltaic panels at the corresponding positions in the later stage.
[0143] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention, characterized in that: The following steps are involved: S1. Data collection, annotation and preprocessing; S2. Use the improved YOLOv11 algorithm to train the model; S3. Predict the input infrared image, accurately locate the defects and mark the categories.
2. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 1 is characterized in that: The specific content of S1 is: S1.1: Obtain infrared imaging images taken by the drone; S1.2: Analyze the internal circuit of the photovoltaic panel and deal with the one-to-one correspondence between the fault type and the manifestation; S1.3: Use annotation tools to carefully annotate defects and faults, including location information and classification groups; S1.4: Use the opencv library to perform data enhancement operations such as rotation, brightness and contrast adjustment, and scaling on the photovoltaic panel images to generate more samples for subsequent model training.
3. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 1, characterized in that: The S2 adaptively improves the YOLOv11 network, passes the image into the pre-trained network, and trains the model. The specific contents are as follows: S2.1: The detection module is divided into the backbone network, the feature pyramid neck, and the final detection head; The downsampling layer in the backbone network uses multi-layer convolution to extract multi-dimensional features from the infrared-photographed photovoltaic panel images that are passed in, and uses the SPPF module to enhance the features of important information to obtain valuable defect information. In the Neck part, a feature pyramid structure is constructed by introducing upsampling layers and multiple convolution modules, and connected with the multi-layer features of the backbone network to achieve multi-scale feature fusion; The detection head contains several Detect modules, in which anchor boxes are drawn and redundant duplicate boxes are removed through non-maximum suppression. The maximum probability category is selected in the output probability distribution, and the defect conditions of small samples and large samples are output in multiple channels. The specific structure of the backbone network is as follows: S2.2: In the backbone network, multiple Conv convolution layers, C3k2 layers and SCC3k2 layers are alternately connected to build a deep network to continuously sample and extract information. The specific connection method is shown in Figure 2, and the gray part in the figure is the improved module. The SCC3k2 module uses the spatial and channel reconstruction convolution SCConv to replace the convolution part in the traditional C3k2 module to enhance the deep feature extraction effect; The SCConv module consists of two core parts, the spatial reconstruction unit (SRU) and the channel reconstruction unit (CRU), which are sequentially composed. The SRU is used to reduce the redundancy in the spatial dimension, and then the CRU is used to reduce the redundancy in the channel dimension. SCConv can be seamlessly integrated into the existing CNN to replace the standard convolution operation. After several Conv convolutional layers, C3k2 layers and SCC3k2 layers are alternately connected, they are processed by SPPF parallel maximum pooling to further enhance the effective defect information; At the end of the backbone network, C2WCGA is used to replace the original C2PSA module. The cascade group attention based on local window parallel processing is used to focus more on learning detailed features with rich information, thus improving the diversity of attention. Since the input image has a high resolution, a local window attention mechanism is introduced before it is input into the module. It is used to split the width and height of the input into multiple sub-regions according to a certain resolution, so that the width and height are reduced, the channels remain unchanged, and the batch size is increased. After dividing multiple local windows, the cascade group attention mechanism is applied independently to each window. This method can reduce the amount of calculation, especially when processing high-resolution input, and can increase the model's attention to local features, thereby improving performance; The CGA cascade group attention part under this module contains more attention heads, which perform attention calculations on different channels and perform cascade operations, that is, the output of the previous head is spliced to the input of the next head. This not only enhances the information flow within each window, but also promotes information fusion across windows, which is conducive to the model gradually generating a global understanding of the entire image; S2.3: The core idea of feature pyramid FPN is to generate feature maps of multiple scales in different convolutional layers and enrich the representation by gradually fusing high-level and low-level features. In the neck layer, the UnSample upsampling layer refines and expands the features and concats them with the C3k2 and SCC3k2 modules in the backbone network to fuse high-level and low-level information to enrich the features. In addition, connecting different SCC3k2 modules at different feature output dimensions can further enhance the effect of multi-scale fusion. S2.4: In the final Head layer, the Anchor-Free concept is adopted. Compared with Anchor-Based, Anchor-Free does not need to generate predefined anchor boxes, making the model more flexible and able to more effectively identify various types of defects in photovoltaic panels; At the same time, this module adopts a decoupled approach, which can maximize the role of multiple Detect modules and output small sample and large sample defects on multiple channels respectively; in addition, the original Detect module is changed to the DTADetect module, and the corresponding dynamic task alignment (DynamicTask Align) structure design is performed on the detection head with reference to the TOOD (task-aligned single-stage target detection) idea, so as to enhance the collaborative performance of the two tasks of photovoltaic panel defect location and category prediction. The extractor can extract different refined features from different convolutional layers and interact with each other, thereby improving the performance of defect classification detection; In this part, multiple anchor boxes are first generated, and then analyzed based on IOU (intersection over union), a key indicator for measuring the degree of overlap of prediction boxes. The larger the IOU value, the more serious the overlap of prediction boxes. Next, the NMS (non-maximum suppression) technology is used to remove redundant anchor boxes with high overlap to ensure the reliability of the final prediction results. On this basis, the corresponding defect categories are assigned to the selected labeled defect anchor boxes according to the generated category probabilities.
4. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 1, characterized in that: The specific content of S3 is: S3.1: specify a specified file directory and use the batch of photovoltaic panel infrared images in it as input; S3.2: Specify the configuration file and load the model trained using the improved YOLOv11 algorithm; S3.3: Traverse each image in the folder, use the model to classify and detect photovoltaic panel defects, and generate anchor boxes in the image to locate the position and size of the defects. Each anchor box will be marked with the category name and confidence level.
5. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 3 is characterized in that: Use the SCConv convolution in the SCC3k2 module to efficiently capture deep spatial and channel information, reduce the extraction of redundant information and improve training efficiency; The SRU module of SCConv includes separation operations and reconstruction operations to reduce the redundancy of spatial information. Although the traditional C3k2 reduces the computational complexity to a certain extent through the residual connection block, the data processing part still includes redundant feature information, which increases the computational complexity and brings unnecessary noise interference; In the SRU separation operation, Group Normalization is used to extract the scaling factor of the feature map. The larger the variance, the more effective information about the photovoltaic panel defects. By comparing the values in the variance, the spatial distribution of the effective information can be known, and it can be divided into two parts according to a specific threshold. The information above the threshold is effective information, and the information below the threshold is redundant information. In order to further integrate these two parts of information, it is necessary to use the reconstruction operation to rearrange and integrate these features by cross reconstruction to enhance the expressiveness of the features; The formula for the separation operation is as follows: W=T(Sigmoid(W γ(GN(X)) )); This formula indicates that in the separation operation, after the input X is processed by Group Normalization, combined with the scaling factor W γ Feature extraction is performed, and then the Sigmoid activation function is applied to make the weight value smoother. Finally, the features are divided into rich features and redundant features through T threshold gating, and the final feature map weight W is obtained. In the SRU module of SCC3k2, the threshold T is changed from the default 0.5 to 0.4 to extract more and richer photovoltaic panel defect features, because the defects caused by categories such as debris covering in photovoltaic panels are usually small targets on photovoltaic panels and there are more of them. If the threshold is not changed, there is a greater probability that they will be identified as redundant features. The relevant formula for threshold allocation is as follows: Among them, W1 is the upper layer rich feature map, and W2 is the shallow layer redundant feature map.
6. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 3 is characterized in that: In backbone network construction and feature extraction, the channel reconstruction unit CRU in SCConv is divided into three steps: segmentation, transformation, and fusion; In the segmentation step, CRU first obtains the spatially refined feature map obtained in the previous step and divides its channels into two parts, including αC channels and (1-α)C channels respectively. Then, the feature map is compressed through 1×1 convolution to improve computational efficiency. In the transformation step, the separated upper half feature map is transformed by group-wise convolution (GWC) and point-wise convolution (PWC) to extract rich high-level features; The lower part uses a cheap 1×1 convolution to perform residual cascade with the original shallow input to supplement the shallow feature information; In the fusion part, the upper and lower parts of the information are weighted and fused through global average pooling and the β parameter related to the soft attention mechanism to extract the final channel refined feature map. Compared with direct connection, fusion is more efficient in extracting information; In the Bottleneck bottleneck structure originally included in C3k2, the second layer Conv is replaced with SCConv. In order to reduce the computational burden, the present invention chooses to replace only the convolutional layer in the second half of the model, thereby achieving performance improvement while effectively controlling the computational cost. Finally, SCBottleneck is formed through residual connection to replace Bottleneck, forming an improved SCC3k2 module.
7. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to claim 3 is characterized by: The cascade group attention mechanism (CGA) is used to enhance the feature representation step. The CGA module is used in the final C2WCGA of the backbone network. By performing different segmentations of the complete features on each head in the multi-head attention, each subtask can extract more accurate and refined features, and finally connect them to obtain more effective feature outputs to provide valuable context information. Specifically, the input feature map is first spatially segmented into multiple local windows, and then multiple sub-regions are formed. In this process, the height and width dimensions are reduced while the number of channels remains unchanged, and the batch size is increased, which can effectively improve the efficiency of parallel computing. Next, a multi-head attention mechanism is applied in parallel on each window. First, the input feature map passes through the QKV layer to calculate the query, key, and value, and then it is divided into multiple sub-features, and each sub-task is processed separately. Each head performs feature modeling through the self-attention mechanism to capture local context information. An attention bias is added to the attention module to guide the attention mechanism to focus more finely on different areas and features, improving the performance of the model. The specific formula is as follows: Among them, d k is the dimension of the key vector, which is used to scale the dot product to avoid excessive values. B is a bias term, which is used to adjust the similarity between the query Q and the key K. The softmax operation converts the similarity into weights, and weights the value vector V by these weights, and finally outputs the weighted value vector V. In the code implementation, a feat feature map array is introduced to record the multi-head output results obtained by parallel calculation. The output results of each head will be cascaded to the input of the next head in turn, and multiple feature fusions and information transfers will be performed to ensure effective information exchange between different heads. The specific cascade formula is as follows: Among them, the actual input features in the first layer are the same as the original input features, and the actual input features f of the i-th layer are i It is necessary to superimpose the current original input x i And the output of the previous layer Get it.
Citation Information
Patent Citations
Green orange detection method based on self-supervised comparative learning and computer device
CN119169472A
Cited By
Chip packaging defect detection method applied to edge device based on YOLOv11m
CN120219388A
A chip packaging defect detection method based on YOLOv11m for edge devices
CN120219388B
Photovoltaic cell defect detection method based on improved YOLOv11, electronic equipment and computer readable storage medium
CN120411737A
Graphite ore grade detection method based on improved YOLO11 model
CN120783075A
A method for detecting the grade of graphite ore based on an improved YOLO11 model
CN120783075B