Photovoltaic panel surface defect detection method based on pruning multi-scale feature fusion network

By introducing FFPEMA, DCCARAFE and PAFIFN modules into the YOLOv8 network, the feature fusion and pruning strategies are optimized, and the problem of insufficient feature extraction in the photovoltaic panel surface defect detection is solved, achieving high-precision and efficient defect detection.

CN120278992APending Publication Date: 2025-07-08ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510441156.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

Smart Images

  • Figure CN120278992A_ABST
    Figure CN120278992A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic panel surface defect detection method based on a pruning multi-scale feature fusion network, and belongs to the technical field of photovoltaic panel defect detection, and the method comprises the following steps: S1, data set processing; s2, constructing a network; s3, model training; and S4, defect detection. According to the method, a feature fusion pruning efficient multi-scale attention (FFPEMA) module is introduced into a backbone network to improve the feature representation capability, and network structure global pruning is combined with a feature fusion technology to enhance feature representation so as to improve the model detection precision; a dilated convolution content awareness (DCCARAFE) module is designed, a dilated convolution strategy is adopted, the receptive field is increased by adjusting the voidage, and the flexibility of model feature extraction is improved; a pruning efficient self-adaptive feature integration fusion strategy (PAFIFN) is provided, and the feature extraction and learning ability of the model for defects of different sizes and scales is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic panel defect detection, and in particular to a method for detecting surface defects of photovoltaic panels based on a pruning multi-scale feature fusion network. Background Art

[0002] Under the background of the rapid growth of the global population and the high-speed development of the economy, the overuse of traditional fossil energy has caused serious environmental and climate problems. This situation has promoted the development of renewable energy as the core strategy for the transformation of the global energy system. As a representative of clean energy, the zero-emission characteristic of solar energy makes it play an important role in the field of renewable energy. As the core device for converting light energy, the performance of a solar photovoltaic panel (PV) directly affects the overall efficiency of a photovoltaic power generation system. It is worth noting that surface damage to PV panel components not only reduces the energy conversion rate, but the hot spot effect among them may also cause safety hazards such as fires. With the continuous expansion of the scale of photovoltaic power stations, ensuring the defect-free operation of equipment has become a key factor in ensuring the stability of the power generation system, improving the energy conversion efficiency, and maintaining safe production.

[0003] Currently, the technology for detecting surface defects of photovoltaic (PV) panel components has become a hot research field. In the initial stage of technology development, the industry generally adopted the method of manual visual inspection, and identified abnormalities such as hot spots, occlusions, or physical damage by observing the appearance of components. This manual inspection mode has limitations such as low detection efficiency, strong subjectivity, and reliance on professionals. Especially during long-term operations, it is prone to misjudgments and missed detections due to visual fatigue, and it is difficult to meet the real-time monitoring and batch detection requirements of large-scale photovoltaic power stations. The existing detection technology system mainly includes two major branches: traditional image processing and deep learning. Traditional methods use a sliding window mechanism to generate candidate regions, and combine HOG feature extraction and a support vector machine (Support Vector Machine, SVM) to form a detection framework. Typical algorithms include the Viola-Jones detector, AdaBoost ensemble learning, etc. There are two major technical bottlenecks in such methods: on the one hand, a large number of false detections are easily generated in complex backgrounds, and on the other hand, the adaptability to target scale changes is insufficient.

[0004] Deep learning technology has achieved a breakthrough improvement in detection efficiency by constructing multi-level feature representations. Its technical route can be divided into two-stage and single-stage architectures: The two-stage algorithms represented by the R-CNN series achieve high-precision detection through two steps: candidate box generation and refined classification regression; while single-stage algorithms such as YOLO and SSD adopt an end-to-end architecture, directly synchronously completing localization and classification on the feature map, with significant speed advantages. With the continuous iterative upgrade of algorithms such as the YOLO series, while ensuring real-time performance, the detection accuracy of single-stage detectors has gradually approached that of two-stage models. For the detection requirements of a large number of components in distributed photovoltaic power stations, intelligent detection algorithms with both high precision and high efficiency have become the key research direction in this field.

[0005] In the field of image detection, the YOLO algorithm has become the mainstream technical framework due to its high inference speed and high detection accuracy. To continuously optimize the performance advantages of this algorithm, researchers have innovatively improved various solutions. Regarding the problem of feature fusion efficiency, some researchers have proposed combining lightweight structures with an efficient multi-scale attention mechanism (EMA) to enhance robustness while controlling the number of model parameters. However, in the process of two-channel feature interaction, there are still technical bottlenecks in the insufficient fusion of cross-scale attention maps. Another research team's developed BiMAF module optimizes multi-level feature correlation through a global-local two-dimensional feature integration mechanism. Although multi-level feature integration has been achieved, it is limited by the fixed receptive field range. In addition, a composite network architecture formed by integrating GSConv lightweight convolution, bidirectional feature pyramid (BiFPN), and depthwise separable convolution (DW-Conv) has made breakthroughs in improving detection accuracy and training efficiency. However, there is still a risk of feature loss when dealing with edge details and small targets, resulting in insufficient integrity of deep network information integration. These technical explorations have reflected that there is still room for optimization in feature extraction and fusion in complex scenarios while improving model performance.

[0006] As an advanced version of the current mainstream detection framework, YOLOv8 demonstrates significant advantages in industrial detection scenarios by integrating a functional architecture that combines classification, detection, and segmentation. While maintaining high classification accuracy, this algorithm exhibits excellent performance metrics (mAP) in detection and segmentation, and its inference speed meets the requirements of real-time monitoring. However, the algorithm has limitations in representing weak feature targets (such as fine cracks), especially in terms of enhancing edge features and adapting to multi-scale targets, leaving room for optimization. The fixed-mode upsampling mechanism in the existing architecture restricts the dynamic extraction of defect features, leading to information attenuation of small-scale abnormal features in the deep network and affecting detection sensitivity. To address these technical bottlenecks, this study proposes a multi-scale feature fusion architecture based on pruning optimization, which is specifically optimized for the detection of surface defects in photovoltaic modules. Summary of the Invention

[0007] The technical problem to be solved by the present invention is as follows: how to solve the problems existing in the existing YOLOv8 detection model, such as the difficulty in extracting features like cracks and the inability to effectively capture targets with different scale features, and further achieve accurate detection of surface defects on photovoltaic panels. A method for detecting surface defects on photovoltaic panels based on a pruned multi-scale feature fusion network is provided.

[0008] As Figure 13 shown, the present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:

[0009] S1: Dataset Processing

[0010] Select a basic dataset and perform partitioning on the dataset to obtain a training set and a test set.

[0011] S2: Network Construction

[0012] Select the YOLOv8 network as the baseline model, and design the FFPEMA module, DCCARAFE module, and PAFIFN module in the baseline model to obtain the pruned multi-scale feature fusion network PMFFN.

[0013] S3: Model Training

[0014] Use the training set to train the pruned multi-scale feature fusion network PMFFN to obtain a trained model, that is, a surface defect detection model for photovoltaic panels.

[0015] S4: Defect Detection

[0016] Input the samples in the test set into the surface defect detection model for photovoltaic panels to output the surface defect results of the photovoltaic panels.

[0017] Furthermore, in the step S2, the specific processes of introducing the FFPEMA module, the DCCARAFE module, and the PAFIFN module are as follows:

[0018] S21: Embed the PEMA module in the C2f module of the backbone network of the baseline model, and construct a multi-level feature fusion architecture at the feature output end to form the FFPEMA module;

[0019] S22: Introduce a dilated convolutional layer in the upsampling kernel prediction and information encoding links of the neck network of the baseline model to form the DCCARAFE module;

[0020] S23: Introduce the PAFIFN module at the output end of the last FFPEMA module in the backbone network of the baseline model.

[0021] Furthermore, in the FFPEMA module of the step S21, the Bottleneck layer is a depthwise separable convolution reconstructed Bottleneck layer. The cross-layer feature splicing mechanism is adopted to perform channel dimension fusion on the global context features output by the PEMA module and the local detail features extracted by the depthwise separable convolution reconstructed Bottleneck layer, and feature interaction enhancement is achieved through gated convolution. The improved ReLU activation function is used in the convolutional layer to eliminate the distribution differences between multi-source features; among them, the depthwise separable convolution reconstructed Bottleneck layer includes two convolutional layers. After the input features pass through the first convolutional layer, they are divided into two branches. After one branch is processed by the second convolutional layer, it is superimposed with the features on the other branch and then output. The convolutional layer includes a two-dimensional convolutional block, a two-dimensional batch normalization unit, and a SiLU activation function layer connected in sequence.

[0022] Furthermore, the PEMA module is obtained by pruning the channels and weight parameters of the EMA module through an unstructured global pruning algorithm. In the EMA module, the input feature map first undergoes a series of convolutional and pooling operations to extract local and global context information, and then an attention map is generated through weighted aggregation. After the attention map is combined with the original input feature map, it is output. The weighted aggregation attention calculation formula is as follows:

[0023]

[0024] Among them, A is the generated attention map, which is used to allocate weights to the input feature map; F i represents the context feature of the i-th channel extracted through convolutional and pooling operations, and the context feature includes local and global features; w iis the learnable weight parameter for the i-th channel, which is used to dynamically adjust the importance of features in different channels; σ is the Sigmoid activation function, which is used to map the weighted sum to the range [0,1] to generate the normalized attention weights; C is the number of channels of the input feature map.

[0025] Furthermore, in the step S21, the calculation formula of the FFPEMA module is as follows:

[0026] FFPEMA[t] = βEMA[t] + (1 - β)y[t]

[0027] where FFPEMA[t] represents the feature fusion output result at the current time step, EMA[t] represents the feature processed by exponential moving average, β represents the fusion weight, and y[t] represents another set of features for fusion.

[0028] Furthermore, in the step S22, the DCCARAFE module includes an upsampling kernel prediction module and a feature reconstruction module. The dilated convolution operation is introduced in the upsampling prediction and information encoding links of the model, that is, a dilated convolution layer is used. Among them, the dilated convolution layer expands the receptive field without increasing the number of parameters by adjusting the dilation rate r in the dilated convolution.

[0029] Furthermore, the calculation formula of the dilated convolution is as follows:

[0030]

[0031] where x is the input feature map, y is the output feature map, w is the convolution kernel, r is the dilation rate, which is used to control the interval between elements in the convolution kernel, k is the size of the convolution kernel, i, j are the position coordinates of the output feature map, representing the value of the i-th row and j-th column in the output feature map y, which is calculated by the weighted sum of the corresponding position of the input feature map; m, n are the row and column indices in the convolution kernel, which are used to traverse each element of the convolution kernel, and the value range is 0 ≤ m, n ≤ k - 1.

[0032] Furthermore, in the step S23, the PAFIFN module dynamically adjusts the weights of features at different scales through an adaptive feature integration strategy, improves the model's detection ability for defects at different scales, and then performs pruning and efficient multi-scale feature fusion operation processing on the feature results generated by the adaptive feature integration to enhance the dynamic nature of the model's feature representation.

[0033] Furthermore, in the PAFIFN module, the weight adjustment formula for adaptive feature integration is as follows:

[0034]

[0035] where zi is the weight score of the i-th feature map, is the final weight of the i-th feature map, and N is the number of feature maps; in pruning-efficient multi-scale feature fusion, the loss function L of the pruning strategy prune is as follows:

[0036]

[0037] where, is the i-th weight, λ is the pruning intensity parameter, and M is the total number of weights.

[0038] Furthermore, in pruning-efficient multi-scale feature fusion, feature optimization is achieved through the following steps:

[0039] S231: Multi-scale feature integration

[0040] Perform channel dimension concatenation on the multi-scale feature maps {F1, F2, …, F n} output by adaptive feature integration to form an integrated feature map F intergrated = Concat{α1F1, α2F2, …, α N F N}, where α i represents the dynamic weight, which is a weight generated based on the input feature content;

[0041] S232: Structured global pruning

[0042] Adopt a structured pruning strategy to sparsify the channel weights of the integrated feature map. The structured global pruning formula is as follows:

[0043]

[0044] where, A c,h,w is the activation value of the c-th channel at the spatial position (h, w); H and W are the height and width of the feature map respectively;

[0045] S233: Efficient feature fusion

[0046] Perform cross-scale fusion on the pruned feature map through gated convolution. The formula is as follows:

[0047] F fused = GatConv(F puned , F skip )

[0048] where, F fused is the fused feature map; F puned is the pruned feature map; F skip retains the fine-grained information of the original feature, and it is related to F punedDynamic weighted fusion through gated convolution; GatConv represents gated convolution.

[0049] The present invention has the following advantages compared with the prior art: In the photovoltaic panel surface defect detection method based on the pruning multi-scale feature fusion network, an FFPEMA module is introduced in the backbone network to improve the feature representation ability, and global pruning of the network structure is combined with the feature fusion technology to enhance the feature expression and improve the model detection accuracy; A dilated convolution content-aware (DCCARAFE) module is designed, which adopts the dilated convolution strategy to increase the receptive field by adjusting the dilation rate and improve the flexibility of the model feature extraction; A pruning efficient adaptive feature integration and fusion strategy (PAFIFN) is proposed to improve the model's feature extraction and learning ability for defects of different scales. Brief Description of the Drawings

[0050] Figure 1 It is a schematic structural diagram of the PEMA module in the embodiment of the present invention;

[0051] Figure 2 It is the overall framework diagram of the PMFFN network in the embodiment of the present invention;

[0052] Figure 3 It is a schematic structural diagram of the FFPMEA module in the embodiment of the present invention;

[0053] Figure 4 It is a schematic structural diagram of the DCCARAFE module in the embodiment of the present invention;

[0054] Figure 5 It is a schematic structural diagram of the PAFIFN module in the embodiment of the present invention;

[0055] Figure 6 It is the overall implementation process framework diagram of the method of the present invention in the embodiment of the present invention;

[0056] Figure 7 It is an example diagram of typical defect types of the PVEL-AD dataset in the embodiment of the present invention. Among them, (a) is crack, (b) is fingerprint, (c) is black core, (d) is short circuit, (e) is thick line, and (f) is horizontal displacement;

[0057] Figure 8 It is the defect distribution diagram in the embodiment of the present invention;

[0058] Figure 9It is a comparison chart of the confusion matrix results of the YOLOv8 baseline model and the PMFFN model in the embodiments of the present invention. Among them, (a) is the YOLOv8 baseline model, and (b) is the PMFFN model;

[0059] Figure 10 It is a comparison chart of the training loss curves of the PMFFN model and the YOLOv8 baseline model in the embodiments of the present invention. Among them, (a) is Box_loss, (b) is cls_loss, and (c) is mAP;

[0060] Figure 11 It is an example of the missed detection and false detection results of the YOLOv8 baseline model and the PMFFN model in the embodiments of the present invention. Among them, (a) is the missed detection result of YOLOv8 relative to PMFFN on the PVEL-AD dataset, (b) is the false detection result of YOLOv8 relative to PMFFN on the PVEL-AD dataset, (c) is the missed detection result of YOLOv8 relative to PMFFN on dataset A, (d) is the false detection result of YOLOv8 relative to PMFFN on dataset A, (e) is the missed detection result of YOLOv8 relative to PMFFN on dataset B, and (f) is the false detection result of YOLOv8 relative to PMFFN on dataset B;

[0061] Figure 12 It is the visualization result of the attention map of the YOLOv8 baseline model and the PMFFN model in the PVEL-AD dataset in the embodiments of the present invention;

[0062] Figure 13 It is the overall flowchart of the photovoltaic panel surface defect detection method based on the pruning multi-scale feature fusion network of the present invention. Detailed implementation mode

[0063] The embodiments of the present invention will be described in detail below. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation modes and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0064] Embodiment 1

[0065] This embodiment describes the design process of the photovoltaic panel surface defect detection method based on the pruning multi-scale feature fusion network, which specifically includes the following content:

[0066] 1 Basic theory

[0067] 1.1 YOLOv8

[0068] YOLOv8 is an important updated version of the YOLO series for defect detection. Due to its advantages such as providing models of various sizes and Anchor-Free detection heads, it significantly improves the detection accuracy, especially for small target objects. The present invention selects YOLOv8 as the baseline model. The network structure of YOLOv8 mainly consists of three parts: Backbone, Neck, and Head. The Backbone extracts image features from the input image and uses the C2f module as the basic building block. Compared with the C3 module of YOLOv5, the C2f module has fewer parameters and better feature extraction capabilities. The Neck part is used to fuse the features output from the Backbone to improve the performance of the model. The Head part is responsible for the final object detection and classification tasks.

[0069] The C2f module of YOLOv8 is designed by referring to the ideas of the C3 module of YOLOv5 and the Efficient Layer Aggregation Network (ELAN) of YOLOv7. This enables YOLOv8 to maintain model lightweight while achieving increased feature fusion. In YOLOv8, the C2f module first processes the input feature map through a convolutional operation to obtain an intermediate feature map. This intermediate feature map is then split into two parts: one part directly enters the Concat layer; the other part undergoes further convolutional, normalization, and activation operations through multiple Bottleneck layers. The C2f module also has a depth adjustment function, which can adjust the depth of the model according to different computational requirements to achieve effective management of computational resources. The loss function of YOLOv8 consists of the objectness loss, the bounding box regression loss, and the class probability. The optimization of the loss function improves the performance of the model in object detection tasks by combining multiple loss functions and optimization algorithms.

[0070] 1.2 PEMA Module

[0071] The Efficient Multi-Scale Attention (EMA) module is a novel attention mechanism optimized based on the traditional Channel Attention (CA) module. EMA is an advanced feature enhancement technology aimed at improving the model's ability to capture key feature information by adaptively adjusting the weight distribution of feature maps. This mechanism can dynamically highlight important features and suppress noise by calculating the exponential moving average of feature maps, thereby enhancing the model's object detection performance in complex scenarios. In the EMA module, the input feature map first undergoes a series of convolution and pooling operations to extract local and global context information, and then an attention map is generated through weighted aggregation. The formula for weighted aggregation attention is as follows:

[0072]

[0073] where A is the generated attention map used to assign weights to the input feature map; F i represents the context feature of the i-th channel (including local and global features) extracted through convolution and pooling operations; w i is the learnable weight parameter of the i-th channel, used to dynamically adjust the importance of different channel features; σ is the Sigmoid activation function that maps the weighted sum to the range [0, 1] to generate normalized attention weights; C is the number of channels of the input feature map.

[0074] The combination of this attention map and the original feature map realizes the fine-tuning of image features. The expression of EMA is as follows:

[0075] EMA[t] = a * x[t] + (1 - α) * EMA[t - 1]

[0076] where the time step t corresponds to the original data x[t], the smoothing factor α (ranging from 0 to 1) determines the weights of the current and historical data, and EMA[t - 1] represents the EMA value at the previous time point.

[0077] The baseline model YOLOv8 is limited by computing resources and inference time, and the model faces the problem of noise interference during transportation. In this embodiment, global pruning operations are performed on EMA. Pruning is divided into two categories: structured pruning and unstructured pruning. Compared with structured pruning that removes weights according to specific patterns or methods, unstructured pruning does not rely on specific patterns or methods for weight removal, and randomly deletes weights in the weight matrix of the neural network. Unstructured global pruning selectively removes weights according to the L1 norm (Manhattan distance) of the weights, rather than removing weights for specific rows, columns, or blocks, which can more flexibly reduce feature loss and improve the robustness and generalization of the model. Therefore, in this embodiment, an unstructured global pruning algorithm is used to prune the channel and weight parameters of the attention mechanism in the model, reducing the computational complexity and model size while ensuring accuracy, thereby obtaining a Pruned Efficient Multi-Scale Attention (PEMA) module with high robustness and strong generalization. The PEMA module is as Figure 1 shown.

[0078] 2 Method of the present invention

[0079] The present invention proposes a method for detecting surface defects of photovoltaic panels based on a pruned multi-scale feature fusion network PMFFN. The overall framework of PMFFN is as Figure 2 shown. Although YOLOv8 has shown good performance in the field of object detection, there are still defects in the task of detecting surface defects of photovoltaic panels. To enhance the compatibility of the model with transplanted devices, the present invention performs global structural pruning on EMA and combines feature fusion technology to form a Feature Fusion Pruned Efficient Multi-Scale Attention (FFPEMA) module. To enhance the multi-scale feature fusion characteristics of the model to enhance feature representation, the present invention designs a Dilated Convolution Content-Aware (DCCARAFE) module, which improves the flexibility of the model's feature extraction by adjusting the dilation rate. To improve the model's feature extraction and learning ability for defects of different scales, a Pruned Adaptive Feature Integration and Fusion Strategy (PAFIFN) is proposed, which extracts feature vectors from the defect information of the feature map input last time, performs feature learning, and then conducts efficient multi-scale feature fusion processing.

[0080] 2.1 FFPEMA module

[0081] The technology for detecting surface defects of photovoltaic modules plays a crucial role in improving the efficiency of solar power generation. The existing detection technologies have the following technical bottlenecks: 1) The uneven scale distribution of defect targets makes it difficult to extract features of small-sized defects; 2) The redundant number of parameters in traditional detection models leads to excessive consumption of computing resources and it is difficult to meet the real-time requirements for mobile deployment; 3) The imperfect multi-scale feature fusion mechanism affects the detection accuracy. In view of the above defects, the present invention proposes a lightweight detection architecture based on an improved FFPEMA module, by embedding the PEMA module in the C2f module of the YOLOv8 backbone network and constructing a multi-level feature fusion architecture at the feature output end, to achieve the collaborative optimization of detection accuracy and inference speed.

[0082] The FFPEMA module of the present invention achieves a performance breakthrough through three innovations:

[0083] First, integrate the PEMA attention sub-module at the front end of the Bottleneck structure of the original C2f module, adopt an improved dual-channel global average pooling and dynamic weight allocation mechanism to achieve cross-scale coupling of macroscopic morphological features and microscopic texture features. The PEMA mechanism dynamically adjusts the weight distribution of the feature map by calculating the exponential moving average of the feature map, thereby highlighting important features and suppressing noise.

[0084] The calculation formula of the FFPEMA module is as follows:

[0085] FFPEMA[t] = βEMA[t] + (1 - β)y[t]

[0086] Among them, FFPEMA[t] represents the feature fusion output result (target feature) at the current time step, EMA[t] represents the feature (main feature) processed by exponential moving average, β represents the fusion weight, and y[t] represents another set of features for fusion.

[0087] Second, innovatively design an activation function group with an adaptive pruning function, and achieve dynamic optimization of feature channels through a differentiable channel compression algorithm, reducing the computing load by 23.7% while improving the recall rate of small target detection;

[0088] Finally, construct a spatio-temporal joint optimization mechanism, and use the exponential weighted moving average algorithm to enhance the temporal continuity of the feature map, improving the model's ability to capture dynamic defects on the surface of the photovoltaic panel by 1.4%.

[0089] The feature fusion architecture of the present invention generates synergistic effects through the following technical means: 1) Reconstruct the Bottleneck layer using depthwise separable convolution; 2) Design a cross-layer feature splicing mechanism to fuse the global context features output by the PEMA module and the local detail features extracted by the Bottleneck layer in the channel dimension, and strengthen feature interaction through gated convolution; 3) Introduce a dynamic feature normalization module (BtachNorm2d) and use an improved ReLU activation function to eliminate the distribution differences between multi-source features. Verified by experiments, the mAP@0.5 on the PVEL-AD dataset reaches 91.4%, which is significantly better than mainstream detection models and meets the deployment requirements of industrial-grade edge computing devices at the same time. The structure diagram of the FFPEMA module is as Figure 3 shown.

[0090] 2.2 DCCARAFE Module

[0091] Feature upsampling methods have been widely used in the field of image processing and are an indispensable part of convolutional neural networks. These include interpolation upsampling, transposed convolution upsampling, and nearest upsampling methods. The benchmark model of the present invention uses bilinear interpolation upsampling method. For each new pixel point in the feature map generated by the previous level, the four pixel points in the original low-resolution feature map that are closest to it are matched. Then, the values of these four pixel points are averaged and weighted, and finally the upsampling operation is completed. However, the upsampling method of the benchmark model has limited learning ability. Since it cannot adaptively make changes well in the face of differences in image scales and contents, problems such as feature information loss and low flexibility exist in feature extraction. Therefore, if we want to expand the receptive field and improve the feature extraction ability, we will face the problem of an increase in model parameter calculation. To address the above problems, the present invention proposes the DCCARAFE module. The network structure of the DCCARAFE module is as Figure 4 shown.

[0092] The DCCARAFE module consists of two functional modules: upsampling kernel prediction and feature reconstruction. Atrous convolution is used to improve the flexibility of upsampling kernel prediction and feature reconstruction. First, atrous convolution operations are introduced in the upsampling prediction and information encoding stages of the model. The dilation parameter of atrous convolution is flexibly adjusted according to actual needs. Without increasing the computational amount of model parameters, the receptive field of the model is effectively expanded, improving the upsampling performance of the model. Second, DCCARAFE automatically selects an appropriate upsampling kernel according to the output feature structure of the previous stage, solving the problem of feature information loss in feature extraction caused by the original upsampling operation. DCCARAFE can reconstruct the feature map more accurately, flexibly receive the information content of the input image, and output a more accurate upsampling result. And it improves the flexibility and accuracy of model feature extraction. Atrous convolution expands the receptive field without increasing the number of parameters by introducing a dilation rate in the convolution kernel. The calculation formula of atrous convolution is as follows:

[0093]

[0094] where x is the input feature map, y is the output feature map, w is the convolution kernel, r is the dilation rate, which is used to control the interval between elements in the convolution kernel, k is the size of the convolution kernel, i and j are the position coordinates of the output feature map, representing the value at the i-th row and j-th column in the output feature map y, which is calculated by the weighted sum of the corresponding positions in the input feature map; m and n are the row and column indices in the convolution kernel, used to traverse each element of the convolution kernel, and the value range is 0 ≤ m, n ≤ k - 1.

[0095] By adjusting the dilation rate r, the DCCARAFE module can flexibly expand the receptive field, thereby improving the feature extraction ability without increasing the model parameters.

[0096] Adjust the dilation rate r according to the actual situation to expand the receptive field without increasing the model parameter calculation, so as to improve the upsampling performance of the model. Compared with the baseline model, DCCARAFE can adaptively adjust and optimize the upsampling kernel for different position content information of the image. Therefore, it is superior to the bilinear interpolation prediction of the baseline model in performance.

[0097] In the detection of photovoltaic panel surface defects, DCCARAFE has brought a significant improvement in the detection performance of the model. To sum up, the use of DCCARAFE in the detection of photovoltaic panel surface defects effectively improves the detection performance of the model. It enhances the flexibility of model feature extraction without increasing the model parameter calculation. This provides strong support for the real-time high-precision detection of photovoltaic panels.

[0098] 2.3 Pruning Adaptive Feature Integration and Fusion Strategy (PAFIFN)

[0099] In the task of detecting surface defects of photovoltaic panels, there are problems such as diverse defect scales and different sizes and shapes of different defect types. As the model extracts features layer by layer, there will be situations where feature extraction information is lost or edge feature information cannot be collected. The SPPF module adopted by the baseline model simply stitches together features of different scales. It does not make adaptive adjustments based on the mutual relationship and dynamic changes between features, and also lacks content-aware information fusion. In the face of these problems, the present invention proposes a pruning adaptive feature integration and fusion strategy (PAFIFN), as Figure 5 shown.

[0100] PAFIFN performs adaptive feature integration processing on the feature maps output by the last FFPEMA module in the backbone, providing feature vectors for each defect information. These feature vectors contain various aspects of information such as the color, texture, and shape matched by the defect, promoting feature recognition in the subsequent operations of the model. PAFIFN learns the distribution of each defect feature information in the upper-level feature map. When facing defects of different scales and types, it will dynamically adjust the weights of adaptive feature integration.

[0101] The PAFIFN module adjusts the weights of features of different scales dynamically through the adaptive feature integration strategy to improve the model's detection ability for defects of different scales. The weight adjustment formula for adaptive feature integration is as follows:

[0102]

[0103] where, z i is the weight score of the i-th feature map, w i1 is the final weight of the i-th feature map, and N is the number of feature maps.

[0104] In this way, the PAFIFN module can dynamically adjust the weights of feature maps at different levels according to the scale and type of defects. For example, when facing small-scale crack defects, the module will increase the weight of the shallow-level feature map to capture more detailed information; when facing large-area black spot defects, the module will increase the weight of the deep-level feature map to utilize richer semantic information.

[0105] When facing small-scale crack defects, the module will increase the fusion weight of the shallow-level feature map. Because more detailed texture information can be obtained in the shallow-level feature map to determine the position and shape of small-scale defect cracks. When facing large-area black spot defects, the module will increase the fusion weight of the deep-level feature map. Because there is richer semantic information in the deep-level feature map to assist in judging the range and position of black spot defects. The dynamic adjustment mechanism of PAFIFN promotes the model to perform adaptive multi-scale feature integration operations when facing diverse defect scales and different shapes and sizes, improving the model's defect detection performance.

[0106] The PAFIFN module further performs pruning for efficient multi-scale feature fusion operation on the feature results generated by adaptive feature integration, enhancing the dynamics of the model's feature representation. In the pruning for efficient multi-scale feature fusion, feature optimization is achieved through the following steps:

[0107] 1. Multi-scale feature integration: Concatenate the multi-scale feature maps {F1, F2, …, F n} in the channel dimension to form an integrated feature map F intergrated = Concat{α1F1, α2F2, …, α N F N}, where α i represents the dynamic weight, which is a weight generated based on the input feature content rather than a fixed parameter;

[0108] 2. Structured global pruning: Adopt a structured pruning strategy to sparsify the channel weights of the integrated feature map. The formula for structured global pruning is as follows:

[0109]

[0110] where A c,h,w is the activation value of the c-th channel at the spatial position (h, w); H and W respectively represent the height and width of the feature map.

[0111] 3. Efficient feature fusion: Perform cross-scale fusion on the pruned feature map through gated convolution. The formula is as follows:

[0112] F fused = GatConv(F puned , F skip )

[0113] where F fused is the fused feature map; F puned is the pruned feature map; F skip retains the fine-grained information of the original feature, which is dynamically weighted and fused with F puned through gated convolution to solve the semantic gap problem between scales; GatConv represents gated convolution, which filters out effective features through a learnable gating coefficient G.

[0114] This module adopts an advanced global average pooling method to extract the defect feature information of photovoltaic panels from an overall perspective and grasp the global feature information. At the same time, with the help of dot product operations, the associations between defect feature information are deeply explored. The Sigmode activation function in the operation is optimized to the ReLU activation function to overcome traditional limitations, enabling the model to capture more defect details. A carefully designed pruning strategy removes redundant parameters to improve the model's detection efficiency. The pruning strategy reduces the computational amount of the model by removing redundant weight parameters while maintaining the detection accuracy of the model. The loss function of the pruning strategy is as follows:

[0115]

[0116] where w i2 is the i-th weight, λ is the pruning intensity parameter, and M is the total number of weights.

[0117] Through the pruning strategy, the PAFIFN module can significantly reduce the computational complexity of the model while ensuring the detection accuracy, thereby improving the inference speed of the model.

[0118] Finally, the efficient integration of the global feature information and local feature information of photovoltaic panel defects is achieved. When the model faces defects of different scales, the feature extraction is more accurate and the learning ability is more powerful, enabling better implementation of the photovoltaic panel surface defect detection task.

[0119] 2.4 Framework of the method of the present invention

[0120] The framework of the proposed method is as Figure 6 shown, and the main process is as follows:

[0121] Step 1: Use a photovoltaic cell inspection drone to collect a photovoltaic panel defect detection dataset at different shooting angles.

[0122] Step 2: Randomly divide the photovoltaic panel defect detection dataset into training samples and test samples.

[0123] Step 3: Use the training samples to train the PMFFN network, and use TS-WIoU (Temperature scaling-WIoU) for temperature classification optimization loss calculation.

[0124] Step 4: Use the test samples to verify the PMFFN model obtained after training.

[0125] Example 2

[0126] This embodiment verifies the method in Example 1 by experimental means, specifically including the following content:

[0127] 3 Experimental results

[0128] 3.1 Dataset

[0129] As Figure 7 shown, the PVEL-AD dataset jointly released by Hebei University of Technology and Beihang University is used as the basic dataset. As Figure 7 shown, the PVEL-AD dataset has 6 types of defects, including crack, finger, black_core, short_circuit, thick_line, and horizontal_displacement. It has a total of 4415 images, and the distribution of each defect is as Figure 8 shown.

[0130] To verify the accuracy and generalization of the proposed model in image detection, especially its superiority in small object detection, image datasets A and B are used for experiments respectively. Dataset A is a photovoltaic panel defect detection dataset, which contains 4415 images and 3 types of defects, including scratch, broken grid, and dirt. Dataset B is a photovoltaic panel small object detection dataset, which contains 4007 images of a single type of small object photovoltaic panel.

[0131] 3.2 Experimental Settings

[0132] In the present invention, the proposed PMFFN network is built using the Pytorch deep learning framework. The PVEL-AD dataset is divided into a training set and a test set, and the ratio of the training set to the test set is 9:1. Specifically, the training set has 3975 images, and the test set contains 440 images. Dataset A is divided in a ratio of 4:1, with 1920 images in the training set and 480 images in the test set. Dataset B is divided in a ratio of 9:1, with 3584 images in the training set and 423 images in the test set.

[0133] In this experiment, YOLOv8 is used as the benchmark model. The number of training rounds is 100, the batch size is set to 32, the optimizer is SGD, the initial learning rate is set to 0.01, 10 worker threads are set, the IoU threshold is 0.5, and the input size is 640×640. This experiment is conducted on a cloud server platform. The GPU used in the experiment platform is NVIDIA Tesla V100-16GB, the CPU is CPU 10x Xeon Platinum 8160T, the memory is 16G, the development environment is Python3.8, CUDA11.3, Pytorch1.10, and the operating system is Windows11.

[0134] 3.3 Experimental Metrics

[0135] The present invention uses standard image detection evaluation metrics, including mAP (Mean Average Precision), Recall, and Precision, as shown in the following formulas:

[0136]

[0137] Among them, TP (True Positive) represents the number of target detection boxes correctly identified by the model; FP (False Positive) refers to the number of non-target detection boxes misjudged as targets by the model; FN (False Negative) refers to the number of actual targets not detected by the model. mAP (Mean Average Precision) is obtained by calculating the average precision of all classes and then taking its average value, which measures the average performance of the model for each class in the dataset. Precision (P) is the proportion of true positive samples among all samples judged as positive by the model, which reflects the false detection rate of the model. Recall (R) is the proportion of samples correctly identified by the model among all samples that are actually positive, which measures the missed detection rate of the model. All detections refer to all positive class samples detected by the model, and all ground truths refer to all positive class samples that actually exist.

[0138] 3.4 Comparison Results with the Baseline Model

[0139] In this embodiment, the proposed PMFFN model is compared with the baseline model YOLOv8 model on the PVEL-AD dataset to obtain the confusion matrix, precision, recall, and visual comparison results. The confusion matrix visually shows the number of true positives, false positives, true negatives, and false negatives to evaluate the performance of the classification model. As Figure 9 shown, the true positives of the PMFFN model proposed by the present invention in the categories of crack, thick_line, black_core, and short_circuit are all higher than those of the YOLOv8 baseline model. This indicates that the PMFFN model has better detection performance in the detection of photovoltaic panel surface defects.

[0140] 3.4.1 Accuracy, Recall, and Mean Average Precision

[0141] Precision, Recall, and mean Average Precision (mAP) are evaluation metrics used to measure the detection performance of a model. mAP is obtained by calculating the average precision of the model for each class and then averaging these average precisions. It reflects the overall performance of the model across all classes in the dataset. Precision represents the ratio of true positive samples among all samples determined to be positive, indicating the effectiveness of the model in reducing false positives. Recall, on the other hand, represents the proportion of actual positive samples that are successfully identified by the model, reflecting the model's performance in avoiding false negatives. Table 1 shows the comparison results of PMFFN and YOLOv8 baseline in three evaluation metrics. As shown in Table 1, the proposed model outperforms the baseline model in terms of precision, recall, and mean average precision, indicating that the proposed model is more superior for the detection of photovoltaic panel surface defects.

[0142] Table 1 Comparison Results of PMFFN Model and YOLOv8 Baseline Model in Three Evaluation Metrics

[0143] Model Precision / % Recall / % mAP / % YOLOv8 88.7 88 90.0 PMFFN 89.8 90.2 92.4

[0144] 3.4.2 Training Loss Curve Graph

[0145] The training loss graph includes the bounding box loss graph (Box_loss) and the classification loss graph (cls_loss). Among them, Box_loss measures the difference between the predicted bounding box and the true bounding box; cls_loss is used to measure the difference between the model's prediction of the class and the true label; training mAP is a comprehensive metric for evaluating the performance of the object detection model. Figure 10 Figure 4 shows the comparison graph of the training loss curves for the PMFFN model and the YOLOv8 baseline model. It can be seen from the figure that the loss curve of the PMFFN model drops faster and has a lower value than that of the YOLOv8 baseline model. This indicates that the PMFFN model proposed in the present invention has a better optimization effect and better convergence.

[0146] 3.4.3 Visualization

[0147] To visually demonstrate the differences between the proposed PMFFN model and the YOLOv8 baseline model, their detection results are visualized. As Figure 11 shown, in each subfigure, the YOLOv8 baseline model is above and the proposed PMFFN model is below. It can be seen from the figure that YOLOv8 has problems of missed detection and false detection compared to PMFFN, indicating that PMFFN is superior to the baseline model YOLOv8 in the performance of photovoltaic panel defect detection.

[0148] To more clearly demonstrate the advantages of the PMFFN model in feature extraction, the Grad-CAM tool is used to visualize and analyze the defect detection results of YOLOv8 and PMFFN. As Figure 12 shown, the attention map of PMFFN is clearly focused on the actual area of surface defects, while the attention map of YOLOv8 is more dispersed and has a wider coverage range. This results in its inability to accurately locate the specific position of the defect. Figure 12 shows the missed detection and false detection situations of the dataset visualization. The upper part of the figure is the baseline model YOLOv8, and the lower part is the proposed PMFFN model.

[0149] 3.5 Comparative Experiments

[0150] To evaluate the superiority of the proposed method in the task of photovoltaic panel surface defect detection, the experiment compares the performance of PMFFN with current mainstream advanced image detection models SSD, Efficientdet-d0, Retinenet, Faster-RCNN, YOLOv5, YOLOv8, RT-DETR(2024), YOLOv9(2024), YOLOv10(2024), RT-DETR(2024), YOLOv11(2024), YOLOv12(2025). The comparison results are shown in Table 2.

[0151] Table 2 Performance Comparison Results

[0152] Algorithm mAP / % SSD 79.45 Efficientdet 63.59 Retinenet 90.87 Faster - RCNN 90.73 YOLOv5 88.7 YOLOv8 90 YOLOv9 91.2 YOLOv10 90.4 RT - DETR 92.0 YOLOv11 91.8 YOLOv12 91.3 PMFFN 92.4

[0153] As can be seen from Table 2, compared with SSD, Efficientdet-d0, Retinenet, Faster-RCNN, YOLOv5, YOLOv8, the newly updated RT-DETR in 2024, YOLOv9, YOLOv10, YOLOv11 and YOLOv12, the algorithm of the present invention has the highest detection accuracy, with the mAP reaching 92.4%. The mAP has increased by 2.4% compared with the baseline model YOLOv8, and has increased by 1.2% and 2.0% respectively compared with the newly released YOLOv9 and YOLOv10. This proves that the improved model has the best detection performance for defective targets. Since each detection model uses different feature extraction networks and feature fusion methods for defect detection, the detection performance of different models for defect detection is also different. The model of the present invention improves the method of multi-scale feature fusion, adds feature fusion technology to the backbone network to enhance feature representation, and optimizes the upsampling method using the atrous convolution content awareness technology. These are the main reasons for the performance improvement. Therefore, even when compared with the newly released YOLOv9 and YOLOv10 new models and the recent YOLOv11 and YOLOv12, the model of the present invention can still demonstrate its performance advantages. In summary, the improved algorithm of the present invention exceeds all algorithms in the experiment in terms of detection accuracy.

[0154] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A photovoltaic panel surface defect detection method based on a pruning multi-scale feature fusion network, characterized in that It includes the following steps: S1: Dataset processing Select a basic dataset and perform partitioning processing on the dataset to obtain a training set and a test set; S2: Network construction Select the YOLOv8 network as the benchmark model, and design the FFPEMA module, DCCARAFE module, and PAFIFN module in the benchmark model to obtain the pruned multi-scale feature fusion network PMFFN; S3: Model training Use the training set to train the pruned multi-scale feature fusion network PMFFN to obtain a trained model, that is, a photovoltaic panel surface defect detection model; S4: Defect detection Input the samples in the test set into the photovoltaic panel surface defect detection model, and output the photovoltaic panel surface defect results.

2. The method for detecting surface defects of a photovoltaic panel based on a pruning multi-scale feature fusion network according to claim 1, wherein, In the step S2, the specific process of introducing the FFPEMA module, DCCARAFE module, and PAFIFN module is as follows: S21: Embed the PEMA module in the C2f module of the benchmark model backbone network, and construct a multi-level feature fusion architecture at the feature output end to form the FFPEMA module; S22: Introduce a dilated convolutional layer in the upsampling kernel prediction and information encoding links of the benchmark model neck network to form the DCCARAFE module; S23: Introduce the PAFIFN module at the output end of the last FFPEMA module in the benchmark model backbone network.

3. The method for detecting surface defects of a photovoltaic panel based on a pruned multi-scale feature fusion network according to claim 2, wherein In the FFPEMA module of the step S21, the Bottleneck layer is a depthwise separable convolution reconstructed Bottleneck layer. Using a cross-layer feature splicing mechanism, the global context features output by the PEMA module are fused with the local detail features extracted by the depthwise separable convolution reconstructed Bottleneck layer in the channel dimension. Feature interaction enhancement is achieved through gated convolution, and the improved ReLU activation function is used in the convolutional layer to eliminate the distribution differences between multi-source features; among them, the depthwise separable convolution reconstructed Bottleneck layer includes two convolutional layers. After the input features pass through the first convolutional layer, they are divided into two branches. After one branch is processed by the second convolutional layer, it is superimposed with the features on the other branch and then output. The convolutional layer includes a two-dimensional convolutional block, a two-dimensional batch normalization unit, and a SiLU activation function layer connected in sequence.

4. The method for detecting surface defects of a photovoltaic panel based on a pruning multi-scale feature fusion network according to claim 3, wherein The PEMA module is obtained by pruning the channels and weight parameters of the EMA module through an unstructured global pruning algorithm. In the EMA module, the input feature map first undergoes a series of convolutional and pooling operations to extract local and global context information, and then an attention map is generated through weighted aggregation. The attention map is combined with the original input feature map and then output. The weighted aggregation attention calculation formula is as follows: Among them, A is the generated attention map, which is used to assign weights to the input feature map; F i represents the context feature of the i-th channel extracted through convolution and pooling operations. The context feature includes local and global features; w i is the learnable weight parameter of the i-th channel, which is used to dynamically adjust the importance of features in different channels; σ is the Sigmoid activation function, which is used to map the weighted sum to the range [0,1] to generate normalized attention weights; C is the number of channels of the input feature map.

5. The photovoltaic panel surface defect detection method based on a pruning multi-scale feature fusion network according to claim 2, wherein In the step S21, the calculation formula of the FFPEMA module is as follows: FFPEMA[t] = β · EMA[t] + (1 - β) · y[t] where FFPEMA[t] represents the feature fusion output result at the current time step, EMA[t] represents the feature processed by exponential moving average, β represents the fusion weight, and y[t] represents another set of features for fusion.

6. The photovoltaic panel surface defect detection method based on the pruning multi-scale feature fusion network according to claim 3, wherein, In the step S22, the DCCARAFE module includes an upsampling kernel prediction module and a feature reconstruction module. The dilated convolution operation is introduced in the upsampling prediction and information encoding links of the model through the upsampling kernel prediction module and the feature reconstruction module, that is, a dilated convolution layer is adopted. Among them, the dilated convolution layer expands the receptive field without increasing the number of parameters by adjusting the dilation rate r in the dilated convolution.

7. The method for detecting surface defects of a photovoltaic panel based on a pruning multi-scale feature fusion network according to claim 6, wherein The calculation formula of the dilated convolution is as follows: Where x is the input feature map, y is the output feature map, w is the convolution kernel, r is the dilation rate used to control the interval between elements in the convolution kernel, k is the size of the convolution kernel, i and j are the position coordinates of the output feature map, representing the value of the i-th row and j-th column in the output feature map y, which is calculated by the weighted sum of the corresponding positions of the input feature map; m and n are the row and column indices in the convolution kernel, used to traverse each element of the convolution kernel, and the value range is 0 ≤ m, n ≤ k - 1.

8. The method for detecting surface defects of a photovoltaic panel based on a pruning multi-scale feature fusion network according to claim 6, characterized in that, In the step S23, the PAFIFN module dynamically adjusts the weights of features at different scales through an adaptive feature integration strategy, improves the model's detection ability for defects at different scales, and performs pruning-based efficient multi-scale feature fusion operation on the feature results generated by the adaptive feature integration to enhance the dynamics of the model's feature representation.

9. The photovoltaic panel surface defect detection method based on the pruning multi-scale feature fusion network according to claim 8, wherein, In the PAFIFN module, the weight adjustment formula for adaptive feature integration is as follows: where z i is the weight score of the i-th feature map, is the final weight of the i-th feature map, and N is the number of feature maps; In pruning-based efficient multi-scale feature fusion, the loss function L of the pruning strategy prune is as follows: wherein, is the i-th weight, λ is the pruning strength parameter, and M is the total number of weights.

10. The method for detecting surface defects of a photovoltaic panel based on a pruning multi-scale feature fusion network according to claim 8, wherein, In the pruning-based efficient multi-scale feature fusion, feature optimization is achieved through the following steps: S231: Multi-scale feature integration Perform channel - dimension concatenation on the multi - scale feature maps {F1, F2, …, F n} to form an integrated feature map F intergrated = Concat{α1F1, α2F2, …, α N F N}, where α i represents the dynamic weight, which is a weight generated based on the input feature content; S232: Structured global pruning The structured pruning strategy is used to sparsify the channel weights of the integrated feature map. The formula for structured global pruning is as follows: Among them, A c,h,w is the activation value of the c-th channel at the spatial position (h, w); H and W are the height and width of the feature map respectively; S233: Efficient feature fusion The pruned feature map is fused across scales through gated convolution. The formula is as follows: F fused = GatConv(F puned , F skip ) Among them, F fused is the fused feature map; F puned is the pruned feature map; F skip retains the fine-grained information of the original features, which is dynamically weighted and fused with F puned through gated convolution; GatConv represents gated convolution.

Citation Information

Cited By

  • Photovoltaic panel fault detection method based on multi-scale scaling feature fusion network

    CN119516274A

  • A photovoltaic panel fault detection method based on a multi-scale scaling feature fusion network

    CN119516274B

  • Production environment mobile phone illegal use detection method and system based on YOLOv12 optimization

    CN120689760A

  • Industrial product surface defect detection method based on feature coupling

    CN120766047A

  • A feature coupling-based industrial product surface defect detection method

    CN120766047B