Solar cell panel defect detection method based on GELAN

By introducing LiteRepLAN, SpdConv, DySample and Enhanced_DDetect modules on the GELAN model, the feature extraction and detection process are optimized, and the problem of high- and multi-scale defect detection in the existing technology is solved, and efficient and low-cost solar panel defect detection is achieved.

CN120147259AActive Publication Date: 2025-06-13CHINA THREE GORGES UNIV

Patent Information

Application Number
CN202510218828.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Existing solar panel defect detection methods are difficult to achieve low-cost and efficient deployment in industrial production environments with limited computing resources, and at the same time, they are difficult to take into account both accuracy and robustness when dealing with multi-scale and morphological defects.

Method used

A solar panel defect detection method based on GELAN model is proposed. By introducing LiteRepLAN module, SpdConv module, DySample module and Enhanced_DDetect detection head, the feature extraction, fusion and detection process are optimized, the calculation overhead is reduced and the detection accuracy is improved.

Benefits of technology

It realizes efficient and high-precision solar panel defect detection in environments with limited computing resources, with a 4.1% increase in detection accuracy and a 65% reduction in parameter quantity and 59% reduction respectively, which is suitable for low-cost deployment in industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147259A_ABST
    Figure CN120147259A_ABST
Patent Text Reader

Abstract

The invention discloses a solar cell panel defect detection method based on a GELAN, and provides a solution through an improved LiteScale-GELAN model for the bottleneck of the existing method under the condition that multi-scale defect detection and computing resources are limited. According to the model, a lightweight multiplexing convolution module is designed, depth separable convolution and a multi-scale feature aggregation mechanism are combined, and the feature extraction and fusion performance is optimized while the calculation amount is remarkably reduced. A space-to-depth conversion strategy is adopted to improve the extraction capability of small targets and multi-scale defects, a dynamic up-sampling mechanism is introduced, and the detection precision and network robustness of the small targets are enhanced through an improved sampling point generation method. In addition, the reconstructed detection head maintains the efficient detection capability, and meanwhile, the parameter quantity and the computing resource consumption are reduced. Experimental results show that the precision of the LiteScale-GELAN on a PVEL-AD data set is 89.1%, which is improved by 4.1% compared with that of a GELAN model, and the parameter quantity and the calculated quantity are respectively reduced by 65% and 59%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, in particular to object detection technology, and specifically to a method for detecting and identifying defects of solar panels based on the GELAN network, and particularly to an efficient and high-precision method for detecting defects of solar panels suitable for environments with limited computing resources. Background Art

[0002] The importance of solar panel defect detection stems from the wide application of photovoltaic power generation systems and their direct impact on power generation efficiency and operating costs. In particular, the defects and problems that are inevitable during the production process make efficient detection a key requirement.

[0003] Although existing solar panel defect detection methods have achieved certain results, they still face the following challenges:

[0004] (1) Dependence on complex network structures, although ensuring detection accuracy, requires high computing resources. Especially in industrial production environments, computing efficiency and low-cost deployment become difficult problems.

[0005] (2) There are various types of solar panel defects, from small target defects such as fine cracks and fingerprints to large-scale defects such as short circuits. Coupled with complex background interference, it is difficult for existing models to balance accuracy and robustness when dealing with multi-scale and morphological defects.

[0006] In order to better detect solar panel defects, scholars have conducted a large number of studies on it. For example, the patent document with the application number CN202410622641.4 discloses a method for detecting solar panel defects based on improved Yolov7-tiny. It replaces the convolutional layer, introduces the Conv-ATT module, modifies the convolutional layer and loss function of the detection head, and fuses multi-level features and multi-scale detection. After training the model, it is used for solar panel defect detection, effectively improving the detection accuracy. Another example is the application number CN202410344938.9, which discloses a method for detecting solar panel defects based on improved YOLOv5. By collecting sample data to construct a data set, data augmentation, building and training the PV-YOLO network model, and finally using the trained model for defect detection. The PV-YOLO network includes feature extraction, feature fusion, and multi-scale detection branches, improving the detection accuracy and helping enterprises achieve automatic and efficient detection of solar panel defects.

[0007] However, these existing technologies still have technical problems such as missed detection of multi-scale defects, poor detection effects under sample imbalance conditions, and computing resource bottlenecks.

[0008] Therefore, in view of the problem of multi-scale defect recognition in the defect detection task of solar panels, especially in the actual situation where computing resources are limited, the present invention proposes a method for defect detection of solar panels based on the GELAN model, LiteScale-GELAN. Summary of the Invention

[0009] The purpose of the present invention is to solve the technical problem that existing deep learning-based defect detection methods for solar panels, although able to ensure high detection accuracy, usually rely on complex network structures, resulting in high computing resource requirements. Especially in industrial production environments, it is difficult to achieve low-cost and efficient deployment. Therefore, the present invention provides an efficient and high-precision defect detection method for solar panels suitable for environments with limited computing resources.

[0010] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0011] A method for defect detection of solar panels based on GELAN, comprising the following steps:

[0012] Step 1: Collect sample images of the defect types of solar panels to construct a data set, label the samples with fewer defect categories to supplement the data set and generate more training samples, and randomly allocate the data set into a training set and a test set according to a certain ratio for model training and evaluation respectively;

[0013] Step 2: Perform preliminary data processing on the data set;

[0014] Step 3: Construct an optimized LiteScale-GELAN network;

[0015] Step 4: Train the optimized LiteScale-GELAN network to identify various defect types on the solar panels;

[0016] Step 5: Through the trained model, perform detection and evaluation on the test set to analyze the performance of the model;

[0017] Step 6: Optimize the model by adjusting the network structure and training parameters;

[0018] Detect various defect types on the solar panels through the above steps.

[0019] In Step 1, collect sample images containing these defect types: linear cracks, star cracks, broken grids, black hearts, scratches, fragments, and fingerprints to construct a data set. Use annotation software to label the samples with fewer defect categories (such as corner defects and fragments) to supplement the data set and generate more training samples. Randomly allocate the data set into a training set and a test set according to a ratio of 8:2 for model training and evaluation respectively.

[0020] In step 2, data augmentation is performed on the dataset, including basic operations such as rotation, flipping, scaling, and cropping, to further increase the sample diversity of the training set and improve the generalization ability of the model.

[0021] In step 3, the optimized LiteScale-GELAN network improves the feature extraction, fusion, and detection efficiency through multiple innovative modules. First, the LiteRepLAN module is adopted to enhance the network's feature extraction and fusion capabilities and improve the capture effect of fine-grained information. Second, SpdConv is used to replace the traditional convolutional layer, effectively retaining more detailed information while maintaining light weight, especially for the detection ability of small-scale defects. To optimize the neck network, the Dysample upsampling mechanism is adopted to ensure the effective retention of features at different scales. Finally, through the reconstructed Enhanced_DDetect detection head, the computational overhead in multi-scale defect detection is significantly reduced, while maintaining the detection accuracy and efficiency.

[0022] In step 4, the defect types include cracks, short circuits, burned boards, fingerprints, etc.

[0023] In step 5, evaluate the accuracy of the model under different defect types and scales;

[0024] In step 6, the tuning is to further improve the detection accuracy and efficiency and ensure the application feasibility of the model in actual industrial production.

[0025] It also includes step 7, which conducts sufficient tests on camera devices or embedded devices to ensure the accuracy and stability of the solar panel defect detection method in industrial production scenarios.

[0026] The obtained optimized LiteScale-GELAN network is specifically as follows:

[0027] The original input image is fed into the first convolutional layer (Conv) of the backbone, and the feature map F1 is output by this layer. The feature map F1 is input into the second convolutional layer to obtain the feature map F2; then, the feature map F2 is input into the third ELAN1 module, and the feature map F3 is output; next, the feature map F3 is input into the fourth spatial-depth convolutional module (SpdConv), and the feature map F4 is output; subsequently, the feature map F4 is input into the fifth lightweight feature aggregation module (LiteRepLAN), and the features are aggregated through 3×3 convolutional operations, and the feature map F5 is output. The feature map F5 is input into the sixth spatial-depth convolutional module (SpdConv), and the feature map F6 is output; then, the feature map F6 is input into the seventh LiteRepLAN to continue optimizing feature fusion, and the feature map F7 is output. The feature map F7 is input into the eighth SpdConv, and the feature map F8 is output; finally, the feature map F8 is input into the ninth LiteRepLAN to further aggregate and extract features, and the final feature map F9 is output. The output F9 of the ninth layer is used as the input of the tenth spatial pyramid pooling module (SPPELAN) to further enhance the multi-scale feature expression ability.

[0028] The output feature map F10 of the tenth layer is used as the input of the first dynamic sampling module (DySample) of the neck module, and the adjusted feature map F11 is output; the feature map F7 of the seventh layer of the backbone and the feature map F11 after the first dynamic sampling of the neck are input into the second concatenation operation (Concat) of the neck, and the concatenated feature map F12 is output; the concatenated feature map F12 is input into the third LiteRepLAN of the neck to obtain the fused feature map F13.

[0029] The fourth layer of the neck module is the dynamic sampling module (DySample). The feature map F13 output by the third layer is input, and the feature map F14 after dynamic sampling is output; the feature map F5 of the fifth layer of the backbone and the feature map F14 after the fourth dynamic sampling of the neck are input into the fifth concatenation operation (Concat) of the neck, and the concatenated feature map F15 is output; the concatenated feature map is input into the sixth LiteRepLAN of the neck to obtain the fused feature map F16.

[0030] Next, input the feature map F16 output from the sixth layer of the neck to the seventh layer SpdConv to obtain the feature map F17; input the feature map F13 of the third layer of the neck and the feature map F17 output from the seventh layer of the neck to the eighth layer of the neck for a concatenation operation (Concat), and then input the concatenated feature map F18 to the ninth layer LiteRepLAN of the neck to output the fused feature map F19; input the feature map F19 output from the ninth layer of the neck to the tenth layer SpdConv to obtain the feature map F20.

[0031] Input the feature map F20 output from the tenth layer of the neck and the feature map F10 output from the tenth layer of the backbone to the eleventh layer of the neck for a concatenation operation (Concat), and then input the concatenated feature map F21 to the twelfth layer LiteRepLAN of the neck to output the fused feature map F22.

[0032] Finally, the outputs of the sixth layer F16, the ninth layer F19, and the twelfth layer F22 of the neck module are combined and input to the detection head (Enhanced_DDetect) module to finally output multi-scale detection results.

[0033] The LiteRepLAN module is specifically as follows:

[0034] Input the original feature map to the first layer LiteRepCSP module. This module first performs preliminary processing on the input feature map through depthwise separable convolution (DSC) to extract compact features and reduce channel redundancy, and outputs the feature map P1; then, divide the feature map P1 into two paths and send them to two RepBottleneck modules respectively. Each path performs deep feature extraction through multiple convolutional layers. Among them, the upper path branch performs feature extraction through a stacked RepBottleneck module, and the number of channels remains unchanged for each layer to further capture deeper feature information and output the feature map P2; the lower path branch retains the original feature information through a 1×1 depthwise separable convolution and performs channel fusion through pointwise convolution to output the feature map P3; then, fuse these two feature maps P2 and P3 through a Concat operation to obtain a more multi-scale information feature map P4; finally, output the feature map P4 to the next layer of depthwise separable convolution for fine-grained feature processing to generate the final feature map P5; this module effectively realizes efficient feature extraction and cross-channel information fusion by combining the RepBottleneck module and depthwise separable convolution.

[0035] The detection head module is the Enhanced_DDetect module, and this module is specifically as follows:

[0036] The input feature map is initially processed by depthwise separable convolution (DSC). Depthwise separable convolution decomposes the traditional convolution operation into two steps: first, depthwise convolution is used for spatial feature extraction, and then pointwise convolution is used for channel fusion. The output feature map P1 is split into two paths and enters the cv2 and cv3 branches respectively for further feature extraction. In the cv2 branch, the input feature map undergoes two depthwise separable convolutions to extract spatial features. Then, an additional convolutional layer is used to further fuse channel information, thereby generating feature map P2. Different from the cv2 branch, the cv3 branch adds a channel attention mechanism on the basis of two depthwise separable convolutions. The channel attention mechanism captures the global information of each channel through adaptive average pooling, and uses two layers of 1x1 convolution and Sigmoid activation function to dynamically adjust the response of each channel, thereby enhancing the sensitivity of the model to key channel features. This mechanism helps to improve the accuracy of target classification. The processed feature map P3 is refined by further 1x1 convolution (Conv2d) to obtain the final feature map P4. Finally, the output feature maps P2 and P4 of the cv2 and cv3 branches are merged through a concatenation operation to form a joint feature map P5. This concatenation operation ensures the effective fusion of features for regression (bounding box) and classification tasks. After regression and classification processing, the output results of object detection are generated, including bounding box regression (BBox) and object category prediction (Cls).

[0037] Compared with the prior art, the present invention has the following technical effects:

[0038] 1) The present invention introduces the LiteRepLAN module: This module combines depthwise separable convolution and multi-scale feature aggregation mechanism, adopts a lightweight structure design, significantly reduces the computational overhead, and at the same time enhances the network's ability in multi-scale defect detection. Whether it is large target defects (such as short-circuited burned boards) or small target defects (such as cracks, fingerprints, etc.), efficient and accurate identification can be achieved.

[0039] 2) The present invention adopts the Enhanced_DDetect detection head: The present invention reconstructs the structure of the detection head and proposes the Enhanced_DDetect detection head. By combining depthwise separable convolution and channel attention mechanism, this detection head reduces the computational overhead in multi-scale defect detection, while maintaining efficient detection ability, ensuring the balance between precision and efficiency.

[0040] 3) The present invention introduces the SpdConv module, which adopts a spatial-to-depth conversion strategy, optimizes the feature extraction process, and retains more fine-grained feature information in the downsampling stage. This design not only improves the detection accuracy of micro-defects, but also ensures the lightweight of the network, facilitating deployment and popularization in practical application environments with limited computing resources.

[0041] 4) The present invention introduces the DySample module: By adopting an optimized upsampling mechanism, this module retains the detailed information in the feature map. Its dynamic sampling point generation strategy further improves the detection accuracy of small target defects, significantly enhances the overall detection effect, and can still ensure efficient detection especially in an environment with limited computing resources.

[0042] Experimental results show that the detection accuracy of LiteScale-GELAN on the PVEL-AD dataset reaches 89.1%, which is 4.1% higher than that of the GELAN model, and the number of parameters and the amount of computation are reduced by 65% and 59% respectively. In addition, the experimental results of the model on the PVMD dataset further verify its effectiveness and adaptability in different environments.

[0043] The LiteScale-GELAN network proposed by the present invention, with its efficient utilization of computing resources and accurate detection ability, provides an efficient solar panel defect detection solution with low computing resource consumption and adaptation to multi-scale features, and has important industrial application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The present invention will be further described below in conjunction with the drawings and embodiments:

[0045] Figure 1 is the flowchart of the solar panel defect detection of the present invention;

[0046] Figure 2 is the structural diagram of the GELAN network in the prior art;

[0047] Figure 3 is the structural diagram of the LiteScale-GELAN network proposed by the present invention;

[0048] Figure 4 is the schematic diagram of the LiteRepLAN module;

[0049] Figure 5 is the schematic diagram of the SpdConv module;

[0050] Figure 6 is the schematic diagram of the dysample module;

[0051] Figure 7 is the schematic diagram of two sampling point generation methods in the dysample;

[0052] Figure 8 is the schematic diagram of the Enhanced_DDetect detection head;

[0053] Figure 9 is the schematic diagram of the channel attention mechanism;

[0054] Figure 10 It is a diagram of the test part for detection results. Specific implementation manners

[0055] A method for detecting defects in solar panels based on the LiteScale-GELAN network. This method first introduces a lightweight feature aggregation module, combines depthwise separable convolution and multi-scale feature aggregation mechanisms, significantly reducing the computational overhead and enhancing the detection ability for multi-scale defects; then adopts a spatial-to-depth conversion strategy to optimize the feature extraction process, improving the adaptability to small targets and complex scenes and retaining more fine-grained feature information; next, the dynamic sampling optimization module optimizes the upsampling process through a dynamic sampling point generation strategy; finally, the optimized detection head module combines depthwise separable convolution and channel attention mechanism to optimize the traditional detection head and reduce the computational overhead. This method greatly improves the detection efficiency while ensuring high accuracy and is applicable to the detection of solar panel defects of different scales.

[0056] The present invention includes the following steps:

[0057] Step 1: Collect sample pictures containing various defect types such as linear cracks, star cracks, broken grids, black hearts, scratches, fragments, fingerprints, etc., and construct a data set therefrom. Samples with fewer defect categories (such as corner defects and fragments) are labeled through annotation software to supplement the data set and generate more training samples. The data set is randomly assigned to a training set and a test set in a ratio of 8:2 for model training and evaluation respectively.

[0058] Step 2: Perform data augmentation on the data set, including basic operations such as rotation, flipping, scaling, and cropping, to further increase the sample diversity of the training set and enhance the generalization ability of the model.

[0059] Step 3: The present invention constructs an optimized LiteScale-GELAN network, which improves the feature extraction, fusion, and detection efficiency through multiple innovative modules. First, the LiteRepLAN module is adopted to enhance the feature extraction and fusion ability of the network and improve the capture effect of fine-grained information. Secondly, SpdConv is used to replace the traditional convolutional layer, effectively retaining more detailed information while maintaining light weight, especially for the detection ability of small-scale defects. In order to optimize the neck network, the Dysample upsampling mechanism is adopted to ensure the effective retention of features at different scales. Finally, through the reconstructed Enhanced_DDetect detection head, the computational overhead in multi-scale defect detection is significantly reduced, while the detection accuracy and efficiency are improved.

[0060] Step 4: Train the optimized LiteScale-GELAN network to identify various types of defects on the solar panels, including cracks, short circuits, burned plates, fingerprints, etc.

[0061] Step 5: Through the trained model, conduct detection and evaluation on the test set, analyze the performance of the model, and evaluate its accuracy under different defect types and scales.

[0062] Step 6: Optimize the model, adjust the network structure and training parameters to further improve the detection accuracy and efficiency, and ensure the application feasibility of the model in actual industrial production.

[0063] Step 7: Conduct sufficient tests on camera devices or embedded devices to ensure the accuracy and stability of the solar panel defect detection method in industrial production scenarios.

[0064] Through the above steps, the present invention can efficiently and accurately detect various types of defects on solar panels, adapt to different defect scales, and is suitable for applications in actual production.

[0065] As Figure 2 shown, it is the network structure of the existing GELAN model.

[0066] As Figure 3 shown, the LiteScale-GELAN network constructed in Step 4 is as follows:

[0067] Input the original image into the first convolutional layer (Conv) of the backbone, and this layer outputs the feature map F1. Input the feature map F1 into the second convolutional layer to obtain the feature map F2; then, input the feature map F2 into the third ELAN1 module to output the feature map F3; next, input the feature map F3 into the fourth spatial depth convolutional module (SpdConv) to output the feature map F4; then, input the feature map F4 into the fifth lightweight feature aggregation module (LiteRepLAN), and aggregate the features through 3×3 convolutional operations to output the feature map F5. Input the feature map F5 into the sixth spatial depth convolutional module (SpdConv) to output the feature map F6; then, input the feature map F6 into the seventh LiteRepLAN to continue optimizing feature fusion and output the feature map F7. Input the feature map F7 into the eighth SpdConv to output the feature map F8; finally, input the feature map F8 into the ninth LiteRepLAN to further aggregate and extract features and output the final feature map F9. The output F9 of the ninth layer is used as the input of the tenth spatial pyramid pooling module (SPPELAN) to further enhance the multi-scale feature expression ability.

[0068] The output feature map F10 of the tenth layer is used as the input of the first layer of the neck module, the dynamic sampling module (DySample), and the adjusted feature map F11 is output; the feature map F7 of the seventh layer of the backbone and the feature map F11 after the first layer of dynamic sampling of the neck are input into the second layer of the neck for concatenation operation (Concat), and the concatenated feature map F12 is output; the concatenated feature map F12 is input into the third layer of the neck, LiteRepLAN, to obtain the fused feature map F13.

[0069] The fourth layer of the neck module is the dynamic sampling module (DySample), which inputs the feature map F13 output by the third layer and outputs the feature map F14 after dynamic sampling; the feature map F5 of the fifth layer of the backbone and the feature map F14 after the fourth layer of dynamic sampling of the neck are input into the fifth layer of the neck for concatenation operation (Concat), and the concatenated feature map F15 is output; the concatenated feature map is input into the sixth layer of the neck, LiteRepLAN, to obtain the fused feature map F16.

[0070] Next, the feature map F16 output by the sixth layer of the neck is input into the seventh layer, SpdConv, to obtain the feature map F17; the feature map F13 of the third layer of the neck and the feature map F17 output by the seventh layer of the neck are input into the eighth layer of the neck for concatenation operation (Concat), and then the concatenated feature map F18 is input into the ninth layer of the neck, LiteRepLAN, and the fused feature map F19 is output; the feature map F19 output by the ninth layer of the neck is input into the tenth layer, SpdConv, to obtain the feature map F20.

[0071] The feature map F20 output by the tenth layer of the neck and the feature map F10 output by the tenth layer of the backbone are input into the eleventh layer of the neck for concatenation operation (Concat), and then the concatenated feature map F21 is input into the twelfth layer of the neck, LiteRepLAN, and the fused feature map F22 is output.

[0072] Finally, the outputs of the sixth layer F16, the ninth layer F19, and the twelfth layer F22 of the neck module are combined and input into the detection head (Enhanced_DDetect) module, and the multi-scale detection results are finally output.

[0073] As Figure 4 shown, the LiteRepLAN module is as follows:

[0074] The original feature map is input into the first - layer LiteRepCSP module. This module first preliminarily processes the input feature map through depth - wise separable convolution (DSC) to extract compact features and reduce channel redundancy, and outputs the feature map P1. Then, the feature map P1 is divided into two paths and fed into two RepBottleneck modules respectively. Each path extracts deep features through multiple convolutional layers. Among them, the upper - path branch extracts features through a multi - layer stacked RepBottleneck module, with the number of channels remaining unchanged in each layer to further capture deeper - level feature information and outputs the feature map P2; the lower - path branch retains the original feature information through a 1×1 depth - wise separable convolution and performs channel fusion through point - wise convolution, outputting the feature map P3. Then, these two feature maps P2 and P3 are fused through a Concat operation to obtain a more multi - scale information - rich feature map P4. Finally, the output feature map P4 enters the next - layer depth - wise separable convolution for fine - grained feature processing to generate the final feature map P5. This module effectively realizes efficient feature extraction and cross - channel information fusion by combining the RepBottleneck module and depth - wise separable convolution.

[0075] As Figure 8 shown, the Enhanced_DDetect detection head is as follows:

[0076] The Enhanced_DDetect module efficiently processes and fuses features through an innovative dual - branch structure (cv2 and cv3), combining depth - wise separable convolution (Depthwise Separable Convolution, DSC), channel attention mechanism (Channel Attention), and feature splicing technology. Specifically as follows:

[0077] The output feature map P1 is split into two paths, entering the cv2 and cv3 branches respectively for further feature extraction. In the cv2 branch, the input feature map goes through two depthwise separable convolutions to extract spatial features. Then, an additional convolutional layer is used to further fuse channel information, thus generating the feature map P2. Different from the cv2 branch, the cv3 branch adds a channel attention mechanism on the basis of two depthwise separable convolutions. The channel attention mechanism captures the global information of each channel through adaptive average pooling, and uses two layers of 1x1 convolutions and the Sigmoid activation function to dynamically adjust the response of each channel, thereby enhancing the model's sensitivity to key channel features. This mechanism helps to improve the accuracy of object classification. The processed feature map P3 is refined by a further 1x1 convolution (Conv2d) to obtain the final feature map P4. Finally, the output feature maps P2 and P4 of the cv2 and cv3 branches are merged through a concatenation operation to form the combined feature map P5. This concatenation operation ensures the effective fusion of the features for regression (bounding box) and classification tasks. After regression and classification processing, the output results of object detection are generated, including bounding box regression (BBox) and object class prediction (Cls).

[0078] Complete the network construction, start training, and save the trained model;

[0079] Use the trained model to test on the test set. The main evaluation metrics include mAP (mean average precision), Recall, and Precision. mAP represents the mean of the average precision (AP) of each defective object, which is used to measure the comprehensive performance of the model on different categories; Precision represents the proportion of actual positive samples among the samples predicted as positive by the model, measuring the detection accuracy of the model; Recall represents the proportion of actual positive samples correctly detected as positive by the model, reflecting the detection recall ability of the model.

[0080] The specific calculation formula for Precision is as follows:

[0081] Precision = TP / (TP + FP)

[0082] Among them, TP represents the number of correctly detected positive samples, and FP represents the number of negative samples misdetected as positive samples.

[0083] The specific calculation formula for Recall:

[0084] Recall = TP / (TP + FN)

[0085] Among them, FN represents the number of missed positive samples.

[0086] AP (Average Precision) is the integral area of precision at different recall rates, with its value ranging from 0 to 1, reflecting the comprehensive detection performance of the model at different detection thresholds. The specific calculation formula is as follows:

[0087] In addition, the number of model parameters (Parameters) and the computational complexity (GFLOPs) are particularly crucial in the detection of solar panel defects, as they directly affect the computational resources required for detection. An efficient model design should minimize the number of parameters and the computational complexity while ensuring high precision, so as to improve the detection speed and deployment feasibility.

[0088] Example:

[0089] The present invention uses the PVEL-AD dataset, which covers various defects such as linear cracks, star cracks, broken grids, black hearts, scratches, and fragments. Since the number of samples for some defect categories (such as corner defects and fragments) is small, to improve data balance and enhance the generalization ability of the model, the present invention uses annotation software to perform additional annotation on them to expand the data volume and improve the detection accuracy. At the same time, through data augmentation strategies such as rotation, flipping, scaling, and cropping, the number of defect class samples is further increased. Finally, the processed dataset contains 4600 images and is randomly divided into a training set and a test set at a ratio of 8:2 to ensure training stability and evaluation reliability.

[0090] The experiments were conducted on an NVIDIA GeForce RTX 3090 GPU (with 24GB of video memory), the operating system was Linux, the Python version was 3.10.14, the deep learning framework used was PyTorch 2.3.1, and the CUDA version was 12.1. To ensure the fairness of the experiments, a unified parameter configuration was used in the training phase: the input image size was 640×640 pixels, the optimizer was selected as SGD, the learning rate was 0.01, the momentum was 0.937, the weight decay coefficient was 0.0005, the number of iterations was 600, and the batch size was 16.

[0091] The results of the ablation experiments are shown in Table 1, and the "√" in the experiments indicates that the corresponding module was introduced.

[0092] Table 1 Ablation Experiments

[0093]

[0094] The results show that after adding LiteRepLAN, the detection accuracy slightly decreases, but the number of parameters is reduced by about 35%, and the GFLOPs drops to 5.1G, indicating that LiteRepLAN effectively reduces the number of model parameters and the occupancy of hardware resources by optimizing the network structure. After introducing SpdConv, the detection accuracy is improved to 88.1%, the number of parameters drops to 1.732M, and the GFLOPs drops to 6.7G, showing that SpdConv further reduces the hardware resource requirements while improving the detection accuracy, thus enhancing the overall efficiency. After adding Enhanced_DDetect, the number of parameters drops to 1.483M and the GFLOPs drops to 5.3G, indicating that this detection head reduces the occupancy of computing resources and enhances the advantages of the model in a low-computing-resource environment. After adding DySample, the detection accuracy is improved to 88.4%, and the computational overhead and the number of parameters remain stable, indicating that DySample further improves the detection accuracy of small targets by optimizing the upsampling mechanism while maintaining the lightweight computing requirements.

[0095] Compared with the baseline model, LiteScale-GELAN improves the detection accuracy by 4.1%, and reduces the number of parameters and the computational volume by 65% and 59% respectively, demonstrating its potential for efficient and low-cost applications in industrial production.

[0096] In addition, the results of the comparative experiments are shown in Table 2.

[0097] Table 2 Comparative Experiments

[0098]

[0099] The present invention conducted 8 groups of comparative experiments and compared them with mainstream object detection models such as YOLOv5n, YOLOv8n, YOLOv10n, YOLOv11n, and DETR. The experimental results are shown in Table 2. The experimental results indicate that LiteScale-GELAN achieves a detection accuracy of 89.1%, which is superior to YOLOv5n (84.4%), YOLOv8n (87.2%), and YOLOv11n (86.7%), and the accuracy is improved by more than 5% based on YOLOv10n and DETR. This model not only has excellent accuracy but also performs well in low-resource environments. The recall rate of LiteScale-GELAN is 97%, which is significantly better than YOLOv5n and YOLOv8n, demonstrating a stronger ability to reduce missed detections. In terms of model size, LiteScale-GELAN is only 1.9MB, which is much lower than YOLOv8 (5.98MB), DETR (63.15MB), and YOLOv11n (5.25MB), effectively reducing the storage requirements and being suitable for deployment on embedded and edge computing devices. Although the accuracy of YOLOv11s is 89.9%, slightly higher than LiteScale-GELAN, the number of parameters (0.682M) and GFLOPs (2.9G) of LiteScale-GELAN are approximately 92.75% and 86.32% less than 9.41M and 21.2G of YOLOv11s respectively, showing its advantage in computational resource occupancy.

[0100] To verify the generalization ability of the LiteScale-GELAN model, the present invention uses the PVMD dataset for experiments. This dataset contains five types of defects: hot_spot, black_border, no_electricity, broken, and scratch, with a total of 1105 samples. The dataset is randomly divided into a training set and a test set at a ratio of 8:2. Except for the experimental settings, other parameters are the same as those in the previous experiment. The experimental results are shown in Table 3.

[0101] Table 3 Results of PVMD Dataset

[0102]

[0103]

[0104] On this dataset, the LiteScale-GELAN model performs excellently, with a detection accuracy of 89.1%, which is very close to 89.4% of the benchmark model, indicating that the improved model performs similarly to the benchmark model in conventional defect detection tasks. In addition, the recall rate of LiteScale-GELAN is 97%, which is significantly improved compared to 93% of the benchmark model, showing its advantage in reducing missed detections.

Claims

1. A solar panel defect detection method based on GELAN, characterized in that: The following steps are involved: Step 1: Collect sample images of defect types of solar panels to build a data set, annotate samples with fewer defect categories to supplement the data set, and generate more training samples. The data set is divided into training set and test set according to a certain ratio, which are used for model training and evaluation respectively; Step 2: Perform preliminary data processing on the dataset; Step 3: Build an optimized LiteScale-GELAN network; Step 4: Train the optimized LiteScale-GELAN network to identify various defect types on solar panels. Step 5: Use the trained model to perform detection and evaluation on the test set to analyze the performance of the model; Step 6: Tune the model, adjust the network structure and training parameters; The above steps are used to detect various defect types on solar panels.

2. The method according to claim 1, characterized in that: In step 1, sample images of defect types such as linear cracks, star-shaped cracks, broken grids, black hearts, scratches, and fragments are collected to construct a data set. Samples with fewer defect categories are annotated using annotation software to supplement the data set and generate more training samples. The data set is randomly divided into a training set and a test set in a ratio of 8:2, which are used for model training and evaluation, respectively.

3. The method according to claim 1, characterized in that In step 2, data enhancement is performed on the dataset, including rotation, flipping, scaling, and cropping, to further increase the sample diversity of the training set and improve the generalization ability of the model.

4. The method according to claim 1, characterized in that: In step 3, the optimized LiteScale-GELAN network improves the efficiency of feature extraction, fusion and detection through multiple innovative modules; first, the LiteRepLAN module is used to enhance the network's feature extraction and fusion capabilities, and improve the capture of fine-grained information; second, SpdConv is used to replace the traditional convolutional layer to effectively retain more detail information while maintaining lightweight, especially for the detection of small-scale defects; in order to optimize the neck network, the Dysample upsampling mechanism is used to ensure the effective retention of features of different scales; finally, through the reconstructed Enhanced_DDetect, the computational overhead in multi-scale defect detection is significantly reduced while maintaining detection accuracy and efficiency.

5. The method according to claim 1, characterized in that: In step 4, the defect types include cracks, short circuits, burnt boards, and fingerprints.

6. The method according to claim 1, characterized in that In step 5, the accuracy of the model is evaluated under different defect types and scales; In step 6, the tuning is to further improve the detection accuracy and efficiency and ensure the feasibility of the model in actual industrial production.

7. The method according to claim 1, characterized in that It also includes step 7, conducting sufficient testing on a camera device or embedded device to ensure the accuracy and stability of the solar panel defect detection method in an industrial production scenario.

8. The method according to claim 1, characterized in that: The optimized LiteScale-GELAN network obtained is as follows: Input the original image to the first convolution layer Conv of the backbone, which outputs the feature map F1; input the feature map F1 to the second convolution layer to obtain the feature map F2; then, input the feature map F2 to the third layer ELAN1 module, output the feature map F3; then, input the feature map F3 to the fourth layer spatial depth convolution module SpdConv, output the feature map F4; next, input the feature map F4 to the fifth layer lightweight feature aggregation module LiteRepLAN, aggregate the features through 3×3 convolution operation, and output the feature map F5; the feature map F 5 Input to the sixth layer of spatial depth convolution module SpdConv, output feature map F6; then, input feature map F6 to the seventh layer LiteRepLAN, continue to optimize feature fusion, output feature map F7; input feature map F7 to the eighth layer SpdConv, output feature map F8; finally, input feature map F8 to the ninth layer LiteRepLAN, further aggregate and extract features, and output the final feature map F9; the output F9 of the ninth layer is used as the input of the tenth layer spatial pyramid pooling module SPPELAN4 to further improve the multi-scale feature expression capability; The output feature map F10 of the tenth layer is used as the input of the dynamic sampling module DySample of the first layer of the neck module, and the adjusted feature map F11 is output; the feature map F7 of the seventh layer of the backbone and the feature map F11 after dynamic sampling of the first layer of the neck are input into the splicing operation of the second layer of the neck, and the spliced ​​feature map F12 is output; the spliced ​​feature map F12 is input into LiteRepLAN of the third layer of the neck, and the fused feature map F13 is obtained; The fourth layer of the neck module is the dynamic sampling module DySample, which inputs the feature map F13 output by the third layer and outputs the dynamically sampled feature map F14; the feature map F5 of the fifth layer of the backbone and the feature map F14 after dynamic sampling of the fourth layer of the neck are input into the fifth layer of the neck for concatenation, and the concatenated feature map F15 is output; the concatenated feature map is input into the sixth layer of the neck LiteRepLAN to obtain the fused feature map F16; Next, the feature map F16 output from the sixth layer of the neck is input to the seventh layer SpdConv to obtain the feature map F17; the feature map F13 output from the third layer of the neck and the feature map F17 output from the seventh layer of the neck are input to the eighth layer of the neck for splicing, and then the spliced ​​feature map F18 is input to the ninth layer LiteRepLAN of the neck to output the fused feature map F19; the feature map F19 output from the ninth layer of the neck is input to the tenth layer SpdConv to obtain the feature map F20; The feature map F20 output from the tenth layer of the neck and the feature map F10 output from the tenth layer of the backbone are input into the eleventh layer of the neck for concatenation, and then the concatenated feature map F21 is input into the twelfth layer of the neck LiteRepLAN, and the fused feature map F22 is output; Finally, the outputs of the sixth layer F16, the ninth layer F19 and the twelfth layer F22 of the neck module are merged and input into the Enhanced_DDetect module, which finally outputs the multi-scale detection results.

9. The method according to claim 8, characterized in that LiteRepLAN modules are as follows: The original feature map is input to the first-layer LiteRepCSP module. The module first performs preliminary processing on the input feature map through deep separable convolution DSC to extract compact features and reduce channel redundancy, and outputs feature map P1. Then, the feature map P1 is divided into two paths and sent to two RepBottleneck modules respectively. Each path performs deep feature extraction through multiple convolutional layers. The upper branch extracts features through multi-layer stacked RepBottleneck modules, and the number of channels in each layer remains unchanged to further capture deeper feature information and output feature map P2. The lower branch retains the original feature information through 1×1 deep separable convolution, and performs channel fusion through point-by-point convolution to output feature map P3. Then, the two feature maps P2 and P3 are fused through the Concat operation to obtain a feature map P4 with more multi-scale information; finally, the output feature map P4 enters the next layer of deep separable convolution for fine-grained feature processing to generate the final feature map P5; this module effectively realizes efficient feature extraction and cross-channel information fusion by combining the RepBottleneck module with deep separable convolution.

10. The method according to claim 8, characterized in that The detection head module is the Enhanced_DDetect module, which is specifically: The output feature map P1 is divided into two paths, entering the cv2 and cv3 branches for deeper feature extraction. In the cv2 branch, the input feature map undergoes two depth-wise separable convolutions to extract spatial features. Then, an additional convolution layer is used to further fuse channel information to generate the feature map P2. Unlike the cv2 branch, the cv3 branch adds a channel attention mechanism on the basis of two depth-wise separable convolutions. The channel attention mechanism captures the global information of each channel through adaptive average pooling, and dynamically adjusts the response of each channel using two layers of 1x1 convolution and Sigmoid activation function, thereby enhancing the model's sensitivity to key channel features. This mechanism helps to improve the accuracy of target classification; the processed feature map P3 is further refined by 1x1 convolution to obtain the final feature map P4; finally, the output feature maps P2 and P4 of the two branches cv2 and cv3 are merged through a splicing operation to form a joint feature map P5; this splicing operation ensures the effective fusion of the bounding box and classification task features, and after regression and classification processing, the output results of target detection are generated, including the bounding box regression BBox and the target category prediction Cls.

Citation Information

Patent Citations

  • Solar cell panel defect detection method and device based on improved YOLOv5

    CN118196045A

  • Solar cell panel defect detection method based on improved Yov7-tiny

    CN118505643A

  • Solar cell panel defect detection method based on lightweight reconstruction network

    CN116758033A

  • Metal surface defect detection method

    CN119251168A

  • Photovoltaic panel defect detection system and method based on unmanned aerial vehicle and machine vision

    CN119515799A

Cited By

  • Improved YOLOv9 lightweight steel surface defect detection model

    CN120747622A