Light-weight and high-precision photovoltaic cell electroluminescence image defect detection model
By constructing the RDS-YOLO detection model and introducing RECBConv, DySample, C2fCG and SCAMHead modules, the shortcomings of existing photovoltaic cell electroluminescent image defect detection technology in lightweight deployment and real-time detection are solved, and high-precision and efficient defect detection are achieved.
Patent Information
- Application Number
- CN202510326467.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-17
AI Technical Summary
Existing electroluminescent image defect detection technology for photovoltaic cells is difficult to achieve lightweight deployment and real-time detection in industrial production, and single-stage algorithms have missed and misdetection phenomena when dealing with complex backgrounds and tiny defects.
A lightweight and high-precision electroluminescent image defect detection model for photovoltaic cells is proposed. By constructing the RDS-YOLO detection model, RECBConv module, DySample module, C2fCG module and SCAMHead module are used to enhance feature extraction, multi-scale feature fusion and object detection accuracy.
The model's feature extraction capability, multi-scale feature fusion capability and object detection accuracy are significantly improved, the missed detection and error detection rates are reduced, and efficient and reliable electroluminescent image defect detection in photovoltaic cells is achieved.
Smart Images

Figure CN120163809A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision detection, and particularly relates to a lightweight and high-precision photovoltaic cell electroluminescence image defect detection model. Background Art
[0002] With the acceleration of the transformation of the global energy structure towards cleaner energy, the installed capacity of photovoltaic power generation has been increasing year by year. As the core carrier of energy conversion, internal defects in photovoltaic modules will reduce the photoelectric conversion efficiency, shorten the service life of photovoltaic panels, and may even cause fires in severe cases. Therefore, researching efficient and accurate defect detection technologies is of great significance for promoting the development of the photovoltaic industry.
[0003] Existing technologies mainly adopt defect detection methods based on deep learning, and their applications focus on two technical routes: two-stage detection algorithms represented by Faster R-CNN, and single-stage detection algorithms represented by the YOLO series. The multi-stage algorithm generates candidate defect regions through the Region Proposal Network (RPN), and then realizes defect localization and classification through ROIPooling and the classification network, with relatively high detection accuracy. However, the computational cost of the two-stage model is very high, making it difficult to meet the lightweight deployment and real-time detection requirements of industrial production; the single-stage algorithm directly outputs defect coordinates and categories on the feature map through end-to-end feature extraction and regression prediction, significantly improving the detection speed. However, due to the large differences in the sizes of photovoltaic defects and the complex background textures, it is difficult for traditional single-stage algorithms to effectively extract feature information, resulting in frequent occurrences of missed detections and false detections. Summary of the Invention
[0004] Aiming at the problem that it is difficult to distinguish between the fault current and inrush current of existing transformers and it is difficult to meet the existing actual application requirements, the present invention proposes a lightweight and high-precision photovoltaic cell electroluminescence image defect detection model to solve the above problems existing in the prior art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A lightweight and high-precision photovoltaic cell electroluminescence image defect detection model, which is trained through the following steps:
[0007] The electroluminescence defect images of photovoltaic cells are accurately labeled by labeling software and divided into a training set, a validation set, and a test set according to a ratio of 8:1:1 to construct a high-quality data set;
[0008] Construct the RDS-YOLO detection model. The RDS-YOLO detection model is constructed based on the YOLOv8 network and includes a Backbone module, a Neck module, and a Head module. In the Backbone module, except for the first two Conv modules, the remaining modules are replaced with the designed RECBConv module to enhance the feature extraction ability. The UpSample module in the Neck module is replaced with the DySample module to improve the upsampling efficiency. At the same time, the C2f module is replaced with the designed C2fCG module to enhance the multi-scale feature fusion ability. The Detect module in the Head module is replaced with the SCAMHead module to improve the object detection accuracy.
[0009] Use the sample data in the dataset to train the RDS-YOLO detection model to obtain a detection model that can accurately identify defects based on the electroluminescence image of the photovoltaic cell. The specific training process is as follows:
[0010] Uniformly adjust the electroluminescence image of the photovoltaic cell to a size of 640×640×3 in resolution to ensure input consistency.
[0011] Use two Conv layers, the RECBConv module, the C2f module, and the SPPF module in the Backbone module to extract and aggregate multi-level feature information.
[0012] Use the multi-level feature information extracted by the Backbone module as the input of the Neck module, and perform efficient feature fusion through the DySample module and the C2fCG module.
[0013] Use the Head module to input the three-scale features output by the Neck module into the SCAMHead module to complete the object detection task.
[0014] As a further solution of the present invention: The Backbone module specifically is that the input image first undergoes preliminary feature extraction through two Conv layers, and then sequentially passes through the RECBConv module and the C2f module for deep feature extraction and optimization. Finally, the features are passed to the SPPF module to further extract and aggregate multi-scale feature information.
[0015] As a further solution of the present invention: The Neck module specifically is that the multi-level features extracted by the Backbone module are used as the input. First, efficient feature upsampling is performed through the DySample module, and then multi-scale feature fusion is performed through the C2fCG module to enhance the feature expression ability.
[0016] As a further solution of the present invention: the Head module specifically inputs the three-scale features output by the Neck module into the SCAMHead module to complete the object detection task and achieve accurate identification and positioning of defects.
[0017] As a further solution of the present invention: the RECBConv module is specifically a convolutional module improved on the basis of the RFCBAMConv module. By replacing the channel attention part in CBAM with the ECA mechanism and combining the recursive feature calibration RFC and the spatial attention mechanism, efficient feature extraction and optimization are achieved. Specifically, the input features first undergo recursive calibration through the RFC module to extract richer context information, then the ECA module adaptively learns the dependencies between channels and assigns weights to each channel, and finally the important regions in the feature map are highlighted through the spatial attention mechanism to output the optimized feature map.
[0018] As a further solution of the present invention: the DySample module is specifically an ultra-lightweight and efficient dynamic upsampler. By bypassing the dynamic convolution and redefining the upsampling process from the perspective of point sampling, the computational burden is significantly reduced. The DySample module first converts the input features into a continuous feature map through bilinear interpolation, then generates content-aware point offsets through linear projection to control the positions of the sampling points. In the sampling point adjustment stage, the DySample changes the initial sampling position from nearest neighbor initialization to bilinear initialization, so that when the offset is zero, the upsampling result is equivalent to bilinear interpolation; a dynamic range factor is introduced to limit the movement range of the sampling points and avoid overlap; in addition, by dividing the feature map into multiple groups along the channel dimension and sharing the same sampling set for each group, the flexibility of upsampling is further increased; finally, the grid_sample function in PyTorch is used to resample the continuous feature map according to the generated sampling points to obtain an efficient and flexible upsampling result.
[0019] As a further solution of the present invention: The C2fCG module is specifically an improved version that introduces the Context Guided Block into the C2f structure of YOLOv8. By integrating the Context Guided Block into C2f, the model's ability to capture local features, surrounding context, and global context information is further enhanced. The C2fCG module first adjusts the number of channels of the input through a 1x1 convolutional layer to reduce the computational amount and extract preliminary features. Subsequently, the adjusted features are fed into the Context Guided Block. This module captures detailed information through a 3x3 convolutional layer of the local feature extractor, and at the same time uses a 3x3 dilated convolutional layer of the surrounding context extractor to expand the receptive field and obtain the context information around the target. The local features and the surrounding context features are concatenated and fused through batch normalization BN and parametric ReLU to form joint features. Then, the global context extractor extracts global context information through global average pooling and a multi-layer perceptron MLP, and weights the joint features to enhance useful features and suppress irrelevant features. Finally, if the number of input and output channels is the same, the module adds the input features to the output of the Context Guided Block through a residual connection to further improve the feature expression ability, and finally outputs a feature map with enhanced context.
[0020] As a further solution of the present invention: The SCAMHead module is specifically an SCAM module added before the original Head module. The SCAM module enhances the ability to distinguish between targets and backgrounds in the feature map through global pooling operations and spatial-channel interactions, and then inputs it into the classification head. First, the SCAMHead module integrates global information through the first branch using global average pooling GAP and global max pooling GMP, guiding the feature map to select channels with significant information and enhancing the context learning ability in the channel dimension. Then, the second branch performs a linear transformation on the input feature map through a 1×1 convolution to generate values, providing a linear mapping for subsequent context information interaction. Next, the third branch also uses a 1×1 convolution to simplify the product operation of queries and keys, reducing the computational complexity while retaining the global context information. After that, the outputs of the first branch and the third branch are respectively matrix-multiplied with the output of the second branch to obtain two new branches, representing cross-channel and cross-space context information respectively. Finally, the outputs of these two branches are fused through a broadcast Hadamard product, effectively enhancing the representation ability of the global context information. Finally, the obtained output is input into the Head module to complete the target detection task.
[0021] Compared with the prior art, the beneficial effects of the present invention are:
[0022] By introducing the RECBConv module, DySample module, C2fCG module, and SCAMHead module, the present invention significantly improves the model's feature extraction ability, multi-scale feature fusion ability, and object detection accuracy. The RECBConv module effectively enhances the model's sensitivity to minute defects through recursive feature calibration and spatial attention mechanisms. The C2fCG module can better distinguish complex backgrounds from real defects through context-guided blocks, reducing false detections. The SCAMHead module further improves the detection accuracy and reduces the missed detection rate through global pooling and spatial-channel interaction mechanisms. In addition, the lightweight design of the present invention significantly reduces the computational complexity while ensuring high precision, providing an efficient and reliable solution for the defect detection of electroluminescence images of photovoltaic cells. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following briefly introduces the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings. Among them:
[0024] Figure 1 It is a flowchart for training a defect detection model for electroluminescence images of photovoltaic cells according to an embodiment of the present invention;
[0025] Figure 2 It is a schematic structural diagram of constructing an RDS-YOLO detection model according to an embodiment of the present invention;
[0026] Figure 3 It is a schematic structural diagram of the RECBConv module according to an embodiment of the present invention;
[0027] Figure 4 It is a visualization result display diagram when the defect detection model trained by the present invention detects an electroluminescence image of a photovoltaic cell to be tested. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0029] Embodiment
[0030] Please refer to Figure 1As shown, the present invention provides a lightweight and high-precision photovoltaic cell electroluminescence image defect detection model, which is trained through the following steps:
[0031] The photovoltaic cell electroluminescence defect images are accurately labeled by a labeling software and divided into a training set, a validation set, and a test set according to a ratio of 8:1:1 to construct a high-quality data set.
[0032] As Figure 2 shown, an RDS-YOLO detection model is constructed. The RDS-YOLO detection model is based on the YOLOv8 network and includes a Backbone module, a Neck module, and a Head module. In the Backbone module, except for the first two Conv modules, the remaining modules are replaced with the designed RECBConv module to enhance the feature extraction ability; the UpSample module in the Neck module is replaced with the DySample module to improve the upsampling efficiency; at the same time, the C2f module is replaced with the designed C2fCG module to enhance the multi-scale feature fusion ability; the Detect module in the Head module is replaced with the SCAMHead module to improve the object detection accuracy.
[0033] The RDS-YOLO detection model is trained using the sample data in the data set to obtain a detection model that can accurately identify defects based on the photovoltaic cell electroluminescence image. The specific training process is as follows:
[0034] The photovoltaic cell electroluminescence images are uniformly adjusted to a size of 640×640×3 in resolution to ensure input consistency.
[0035] Using the two Conv layers, RECBConv module, C2f module, and SPPF module in the Backbone module, multi-level feature information is extracted and aggregated.
[0036] The multi-level feature information extracted by the Backbone module is used as the input of the Neck module, and efficient feature fusion is performed through the DySample module and the C2fCG module.
[0037] Using the Head module, the three-scale features output by the Neck module are input into the SCAMHead module to complete the object detection task.
[0038] Furthermore, the Neck module specifically uses the multi-level features extracted by the Backbone module as the input. First, efficient feature upsampling is performed through the DySample module, and then multi-scale feature fusion is performed through the C2fCG module to enhance the feature expression ability.
[0039] Furthermore, the Neck module specifically takes the multi-level features extracted by the Backbone module as input. First, it performs efficient feature upsampling through the DySample module, and then conducts multi-scale feature fusion through the C2fCG module to enhance the feature expression ability.
[0040] Furthermore, the Head module specifically inputs the three-scale features output by the Neck module into the SCAMHead module to complete the object detection task and achieve precise identification and localization of defects.
[0041] Furthermore, as Figure 3 shown, the RECBConv module is a convolutional module improved on the basis of the RFCBAMConv module. By replacing the channel attention part in CBAM with the ECA mechanism and combining the recursive feature calibration RFC and the spatial attention mechanism, it realizes efficient feature extraction and optimization. Specifically, the input feature first undergoes recursive calibration through the RFC module to extract richer context information, then adaptively learns the channel-to-channel dependencies through the ECA module to assign weights to each channel, and finally highlights the important regions in the feature map through the spatial attention mechanism to output the optimized feature map.
[0042] Furthermore, the DySample module is a super lightweight and efficient dynamic upsampler. By bypassing dynamic convolution and redefining the upsampling process from the perspective of point sampling, it significantly reduces the computational burden. The DySample module first converts the input feature into a continuous feature map through bilinear interpolation, then generates content-aware point offsets through linear projection to control the positions of the sampling points. In the sampling point adjustment stage, DySample changes the initial sampling position from nearest neighbor initialization to bilinear initialization, so that when the offset is zero, the upsampling result is equivalent to bilinear interpolation; a dynamic range factor is introduced to limit the movement range of the sampling points and avoid overlap; in addition, by dividing the feature map into multiple groups along the channel dimension and sharing the same sampling set for each group, the flexibility of upsampling is further increased; finally, the grid_sample function in PyTorch is used to resample the continuous feature map according to the generated sampling points to obtain an efficient and flexible upsampling result.
[0043] Furthermore, the C2fCG module is specifically an improved version that introduces the Context Guided Block into the C2f structure of YOLOv8. By integrating the Context Guided Block into C2f, the model's ability to capture local features, surrounding context, and global context information is further enhanced. The C2fCG module first adjusts the number of channels of the input through a 1x1 convolutional layer to reduce the computational load and extract preliminary features. Subsequently, the adjusted features are fed into the Context Guided Block, which captures detailed information through a 3x3 convolutional layer of the local feature extractor and expands the receptive field using a 3x3 dilated convolutional layer of the surrounding context extractor to obtain the context information around the target. The local features and the surrounding context features are concatenated and fused through batch normalization BN and parametric ReLU to form joint features. Then, the global context extractor extracts global context information through global average pooling and a multi-layer perceptron MLP and weights the joint features to enhance useful features and suppress irrelevant features. Finally, if the number of input and output channels is the same, the module adds the input features to the output of the Context Guided Block through a residual connection to further improve the feature expression ability, and finally outputs a context-enhanced feature map.
[0044] Furthermore, the SCAMHead module is specifically a module that adds the SCAM module in front of the original Head module. The SCAM module enhances the discrimination ability between the target and the background in the feature map through global pooling operations and spatial-channel interactions, and then inputs it into the classification head. First, the SCAMHead module integrates global information through the first branch using global average pooling GAP and global max pooling GMP to guide the feature map to select channels with significant information and enhance the context learning ability in the channel dimension. Then, the second branch performs a linear transformation on the input feature map through a 1×1 convolution to generate values, providing a linear mapping for subsequent context information interaction. Next, the third branch also uses a 1×1 convolution to simplify the product operation of the query and the key, reducing the computational complexity while retaining the global context information. After that, the outputs of the first branch and the third branch are respectively matrix-multiplied with the output of the second branch to obtain two new branches, representing cross-channel and cross-space context information respectively. Finally, the outputs of these two branches are fused through a broadcast Hadamard product, effectively enhancing the representation ability of the global context information. Finally, the obtained output is input into the Head module to complete the object detection task.
[0045] In summary, in the embodiments of the present invention, first, the RECBConv module is introduced into the Backbone module, which significantly enhances the feature extraction ability by combining the ECA mechanism, recursive feature calibration, and spatial attention mechanism; second, the DySample module and the C2FCG module are adopted in the Neck module. The former improves the efficiency through dynamic upsampling, and the latter enhances the multi-scale feature fusion ability through the context-guided block; third, the SCAMHead module is introduced into the Head module, which improves the accuracy and robustness of object detection through global pooling and the spatial-channel interaction mechanism.
[0046] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A lightweight and high-precision photovoltaic cell electroluminescent image defect detection model, characterized in that: The detection model is trained through the following steps: The electroluminescent defect images of photovoltaic cells are accurately annotated by annotation software and divided into training set, validation set and test set in a ratio of 8:1:1 to construct a high-quality data set; Construct an RDS-YOLO detection model, which is based on a YOLOv8 network and includes a Backbone module, a Neck module, and a Head module. In the Backbone module, except for the first two Conv modules, the remaining modules are replaced with the designed RECBConv modules to enhance feature extraction capability; The UpSample module in the Neck module is replaced by the DySample module to improve the upsampling efficiency; at the same time, the C2f module is replaced by the designed C2fCG module to enhance the multi-scale feature fusion capability; the Detect module in the Head module is replaced by the SCAMHead module to improve the target detection accuracy; The RDS-YOLO detection model is trained using the sample data in the data set to obtain a detection model that can accurately identify defects based on the electroluminescent image of the photovoltaic cell. The specific training process is as follows: The photovoltaic cell electroluminescence images are uniformly resized to a resolution of 640×640×3 to ensure input consistency; Utilize the two Conv layers, RECBConv module, C2f module and SPPF module in the Backbone module to extract and aggregate multi-level feature information; The multi-level feature information extracted by the Backbone module is used as the input of the Neck module, and efficient feature fusion is performed through the DySample module and the C2fCG module; Using the Head module, the three scale features output by the Neck module are input into the SCAMHead module to complete the target detection task.
2. According to claim 1, a lightweight and high-precision photovoltaic cell electroluminescent image defect detection model is characterized by: Specifically, the input image first passes through two Conv layers for preliminary feature extraction, and then passes through the RECBConv module and the C2f module in sequence for deep feature extraction and optimization. Finally, the features are passed to the SPPF module to further extract and aggregate multi-scale feature information.
3. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 2 is characterized by: Specifically, the Neck module takes the multi-level features extracted by the Backbone module as input, first performs efficient feature upsampling through the DySample module, and then performs multi-scale feature fusion through the C2fCG module to enhance the expressiveness of the features.
4. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 3 is characterized by: The Head module specifically inputs the three scale features output by the Neck module into the SCAMHead module to complete the target detection task and realize the accurate identification and positioning of defects.
5. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 2 is characterized by: The RECBConv module is specifically a convolution module improved on the basis of the RFCBAMConv module. By replacing the channel attention part in CBAM with the ECA mechanism, combined with the recursive feature calibration RFC and the spatial attention mechanism, efficient feature extraction and optimization are achieved. Specifically, the input features are first recursively calibrated by the RFC module to extract richer contextual information, and then the dependencies between channels are adaptively learned through the ECA module, and weights are assigned to each channel. Finally, the important areas in the feature map are highlighted through the spatial attention mechanism, and the optimized feature map is output.
6. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 3 is characterized by: The DySample module is specifically an ultra-lightweight and efficient dynamic upsampler. By bypassing dynamic convolution and redefining the upsampling process from the perspective of point sampling, the computational burden is significantly reduced. The DySample module first converts the input features into continuous feature maps through bilinear interpolation, and then generates content-aware point offsets through linear projection to control the positions of sampling points. In the sampling point adjustment stage, DySample changes the initial sampling position from nearest neighbor initialization to bilinear initialization, so that when the offset is zero, the upsampling result is equivalent to bilinear interpolation; a dynamic range factor is introduced to limit the movement range of the sampling points to avoid overlap; in addition, by dividing the feature map into multiple groups along the channel dimension, each group shares the same sampling set, which further increases the flexibility of upsampling; finally, the grid_sample function in PyTorch is used to resample the continuous feature map according to the generated sampling points to obtain efficient and flexible upsampling results.
7. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 3 is characterized by: The C2fCG module is specifically an improved version of the Context Guided Block introduced into the C2f structure of YOLOv8. By integrating the Context Guided Block into C2f, the model's ability to capture local features, surrounding context and global context information is further enhanced. The C2fCG module first adjusts the number of channels of the input through a 1x1 convolution layer to reduce the amount of calculation and extract preliminary features. Subsequently, the adjusted features are sent to the Context GuidedBlock. The module captures detail information through a local feature extractor 3x3 convolution, and uses a surrounding context extractor 3x3 hole convolution to expand the receptive field and obtain context information around the target. The local features and surrounding context features are concatenated and fused through batch normalization BN and parameterized ReLU to form a joint feature. Then, the global context extractor extracts global context information through global average pooling and multi-layer perceptron MLP, and weights the joint features to enhance useful features and suppress irrelevant features. Finally, if the number of input and output channels is the same, the module connects the input features with the Context Guided through a residual connection. The outputs of the blocks are added to further improve the feature expression capability, and finally a context-enhanced feature map is output.
8. The lightweight and high-precision photovoltaic cell electroluminescent image defect detection model according to claim 4 is characterized by: The SCAMHead module is specifically a SCAM module added before the original Head module. The SCAM module enhances the ability to distinguish between targets and backgrounds in the feature map through global pooling operations and space-channel interactions, and then inputs it into the classification head. First, the SCAMHead module integrates global information through the first branch using global average pooling GAP and global maximum pooling GMP to guide the feature map to select channels with significant information and enhance the context learning ability of the channel dimension. Then, the second branch linearly transforms the input feature map through 1×1 convolution to generate values to provide linear mapping for subsequent context information interaction. Then, the third branch also uses 1×1 convolution to simplify the query and key product operations, reducing the computational complexity while retaining the global context information. After that, the outputs of the first branch and the third branch are matrix multiplied with the output of the second branch respectively to obtain two new branches, which represent cross-channel and cross-space context information respectively. Finally, the outputs of the two branches are fused through broadcast Hadamard products, thereby effectively enhancing the representation ability of global context information. Finally, the obtained output is input into the Head module to complete the target detection task.
Citation Information
Cited By
Intelligent surface defect analysis method and system for full-period coverage of vehicle part production
CN120931649A