Photovoltaic EL defect detection method based on YOLOv11-MSFF model
By introducing local feature extraction units, deformable attention mechanism and feature fusion units in the YOLOv11-MSFF model, the problem of insufficient local feature extraction and insufficient fusion of scale features in photovoltaic module detection is solved, and defect detection with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510437974.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing single-stage object detection algorithm based on deep learning has problems such as insufficient local feature extraction, lack of effective attention mechanism and insufficient fusion of features at different scales in photovoltaic module defect detection, resulting in insufficient detection accuracy and robustness.
The YOLOv11-MSFF model is adopted, and the feature extraction and fusion mechanism is optimized by introducing the C3k2 module, deformable attention mechanism and five feature fusion units in the local feature extraction unit, the feature extraction and fusion mechanism is improved, and the local feature capture capability and defect area positioning accuracy are achieved to achieve effective fusion of multi-scale features.
It significantly improves the accuracy and reliability of surface defect detection of photovoltaic modules, can more accurately identify small defects, reduce background noise interference, and improve the detection speed and generalization ability of the model.
Smart Images

Figure CN120298378A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of photovoltaic cell defect detection, and specifically is a photovoltaic EL defect detection method based on the YOLOv11-MSFF model. Background Art
[0002] In the photovoltaic industry, the quality of photovoltaic modules directly determines their photoelectric conversion efficiency and service life. However, defects on the surface of the modules, such as cracks, hot spots, and hidden cracks, may not only become potential performance hazards, but may also affect the stability and safety of the entire photovoltaic system. Therefore, high-precision surface defect detection of photovoltaic modules is an important part of ensuring the healthy development of the photovoltaic industry. However, in actual industrial scenarios, traditional target detection methods rely on manually designed features, which are inefficient and subjective, and cannot meet the strict requirements of modern industry for detection efficiency and accuracy.
[0003] In contrast, deep learning-based object detection algorithms can automatically learn complex image features, thereby significantly reducing the complexity of manually designed features and improving the versatility and scalability of the model. With these advantages, deep learning-based object detection algorithms have become the mainstream choice in the current field. At present, deep learning-based object detection algorithms are mainly divided into two categories: two-stage object detection algorithms and single-stage object detection algorithms. Two-stage object detection algorithms, such as R-CNN and its variants (Fast R-CNN, Faster R-CNN, etc.), are usually divided into two stages: first, candidate regions are generated, and then these candidate regions are classified and bounding box regressed. Although it has significant advantages in detection accuracy, the model is complex and the detection speed is slow, which makes it difficult to meet the needs of real-time detection. In contrast, single-stage object detection algorithms, such as the YOLO series and SSD, directly perform detection on the input image without generating candidate regions. These algorithms have the advantages of fast detection speed and simple model structure, and can meet the needs of real-time detection. However, the single-stage object detection algorithm currently used to detect surface defects in EL images has the advantage of fast detection, but there are many problems that need to be solved. These problems are mainly reflected in the following aspects: First, the local feature extraction is not fine enough, and it is difficult to accurately capture the subtle features of the defects; second, there is a lack of effective attention mechanism, which makes it impossible to highlight the defect area, resulting in susceptibility to background noise interference during detection; third, the fusion of features of different scales is insufficient, and the fusion of high-level semantic information and underlying position information is insufficient, which limits the detection performance of the model when dealing with multi-scale targets. Summary of the invention
[0004] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a photovoltaic EL defect detection method based on the YOLOv11-MSFF model, in order to significantly improve the detection accuracy of surface defects of photovoltaic modules by introducing an effective attention mechanism, optimizing the feature extraction and fusion mechanism, and making full use of multi-scale feature information.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A photovoltaic EL defect detection method based on the YOLOv11-MSFF model of the present invention is characterized by including the following steps:
[0007] Step 1: Use a high-resolution CCD camera to capture images of the photovoltaic module under near-infrared light and form an EL image set. Let any EL image in the EL image set be denoted as I el ; and the size of I el is M×N, where M represents the length of the EL image and N represents the width of the EL image;
[0008] Label the positions and categories of s defects in the EL image I el to obtain s defect position labels and s defect category labels; let the position label of the i-th defect in I el be denoted as and = ([[]] , [[[]] , [[[]] ), where represents the center point of the true bounding box of the i-th defect in I el , is the width of the true bounding box of the i-th defect in I el , is the height of the true bounding box of the i-th defect in I el . Denote the category label of the i-th defect in I el as ;
[0009] Step 2: Construct a YOLOv11-MSFF model, which sequentially includes: a feature extraction network, a feature fusion network, and a prediction network, and process I el to obtain the category prediction label el of the I el image and the position prediction label el of the I n+1 image;
[0010] Step 3: Construct a total loss function Loss;
[0011] Step 4: Iteratively train the YOLOv11-MSFF model using the SGD optimizer, and adjust the model parameters according to the total loss function Loss until the total loss function Loss converges, thereby obtaining a trained EL image defect detection model for detecting defects in EL images.
[0012] Another feature of the photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to the present invention is that the step 2 includes:
[0013] Step 2.1: The feature extraction network consists of an initial convolutional block Conv0, n local feature extraction units, a fast spatial pyramid pooling module, a C2PSA module, and a deformable attention module in sequence, and processes I el to generate a set of coarse features { | = 1, 2, …, } by the n local feature extraction units, and outputs an attention feature I n+1 × N / 2 n+1 of size M / 2 by the deformable attention module; where DAttetion ; among them, represents the u-th coarse feature output by the u-th local feature extraction unit, and is of size M / 2 u+1 × N / 2 u+1 ;
[0014] Step 2.2: The feature fusion network includes five feature fusion units, and processes I DAttetion , the (n - 1)-th coarse feature and the (n - 2)-th coarse feature to obtain a third fusion feature F fused,3 , a fourth fusion feature F fused,4 and a fifth fusion feature F fused,5 ;
[0015] Step 2.3: The prediction network consists of two parallel branches, and processes F fused,3 , F fused,4 and F fused,5 to obtain the class prediction label el of the I image and the position prediction label el of the I image.
[0016] Furthermore, step 2.1 includes the following steps:
[0017] Step 2.1.1: The initial convolutional block Conv0 processes I elProcess to obtain the initial convolutional feature I with dimensions reduced to M / 2 × N / 2 conv0 ;
[0018] Step 2.1.2: The initial convolutional feature I conv0 After being processed by n local feature extraction units in sequence, a set of coarse features { | = 1, 2, …, } is correspondingly generated, where represents the u-th coarse feature output by the u-th local feature extraction unit, and has dimensions M / 2 u+1 × N / 2 u+1 ;
[0019] Step 2.1.3: Input the n-th coarse feature output by the n-th local feature extraction unit into the fast spatial pyramid pooling module for preliminary enhancement processing to obtain the preliminary enhanced feature I SPPF ;
[0020] Step 2.1.4: Input the preliminary enhanced feature I SPPF into the C2PSA module for further enhancement to obtain the enhanced and enhanced feature I C2PSA ;
[0021] Step 2.1.5: Input the enhanced and enhanced feature I C2PSA feature into the deformable attention module for processing to obtain the attention feature I with dimensions M / 2 n+1 × N / 2 n+1 DAttetion .
[0022] Furthermore, in Step 2.2:
[0023] Step 2.2.1: The first feature fusion unit consists of a first convolutional module Conv1, a second convolutional module Conv2, a first splicing module, and a first CSPStage module in sequence, and processes I DAttetion and the (n - 1)-th coarse feature to obtain the first fusion feature F fused,1 ;
[0024] After processing the attention feature I DAttetion through the first convolutional block Conv1, a first convolutional feature I with dimensions M / 2 n+1 × N / 2 n +1 is obtained conv,1 ;
[0025] The (n - 1)-th coarse feature output by the (n - 1)-th local feature extraction unit After being processed in the second convolutional block Conv2, the second convolutional feature I with a size of M / 2 n+1 × N / 2 n+1 is obtained; conv,2 ;
[0026] The first convolutional feature I conv,1 and the second convolutional feature I conv,2 are concatenated through the first concatenation module, and the first concatenated feature I with a size of M / 2 n+1 × N / 2 n+1 is obtained; concat,1 ;
[0027] The first concatenated feature I concat,1 is processed through the first CSPStage module, and the first fused feature F with a size of M / 2 n+1 × N / 2 n+1 is obtained; fused,1 ;
[0028] Step 2.2.2: The second feature fusion unit consists of a first upsampling module, a third convolutional module Conv3, a second concatenation module, and a second CSPStage module, and processes F fused,1 , the (n - 2)-th coarse feature and the (n - 1)-th coarse feature to obtain the second fused feature F with a size of M / 2 n × N / 2 n ; fused,2 ;
[0029] The first fused feature F fused,1 is input into the first upsampling module to obtain the first upsampled feature I with a size of M / 2 n × N / 2 n ; unsample,1 ;
[0030] The (n - 2)-th coarse feature output by the (n - 2)-th local feature extraction unit is input into the third convolutional module Conv3, and the third convolutional feature I with a size of M / 2 n × N / 2 n is obtained; conv,3 ;
[0031] The first upsampled feature I unsample,1 with the same size, the third convolutional feature I conv,3 and the (n - 1)-th coarse feature are jointly input into the second concatenation module for processing, and the second concatenated feature I with a size of M / 2 n × N / 2 n is obtained.concat,2 ;
[0032] The second splicing feature I concat,2 Through the processing of the second CSPStage module, the second fused feature F with a size of M / 2 n × N / 2 n is obtained; fused,2 ;
[0033] Step 2.2.3: The third feature fusion unit consists of a second upsampling module, a third splicing module, and a third CSPStage module, and processes F fused,2 and the (n - 2)-th rough feature to obtain the third fused feature F with a size of M / 2 n-1 × N / 2 n-1 ; fused,3 ;
[0034] The second fused feature F fused,2 is input into the second upsampling module for processing, and the second upsampled feature I with a size of M / 2 n-1 ×N / 2 n-1 is obtained; unsample,2 ;
[0035] The second upsampled feature I unsample,2 and the (n - 2)-th rough feature are spliced through the feature splicing of the third splicing module, and the third splicing feature I with a size of M / 2 n-1 × N / 2 n-1 is obtained; concat,3 ;
[0036] The third splicing feature I concat,3 is processed through the third CSPStage module to obtain the third fused feature F with a size of M / 2 n-1 × N / 2 n-1 ; fused,3 ;
[0037] Step 2.2.4: The fourth feature fusion unit consists of a fourth convolution module Conv4, a fourth splicing module, and a fourth CSPStage module; and processes F fused,3 and F fused,2 to obtain the fourth fused feature F with a size of M / 2 n × N / 2 n ; fused,4 ;
[0038] The third fused feature F fused,3 is input into the fourth convolution module Conv4 for processing, and the fourth convolution feature I with a size of M / 2 n ×N / 2 n is obtainedconv,4 ;
[0039] The fourth convolutional feature I conv,4 and the second fusion feature F fused,2 After feature concatenation through the fourth concatenation module, the fourth concatenated feature I with a size of M / 2 n × N / 2 n is obtained; concat,4 ;
[0040] The fourth concatenated feature I concat,4 After being processed by the fourth CSPStage module, the fourth fusion feature F with a size of M / 2 n × N / 2 n is obtained; fused,4 ;
[0041] Step 2.2.5: The fifth feature fusion unit consists of a fifth convolutional module Conv5, a sixth convolutional module Conv6, a fifth concatenation module, and a fifth CSPStage module, and processes F fused,2 and F fused,4 as well as F fused,1 to obtain the fifth fusion feature F with a size of M / 2 n+1 ×N / 2 n+1 ; fused,5 ;
[0042] The second fusion feature F fused,2 is input into the fifth convolutional module Conv5 for processing to obtain the fifth convolutional feature I conv,5 ;
[0043] The fourth fusion feature F fused,4 is input into the sixth convolutional module Conv6 for processing to obtain the sixth convolutional feature I conv,6 ;
[0044] The fifth convolutional feature I conv,5 , the sixth convolutional feature I conv,6 and the first fusion feature F fused,1 are input into the fifth concatenation module for processing to obtain the fifth concatenated feature I concat,5 ;
[0045] The fifth concatenated feature I concat,5 After being processed by the fifth CSPStage module, the fifth fusion feature F with a size of M / 2 n+1 ×N / 2 n+1 is obtained; fused,5 .
[0046] Furthermore, the said step 2.3 includes:
[0047] Step 2.3.1: The first branch consists of three parallel category prediction units, which process F fused,3 , F fused,4 , and F fused,5 respectively to obtain the category prediction label el of the I image;
[0048] Among them, the f-th category prediction unit consists of the f-th first convolutional module Conv f,1 , the f-th second convolutional module Conv f,2 , and the f-th first two-dimensional convolutional module Conv2d f,1 in sequence; f = 1, 2, 3;
[0049] Input any (f + 2)-th fusion feature F fused,3 , F fused,4 , and F fused,5 in F fused,f+2 into the f-th category prediction unit. First, after being processed by the first convolutional module Conv f,1 , the first convolutional feature I conv,f,1 output by the f-th category prediction unit is obtained. Then, I conv,f,1 is processed by the second convolutional module Conv f,2 to obtain the second convolutional feature I conv,f,2 output by the f-th category prediction unit. Finally, I conv,f,2 is processed by the first two-dimensional convolutional module Conv2d f,1 to obtain the first two-dimensional convolutional feature I conv2d,f,1 output by the f-th category prediction unit;
[0050] Step 2.3.2: The second branch consists of three parallel position prediction units, which process F fused,3 , F fused,4 , and F fused,5 respectively to obtain the position prediction label el of the I image;
[0051] Among them, the f-th position prediction unit consists of the f-th first depthwise separable convolutional module DWConv f,1 , the f-th third convolutional module Conv f,3 , the f-th second separable convolutional module DWConv f,2 , the f-th fourth convolutional module Conv f,4 , and the f-th second two-dimensional convolutional module Conv2d f,2 in sequence;
[0052] Input F fused,f+2Input into the f-th position prediction unit, first through the processing of the f-th first depthwise separable convolution module DWConv f,1 to obtain the first depthwise separable convolution feature I output by the f-th position prediction unit dwconv,f,1 , secondly, I dwconv,f,1 after being processed by the f-th third convolution module Conv f,3 to obtain the third convolution feature I output by the f-th position prediction unit conv,f,3 , then, I conv,f,3 after being processed by the f-th second separable convolution module DWConv f,2 to obtain the second depthwise separable convolution feature I output by the f-th position prediction unit dwconv,f,2 , then, I dwconv,f,2 after being processed by the f-th fourth convolution module Conv f,4 to obtain the fourth convolution feature I output by the f-th class prediction unit conv,f,4 , finally, I conv,f,4 after being processed by the f-th second two-dimensional convolution module Conv2d f,2 to obtain the second two-dimensional convolution feature I output by the f-th position prediction unit conv2d,f,2 ;
[0053] Step 2.3.3: Concatenate I conv2d,f,1 and I conv2d,f,2 along the channel dimension respectively to obtain the f-th predicted concatenated feature set I f ×H f ×W f ={ I conv2d,f , I conv2d,1 , I conv2d,2}, where I conv2d,3 represents the first predicted concatenated feature, I conv2d,1 represents the second predicted concatenated feature, I conv2d,2 represents the third predicted concatenated feature, A conv2d,3 represents the number of channels of the f-th predicted concatenated feature, H f represents the height of the f-th predicted concatenated feature, W f represents the width of the f-th predicted concatenated feature; f
[0054] Step 2.3.4: After reshaping and concatenating I conv2d,1 , I conv2d,2 , I conv2d,3 , obtain the position prediction label el of the I image and the class prediction label , and let the position prediction label of the i-th defect in be denoted as =( , , ), where represents the center point of the predicted bounding box at the location of the i-th defect in I el , is the width of the predicted bounding box at the location of the i-th defect in I el , is the height of the predicted bounding box at the location of the i-th defect in I el ; Let represent the predicted category of the i-th defect in
[0055] Furthermore, step 3 includes:
[0056] Step 3.1: Construct the loss function of the i-th defect using Equation (1) :
[0057] (1)
[0058] In Equation (1), and are two weights; represents the category loss of the i-th defect, represents the bounding box loss of the i-th defect, and there is:
[0059] (2)
[0060] (3)
[0061] In Equation (3), represents the Euclidean distance, represents the diagonal length of the minimum bounding rectangle of the i-th defect prediction box and the i-th defect ground truth box, represents the overlap degree between the i-th defect prediction box and the i-th defect ground truth box, is the aspect ratio similarity between the i-th defect prediction box and the i-th defect ground truth box, is the weight function of the i-th defect, and there is:
[0062] (4)
[0063] (5)
[0064] Step 3.2: Calculate the total loss function Loss using Equation (6):
[0065] (6).
[0066] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program that supports the processor to execute any one of the photovoltaic EL defect detection methods in claims 1-6, and the processor is configured to execute the program stored in the memory.
[0067] A computer-readable storage medium according to the present invention, characterized in that when the computer program stored on the computer-readable storage medium is run by a processor, it executes the steps of the photovoltaic EL defect detection method.
[0068] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0069] 1. Significantly improve the local feature extraction ability: The present invention ingeniously integrates a local feature extraction module (LFE) into the C3k2 module in the local feature extraction unit. The LFE module enhances feature interaction through shift operations by introducing a shift convolution module, improving the local feature extraction effect. This improvement enables the model to capture local detail features in the EL image more sensitively, thereby more accurately identifying tiny defects, effectively solving the problem of insufficiently fine local feature extraction in the prior art.
[0070] 2. Accurately locate the defect area: The present invention introduces a deformable attention mechanism into the backbone network. This mechanism can adaptively adjust the focus area of attention according to the image content, thereby more accurately locating the position of the defect and improving the accuracy of defect detection. Through this attention mechanism, the model can highlight the defect area, reduce the interference of background noise, effectively solve the problem of the lack of an effective attention mechanism in the prior art, and further improve the reliability and accuracy of detection.
[0071] 3. Optimize multi-scale feature fusion: The feature fusion network of the present invention realizes efficient multi-scale feature fusion through the cascaded processing of five feature fusion units. This improvement enables different-scale features to be more effectively fused, thereby improving the model's detection ability for defects of different sizes, effectively solving the problem of insufficient fusion of different-scale features in the prior art, and significantly enhancing the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a structural diagram of the C2k2-LFE module in the local feature extraction unit of the present invention;
[0073] Figure 2a It is a detection effect diagram of the model of the present invention using the local feature extraction module;
[0074] Figure 2b It is a detection effect diagram of the model of the present invention without using the local feature extraction module;
[0075] Figure 3a is an image in the EL image dataset of the present invention;
[0076] Figure 3b is the heatmap when the deformable attention mechanism is used in the model of the present invention;
[0077] Figure 3c is the heatmap when the deformable attention mechanism is not used in the model of the present invention;
[0078] Figure 4a is the detection effect diagram when five feature fusion units are used in the model of the present invention;
[0079] Figure 4b is the detection effect diagram when five feature fusion units are not used in the model of the present invention. Detailed implementation manners
[0080] In this embodiment, a photovoltaic EL defect detection method based on the YOLOv11-MSFF model includes the following steps:
[0081] Step 1: Use a high-resolution CCD camera to capture images of the photovoltaic module under near-infrared light and form an EL image set. Denote any EL image in the EL image set as I el ; and the size of I el is M×N, where M represents the length of the EL image and N represents the width of the EL image;
[0082] Label the positions and categories of s defects in the EL image I el to obtain the position labels and category labels of s defects. Denote the position label of the i-th defect in I el as , and =( , , ), where represents the center point of the true bounding box of the i-th defect at the position in I el , is the width of the true bounding box of the i-th defect at the position in I el , is the height of the true bounding box of the i-th defect at the position in I el . Denote the category label of the i-th defect in I el as ;
[0083] Step 2: Construct a YOLOv11-MSFF model, which sequentially includes: a feature extraction network, a feature fusion network, and a prediction network, and perform operations on I elProcess it to obtain I el Class prediction label of the image And I el Position prediction label g' of the image;
[0084] Step 2.1: The feature extraction network consists of an initial convolutional block Conv0, n local feature extraction units, a fast spatial pyramid pooling module, a C2PSA module, and a deformable attention module in sequence, and processes I el to generate a set of coarse features { | = 1, 2, …, } from the n local feature extraction units, and outputs attention feature I n+1 of size M / 2 n+1 × N / 2 DAttetion by the deformable attention module; where, represents the u-th coarse feature output by the u-th local feature extraction unit, and is of size M / 2 u+1 × N / 2 u+1 .
[0085] Step 2.1.1: The initial convolutional block Conv0 processes I el to obtain the initial convolutional feature I conv0 with the size reduced to M / 2 × N / 2;
[0086] Step 2.1.2: After the initial convolutional feature I conv0 is processed by n local feature extraction units in sequence, a set of coarse features { | = 1, 2, …, } is generated accordingly, where, represents the u-th coarse feature output by the u-th local feature extraction unit, and is of size M / 2 u+1 × N / 2 u+1 ; Any one of the local feature extraction units consists of a convolutional module Conv and a C3k2-LFE module, and the C3k2-LFE structure is as Figure 1 shown; A local feature extraction (LFE) module is introduced in the local feature extraction unit, which cleverly uses the shift operation to enhance the interaction between features, enabling the model to more sensitively capture the local detailed features in the EL image, thereby significantly improving the detection accuracy of the model for tiny defects. Figure 2a And Figure 2b show the relevant comparison effects: Figure 2a is the detection effect diagram after using the local feature extraction module, Figure 2bThis is the detection effect diagram when the module is not used.
[0087] Step 2.1.3: The nth coarse feature output by the nth local feature extraction unit Input to the fast spatial pyramid pooling module for preliminary enhancement processing to obtain the preliminary enhanced feature I SPPF ;
[0088] Step 2.1.4: Preliminary Enhancement of Features I SPPF The input is further enhanced in the C2PSA module to obtain the enhanced feature I C2PSA ;
[0089] Step 2.1.5: Strengthen Enhancement Feature I C2PSA The feature is input into the deformable attention module for processing, and the size is M / 2 n+1 ×N / 2 n+1 Attention Features I DAttetion After introducing the deformable attention mechanism into the backbone network, the model can flexibly adjust the focus area according to the image content, thereby accurately locking the defect location and significantly improving the defect detection accuracy. Figure 3a , Figure 3b and Figure 3c This effect is clearly demonstrated: Figure 3a This is an example image in the EL image dataset; Figure 3b The model heat map after applying the deformable attention mechanism is presented. Figure 3c This is the model heat map when the mechanism is not applied.
[0090] Step 2.2: The feature fusion network consists of five feature fusion units and performs DAttetion , the n-1th coarse feature And the n-2th coarse feature After processing, the third fusion feature F is obtained fused,3 , the fourth fusion feature F fused,4 and the fifth fusion feature F fused,5 ; The feature fusion network can effectively take into account the high-level abstract information and low-level location information of defect features, thereby significantly improving the model's detection ability for defects of different scales. Figure 4a and Figure 4b The relevant comparison effects are shown: Figure 4a This is the detection effect diagram after the model uses five feature fusion units. Figure 4b This is the detection effect diagram when the model does not use five feature fusion units.
[0091] Step 2.2.1: The first feature fusion unit consists of a first convolutional module Conv1, a second convolutional module Conv2, a first splicing module, and a first CSPStage module in sequence, and processes I DAttetion and the (n - 1)-th rough feature to obtain a first fusion feature F fused,1 ;
[0092] After processing the attention feature I DAttetion through the first convolutional block Conv1, a first convolutional feature I n+1 with a size of M / 2 n +1 × N / 2 conv,1 is obtained;
[0093] The (n - 1)-th rough feature output by the (n - 1)-th local feature extraction unit is input into the second convolutional block Conv2 for processing, and a second convolutional feature I n+1 with a size of M / 2 n+1 × N / 2 conv,2 is obtained;
[0094] The first convolutional feature I conv,1 and the second convolutional feature I conv,2 are spliced through the first splicing module to obtain a first spliced feature I n+1 with a size of M / 2 n+1 × N / 2 concat,1 ;
[0095] The first spliced feature I concat,1 is processed through the first CSPStage module to obtain a first fusion feature F n+1 with a size of M / 2 n+1 × N / 2 fused,1 .
[0096] Step 2.2.2: The second feature fusion unit consists of a first upsampling module, a third convolutional module Conv3, a second splicing module, and a second CSPStage module, and processes F fused,1 , the (n - 2)-th rough feature and the (n - 1)-th rough feature to obtain a second fusion feature F n with a size of M / 2 n × N / 2 fused,2 ;
[0097] Input the first fusion feature F fused,1 into the first upsampling module to obtain a size of M / 2 n × N / 2 nThe first upsampling feature I unsample,1 ;
[0098] The (n - 2)-th coarse feature output by the (n - 2)-th local feature extraction unit After being input into the third convolutional module Conv3, the third convolutional feature I with a size of M / 2 n × N / 2 n is obtained conv,3 ;
[0099] The first upsampling feature I with the same size unsample,1 , the third convolutional feature I conv,3 and the (n - 1)-th coarse feature are jointly input into the second splicing module for processing, and the second splicing feature I with a size of M / 2 n × N / 2 n is obtained concat,2 ;
[0100] The second splicing feature I concat,2 is further processed by the second CSPStage module to obtain the second fusion feature F with a size of M / 2 n × N / 2 n fused,2 .
[0101] Step 2.2.3: The third feature fusion unit consists of a second upsampling module, a third splicing module, and a third CSPStage module, and processes F fused,2 and the (n - 2)-th coarse feature to obtain the third fusion feature F with a size of M / 2 n-1 × N / 2 n-1 fused,3 ;
[0102] The second fusion feature F fused,2 is input into the second upsampling module for processing to obtain the second upsampling feature I with a size of M / 2 n-1 ×N / 2 n-1 unsample,2 ;
[0103] The second upsampling feature I unsample,2 and the (n - 2)-th coarse feature are subjected to feature splicing through the third splicing module to obtain the third splicing feature I with a size of M / 2 n-1 × N / 2 n-1 concat,3 ;
[0104] The third splicing feature I concat,3 is processed by the third CSPStage module to obtain a size of M / 2 n-1 × N / 2 n-1 The third fusion feature F fused,3 .
[0105] Step 2.2.4: The fourth feature fusion unit consists of a fourth convolutional module Conv4, a fourth splicing module, and a fourth CSPStage module; and processes F fused,3 and F fused,2 to obtain a fourth fusion feature F with a size of M / 2 n × N / 2 n ; fused,4 ;
[0106] Input the third fusion feature F fused,3 into the fourth convolutional module Conv4 for processing to obtain a fourth convolutional feature I with a size of M / 2 n ×N / 2 n ; conv,4 ;
[0107] After the fourth convolutional feature I conv,4 is feature - spliced with the second fusion feature F fused,2 through the fourth splicing module, a fourth spliced feature I with a size of M / 2 n × N / 2 n is obtained; concat,4 ;
[0108] After the fourth spliced feature I concat,4 is processed by the fourth CSPStage module, a fourth fusion feature F with a size of M / 2 n × N / 2 n is obtained; fused,4 .
[0109] Step 2.2.5: The fifth feature fusion unit consists of a fifth convolutional module Conv5, a sixth convolutional module Conv6, a fifth splicing module, and a fifth CSPStage module, and processes F fused,2 and F fused,4 and F fused,1 to obtain a fifth fusion feature F with a size of M / 2 n+1 ×N / 2 n+1 ; fused,5 ;
[0110] Input the second fusion feature F fused,2 into the fifth convolutional module Conv5 for processing to obtain a fifth convolutional feature I conv,5 ;
[0111] Input the fourth fusion feature F fused,4 into the sixth convolutional module Conv6 for processing to obtain a sixth convolutional feature I conv,6 ;
[0112] Input the fifth convolutional feature I conv,5 , the sixth convolutional feature I conv,6 , and the first fusion feature F fused,1 into the fifth splicing module for processing to obtain the fifth spliced feature I concat,5 ;
[0113] The fifth spliced feature I concat,5 is processed by the fifth CSPStage module to obtain the fifth fusion feature F n+1 with a size of M / 2 n+1 × N / 2 fused,5 .
[0114] Step 2.3: The prediction network consists of two parallel branches and processes F fused,3 , F fused,4 , and F fused,5 to obtain the class prediction label of the I el image and the position prediction label g' of the I el image
[0115] Step 2.3.1: The first branch consists of three parallel class prediction units, which respectively process F fused,3 , F fused,4 , and F fused,5 to obtain the class prediction label of the I el image ;
[0116] Among them, the f-th class prediction unit consists of the f-th first convolutional module Conv f,1 , the f-th second convolutional module Conv f,2 , and the f-th first two-dimensional convolutional module Conv2d f,1 ; f = 1, 2, 3;
[0117] Input any (f + 2)-th fusion feature F fused,3 , F fused,4 , and F fused,5 in F fused,f+2 into the f-th class prediction unit. First, after being processed by the first convolutional module Conv f,1 , the first convolutional feature I conv,f,1 output by the f-th class prediction unit is obtained. Then, I conv,f,1 is processed by the second convolutional module Conv f,2 to obtain the second convolutional feature I conv,f,2 output by the f-th class prediction unit. Finally, I conv,f,2 is processed by the first two-dimensional convolutional module Conv2df,1 After processing, the first two-dimensional convolution feature I output by the f-th category prediction unit is obtained conv2d,f,1 .
[0118] Step 2.3.2: The second branch consists of three parallel location prediction units, and processes F fused,3 , F fused,4 , and F fused,5 respectively to obtain the location prediction label g' of the I el image;
[0119] Among them, the f-th location prediction unit is successively composed of the f-th first depthwise separable convolution module DWConv f,1 , the f-th third convolution module Conv f,3 , the f-th second depthwise separable convolution module DWConv f,2 , the f-th fourth convolution module Conv f,4 , and the f-th second two-dimensional convolution module Conv2d f,2 .
[0120] Input F fused,f+2 into the f-th location prediction unit. First, after being processed by the f-th first depthwise separable convolution module DWConv f,1 , the first depthwise separable convolution feature I dwconv,f,1 output by the f-th location prediction unit is obtained. Secondly, I dwconv,f,1 is processed by the f-th third convolution module Conv f,3 to obtain the third convolution feature I conv,f,3 output by the f-th location prediction unit. Then, I conv,f,3 is processed by the f-th second depthwise separable convolution module DWConv f,2 to obtain the second depthwise separable convolution feature I dwconv,f,2 output by the f-th location prediction unit. Next, I dwconv,f,2 is processed by the f-th fourth convolution module Conv f,4 to obtain the fourth convolution feature I conv,f,4 output by the f-th category prediction unit. Finally, I conv,f,4 is processed by the f-th second two-dimensional convolution module Conv2d f,2 to obtain the second two-dimensional convolution feature I conv2d,f,2 output by the f-th location prediction unit.
[0121] Step 2.3.3: Concatenate I conv2d,f,1 and I conv2d,f,2 respectively in the channel dimension to obtain a size of A f ×H f ×W fThe f-th predicted splicing feature set I conv2d,f ={ I conv2d,1 , I conv2d,2 ,I conv2d,3}, where I conv2d,1 represents the 1st predicted splicing feature, I conv2d,2 represents the 2nd predicted splicing feature, I conv2d,3 represents the 3rd predicted splicing feature, A f represents the number of channels of the f-th predicted splicing feature, H f represents the height of the f-th predicted splicing feature, W f represents the width of the f-th predicted splicing feature.
[0122] Step 2.3.4: After reshaping and splicing I conv2d,1 , I conv2d,2 ,I conv2d,3 , the position prediction label g’ and the class prediction label c’ of the I el image can be obtained, and let the position prediction label of the i-th defect in g’ be denoted as g i ’ =( ,w i ’,h i ’), where represents the center point of the predicted bounding box where the i-th defect is located in I el , w i ’ is the width of the predicted bounding box where the i-th defect is located in I el , h i ’ is the height of the predicted bounding box where the i-th defect is located in I el ; let represent the predicted class of the i-th defect in c’.
[0123] Step 3: Construct the total loss function Loss;
[0124] Step 3.1: Use Equation (1) to construct the loss function of the i-th defect :
[0125] (1)
[0126] In Equation (1), and are 2 weights; represents the class loss of the i-th defect, represents the bounding box loss of the i-th defect, and there is:
[0127] (2)
[0128] (3)
[0129] In formula (3), represents the Euclidean distance, represents the diagonal length of the minimum bounding rectangle of the i-th predicted defect box and the i-th ground truth defect box, represents the overlap degree between the i-th predicted defect box and the i-th ground truth defect box, is the similarity of the length and width between the i-th predicted defect box and the i-th ground truth defect box, is the weight function of the i-th defect, and there is:
[0130] (4)
[0131] (5)
[0132] Step 3.2: Calculate the total loss function Loss using formula (6):
[0133] (6)
[0134] Step 4: Iteratively train the YOLOv11-MSFF model using the SGD optimizer, and adjust the model parameters according to the total loss function Loss until the total loss function Loss converges, so as to obtain a trained EL image defect detection model for detecting defects in EL images.
[0135] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0136] In this embodiment, a computer-readable storage medium stores a computer program. When the computer program is run by a processor, it executes the steps of the above method.
Claims
1. A photovoltaic EL defect detection method based on the YOLOv11-MSFF model, characterized in that, including the following steps: Step 1: Use a high-resolution CCD camera to capture images of photovoltaic modules under near-infrared light and form an EL image set. Denote any EL image in the EL image set as I el ; and the size of I el is M×N, where M represents the length of the EL image and N represents the width of the EL image; Label the positions and categories of s defects in the EL image I, obtaining the position labels and category labels of the s defects; let the position label of the i-th defect in I el be denoted as el , and =( , , , ), where represents the center point of the true bounding box of the i-th defect at the position in I el , is the width of the true bounding box of the i-th defect at the position in I el , is the height of the true bounding box of the i-th defect at the position in I el , and denote the category label of the i-th defect in I el as ; Step 2: Construct the YOLOv11-MSFF model, which sequentially includes: a feature extraction network, a feature fusion network, and a prediction network, and process I el to obtain I el the class prediction label of the image and I el the location prediction label of the image ; Step 3: Construct the total loss function Loss; Step 4: Iteratively train the YOLOv11-MSFF model using the SGD optimizer and adjust the model parameters according to the total loss function Loss until the total loss function Loss converges, thereby obtaining a trained EL image defect detection model for detecting defects in EL images.
2. The photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to claim 1, characterized in that, The said Step 2 includes: Step 2.1: The feature extraction network consists of an initial convolutional block Conv0, n local feature extraction units, a fast spatial pyramid pooling module, a C2PSA module, and a deformable attention module in sequence, and processes I el to generate a set of coarse features { | = 1, 2, …, } from the n local feature extraction units, and outputs attention feature I n+1 with a size of M / 2 n+1 × N / 2 DAttetion by the deformable attention module; where represents the u-th coarse feature output by the u-th local feature extraction unit, and the size of is M / 2 u+1 × N / 2 u+1 ; Step 2.2: The feature fusion network includes five feature fusion units, and processes I DAttetion , the (n - 1)-th coarse feature and the (n - 2)-th coarse feature to obtain the third fusion feature F fused,3 , the fourth fusion feature F fused,4 and the fifth fusion feature F fused,5 ; Step 2.3: The prediction network consists of two parallel branches and processes F fused,3 , F fused,4 and F fused,5 to obtain the class prediction label of the I el image and the position prediction label of the I el image .
3. The photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to claim 2, wherein, Step 2.1 includes the following steps: Step 2.1.1: The initial convolution block Conv0 processes I el to obtain an initial convolution feature I conv0 with its size reduced to M / 2 × N / 2 conv0 ; Step 2.1.2: Initial Convolution Feature I conv0 After being processed by n local feature extraction units in sequence, a set of coarse features { | = 1, 2, …, } is correspondingly generated, where represents the u-th coarse feature output by the u-th local feature extraction unit, and has a size of M / 2 u+1 × N / 2 u+1 ; Step 2.1.3: Input the nth rough feature output by the nth local feature extraction unit into the fast spatial pyramid pooling module for preliminary enhancement processing to obtain the preliminary enhanced feature I SPPF ; Step 2.1.4: Preliminary enhanced feature I SPPF Input it into the C2PSA module for further enhancement to obtain the enhanced enhanced feature I C2PSA ; Step 2.1.5: Strengthen the enhanced feature I C2PSA The feature is input into the deformable attention module for processing to obtain the attention feature I with a size of M / 2 n+1 ×N / 2 n+1 DAttetion . 4. A photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to claim 3, characterized in that, The said Step 2.2: Step 2.2.1: The first feature fusion unit consists of a first convolutional module Conv1, a second convolutional module Conv2, a first splicing module, and a first CSPStage module in sequence, and processes I DAttetion and the (n - 1)-th coarse feature to obtain a first fusion feature F fused,1 ; Attention feature I DAttetion After being processed by the first convolutional block Conv1, a first convolutional feature I with a size of M / 2 n+1 × N / 2 n+1 is obtained conv,1 ; The (n - 1)-th rough feature output by the (n - 1)-th local feature extraction unit After being input into the second convolutional block Conv2 for processing, a second convolutional feature I with a size of M / 2 n+1 × N / 2 n+1 is obtained conv,2 ; The first convolutional feature I conv,1 and the second convolutional feature I conv,2 are concatenated through the first concatenation module to obtain the first concatenated feature I n+1 with a size of M / 2 n+1 × N / 2 concat,1 ; The first splicing feature I concat,1 After being processed by the first CSPStage module, a first fusion feature F with a size of M / 2 n+1 × N / 2 n+1 is obtained fused,1 ; Step 2.2.2: The second feature fusion unit consists of a first upsampling module, a third convolutional module Conv3, a second splicing module, and a second CSPStage module, and processes F fused,1 , the (n - 2)-th coarse feature and the (n - 1)-th coarse feature to obtain a second fusion feature F n with a size of M / 2 n × N / 2 fused,2 ; Input the first fusion feature F fused,1 into the first upsampling module to obtain the first upsampling feature I n with a size of M / 2 n × N / 2 unsample,1 ; The (n - 2)-th rough feature output by the (n - 2)-th local feature extraction unit After being input into the third convolution module Conv3, a third convolution feature I with a size of M / 2 n × N / 2 n is obtained conv,3 ; Input the first upsampling feature I with the same size unsample,1 , the third convolutional feature I conv,3 and the (n - 1)-th coarse feature into the second splicing module for processing together to obtain the second splicing feature I with a size of M / 2 n × N / 2 n ; concat,2 ; Second splicing feature I concat,2 After further processing by the second CSPStage module, a second fusion feature F with a size of M / 2 n × N / 2 n is obtained fused,2 ; Step 2.2.3: The third feature fusion unit consists of a second upsampling module, a third splicing module, and a third CSPStage module, and processes F fused,2 and the (n - 2)-th coarse feature to obtain a third fused feature F n-1 with a size of M / 2 n-1 × N / 2 fused,3 ; Input the second fusion feature F fused,2 into the second upsampling module for processing to obtain a second upsampled feature I n-1 with a size of M / 2 n-1 × N / 2 unsample,2 ; Second upsampling feature I unsample,2 and the (n - 2)-th coarse feature After feature splicing by the third splicing module, the third spliced feature I with a size of M / 2 n-1 × N / 2 n-1 is obtained concat,3 ; The third splicing feature I concat,3 Through the processing of the third CSPStage module, a third fusion feature F with a size of M / 2 n-1 × N / 2 n-1 is obtained fused,3 ; Step 2.2.4: The fourth feature fusion unit consists of a fourth convolutional module Conv4, a fourth splicing module, and a fourth CSPStage module; and processes F fused,3 and F fused,2 to obtain a fourth fused feature F n × N / 2 n with a size of M / 2 fused,4 ; Input the third fusion feature F fused,3 into the fourth convolutional module Conv4 for processing, to obtain a fourth convolutional feature I n with a size of M / 2 n × N / 2 conv,4 ; The fourth convolutional feature I conv,4 and the second fusion feature F fused,2 are subjected to feature splicing through a fourth splicing module, and a fourth splicing feature I with a size of M / 2 n × N / 2 n is obtained; concat,4 ; Fourth splicing feature I concat,4 After being processed by the fourth CSPStage module, the obtained size is M / 2 n × N / 2 n of the fourth fusion feature F fused,4 ; Step 2.2.5: The fifth feature fusion unit consists of a fifth convolutional module Conv5, a sixth convolutional module Conv6, a fifth splicing module, and a fifth CSPStage module, and processes F fused,2 and F fused,4 as well as F fused,1 to obtain the fifth fusion feature F n+1 with a size of M / 2 n+1 × N / 2 fused,5 ; The second fusion feature F fused,2 is input into the fifth convolutional module Conv5 for processing to obtain the fifth convolutional feature I conv,5 ; The fourth fusion feature F fused,4 is input into the sixth convolutional module Conv6 for processing to obtain the sixth convolutional feature I conv,6 ; Input the fifth convolutional feature I conv,5 , the sixth convolutional feature I conv,6 , and the first fusion feature F fused,1 into the fifth splicing module for processing to obtain the fifth spliced feature I concat,5 ; The Fifth Splicing Feature I concat,5 Processed by the fifth CSPStage module to obtain the fifth fusion feature F with a size of M / 2 n+1 ×N / 2 n+1 fused,5 . 5. A photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to claim 4, characterized in that The said Step 2.3 includes: Step 2.3.1: The first branch consists of three parallel class prediction units, and processes F fused,3 , F fused,4 , and F fused,5 respectively to obtain the class prediction label of the I el image ; Among them, the f-th category prediction unit is successively composed of the f-th first convolution module Conv f,1 , the f-th second convolution module Conv f,2 , and the f-th first two-dimensional convolution module Conv2d f,1 ; f = 1, 2, 3; Input F fused,3 、F fused,4 and F fused,5 among any of the (f + 2)-th fusion features F fused,f+2 into the f-th class prediction unit. First, after being processed by the first convolutional module Conv f,1 , the first convolutional feature I conv,f,1 output by the f-th class prediction unit is obtained. Then, I conv,f,1 is processed by the second convolutional module Conv f,2 , and the second convolutional feature I conv,f,2 output by the f-th class prediction unit is obtained. Finally, I conv,f,2 is processed by the first two-dimensional convolutional module Conv2d f,1 , and the first two-dimensional convolutional feature I conv2d,f,1 output by the f-th class prediction unit is obtained; Step 2.3.2: The second branch consists of three parallel location prediction units, which respectively process F fused,3 , F fused,4 and F fused,5 to obtain the location prediction label of the I el image ; Among them, the f-th position prediction unit is successively composed of the f-th first depthwise separable convolution module DWConv f,1 , the f-th third convolution module Conv f,3 , the f-th second separable convolution module DWConv f,2 , the f-th fourth convolution module Conv f,4 , and the f-th second two-dimensional convolution module Conv2d f,2 ; Input F fused,f+2 into the f-th position prediction unit. First, after being processed by the f-th first depthwise separable convolution module DWConv f,1 , the first depthwise separable convolution feature I output by the f-th position prediction unit is obtained dwconv,f,1 . Secondly, I dwconv,f,1 is processed by the f-th third convolution module Conv f,3 , and the third convolution feature I output by the f-th position prediction unit is obtained conv,f,3 . Then, I conv,f,3 is processed by the f-th second separable convolution module DWConv f,2 , and the second depthwise separable convolution feature I output by the f-th position prediction unit is obtained dwconv,f,2 . Next, I dwconv,f,2 is processed by the f-th fourth convolution module Conv f,4 , and the fourth convolution feature I output by the f-th class prediction unit is obtained conv,f,4 . Finally, I conv,f,4 is processed by the f-th second two-dimensional convolution module Conv2d f,2 , and the second two-dimensional convolution feature I output by the f-th position prediction unit is obtained conv2d,f,2 ; Step 2.3.3: Concatenate I conv2d,f,1 and I conv2d,f,2 along the channel dimension respectively to obtain the f-th predicted concatenated feature set I f ×H f ×W f where I conv2d,f ={ I conv2d,1 , I conv2d,2 ,I conv2d,3}, where I conv2d,1 represents the 1st predicted concatenated feature, I conv2d,2 represents the 2nd predicted concatenated feature, I conv2d,3 represents the 3rd predicted concatenated feature, A f represents the number of channels of the f-th predicted concatenated feature, H f represents the height of the f-th predicted concatenated feature, and W f represents the width of the f-th predicted concatenated feature; Step 2.3.4: Reshape and splice I conv2d,1 , I conv2d,2 ,I conv2d,3 After reshaping and splicing, obtain the position prediction label of I el image and the class prediction label , and let The position prediction label of the i-th defect in be denoted as =( , , ), where represents the center point of the predicted bounding box of the i-th defect in I el image, is the width of the predicted bounding box of the i-th defect in I el image, is the height of the predicted bounding box of the i-th defect in I el image; Let represent the predicted class of the i-th defect in 6. The photovoltaic EL defect detection method based on the YOLOv11-MSFF model according to claim 5, characterized in that The said Step 3 includes: Step 3.1: Construct the loss function of the \(i\)-th defect using Equation (1) : (1) In formula (1), and are two weights; represents the class loss of the i-th defect, represents the bounding box loss of the i-th defect, and there is: (2) (3) In formula (3), represents the Euclidean distance, represents the diagonal length of the minimum bounding rectangle of the i-th predicted defect box and the i-th true defect box, represents the overlap degree of the i-th predicted defect box and the i-th true defect box, is the similarity of the length and width of the i-th predicted defect box and the i-th true defect box, is the i-th defect weight function, and there is: (4) (5) Step 3.2: Calculate the total loss function Loss using Equation (6): (6) 。 7. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute any one of the photovoltaic EL defect detection methods recited in claims 1-6, and the processor is configured to execute the program stored in the memory.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by the processor, it executes the steps of any one of the photovoltaic EL defect detection methods recited in claims 1-6.