Photovoltaic panel defect detection method based on improved YOLOv7
By improving the YOLOv7 network structure, introducing the SimAM attention mechanism, SPD-Conv and NWD, and optimizing the feature extraction and regression loss functions, the problems of low accuracy and slow speed in photovoltaic panel defect detection are solved, and the ability to recognize small target defects is improved.
Patent Information
- Application Number
- CN202510625691.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-19
AI Technical Summary
Existing photovoltaic panel defect detection technology has problems such as low detection accuracy, slow speed and difficulty in identifying multiple types of defects, especially small target defects are difficult to identify under complex background interference.
An improved YOLOv7 network structure is adopted. By introducing the SimAM attention mechanism, SPD-Conv and distribution shift convolution (DSConv) and normalized Gauss-Wasserstein distance (NWD), the model's ability to detect small targets is enhanced, and the feature extraction and regression loss functions are optimized.
The accuracy and speed of photovoltaic panel defect detection are improved, especially the ability to identify small target defects, which reduces the model volume and speeds up the detection speed.
Smart Images

Figure CN120672660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photovoltaic panel defect detection, and more particularly to a photovoltaic panel defect detection method based on an improved YOLOv7. Background Art
[0002] Photovoltaic power generation technology continues to develop rapidly in my country's new energy system development. In the first half of 2024, newly connected grid capacity exceeded 102.48 GW, with distributed photovoltaics contributing over 50%, marking a new stage of large-scale application in the industry. As the core component of a photovoltaic system, the quality of photovoltaic panels directly determines power generation efficiency and system stability. However, they face significant challenges throughout their lifecycle: process defects in production, mechanical stress during transportation, and environmental factors such as continuous UV radiation, temperature and humidity fluctuations, and chemical corrosion during operation can all lead to various defects in panels, including cracks, broken grids, hidden cracks, and black cores.
[0003] The current mainstream defect detection technologies mainly include three categories: (1) detection methods based on physical parameters; (2) traditional image processing technology; and (3) deep learning algorithms. From the perspective of technological evolution, the physical parameter detection method mainly makes judgments through analysis of current and voltage characteristics, but it has two significant defects: first, the threshold setting of the measurement parameters depends on the operator's experience, and there is a 5%-15% interpretation error; second, contact measurement may cause secondary damage to the surface of the solar panel. The second-generation image processing method uses image segmentation and feature extraction technology. Although it can achieve non-contact detection, it is limited by the algorithm design. The types of defects that can be detected are limited, and the detection effect is poor, which has great limitations. The third generation detection method is based on deep learning. This method uses convolutional networks to extract features and detect targets from images. The advantages are high accuracy, fast speed, and less manual participation. However, it has problems such as the identification, classification and positioning of multiple types of defects and the difficulty in identifying small target defects under complex background interference. Summary of the Invention
[0004] To address the above problems, the present invention proposes a photovoltaic panel defect detection method based on improved YOLOv7, which strives to reduce the model volume and speed up the detection speed while ensuring that the model detection accuracy reaches a high standard.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a photovoltaic panel defect detection method based on improved YOLOv7, the detection method specifically comprising:
[0006] S1. Filter and enhance the collected EL images of photovoltaic panels to identify defective EL images. The defects in the images are classified into seven categories: cracks, black cores, dislocations, failures, thick lines, broken grids, and star-shaped cracks. Defects are then annotated on the EL images to create a dataset of defective EL images of photovoltaic panels. Finally, the dataset of defective EL images of photovoltaic panels is divided into a training set, a test set, and a validation set in proportion.
[0007] S2. Using the YOLOv7 network structure as the basic framework, the photovoltaic panel defect electroluminescence image dataset established in step S1 is used for training, verification, and testing to obtain a photovoltaic panel defect detection model. During the training process, the SimAM attention mechanism is introduced into the YOLOv7 model to improve the model's attention to positive samples and enhance its key feature extraction capabilities. The MP structure is optimized, and Space-to-depth-Conv (SPD-Conv) is introduced to improve the small target detection capability of the YOLOv7 model. The head network of the YOLOv7 model is improved through distributed shift convolution (DSConv). The normalized Gauss-Wasserstein distance is introduced into the regression loss function of the improved YOLOv7 model to measure the similarity between the predicted bounding box and the true target bounding box.
[0008] S3. Input the newly collected photovoltaic panel electroluminescence image data into the photovoltaic panel defect detection model to perform defect detection and output the detection results.
[0009] A better technical solution of the present invention: in the step S1, the electroluminescent images of photovoltaic panels that are not clear or severely distorted are screened and eliminated, and then the filtered original images are mirrored, rotated, brightness adjusted and contrast adjusted, and the image data is divided into defective and non-defective positive and negative samples in proportion.
[0010] A preferred technical solution of the present invention: In the S1 step, LabelImg software is used to label the photovoltaic panel defect electroluminescent images after screening and data enhancement processing; the photovoltaic panel defect electroluminescent image data set is divided into a training set, a test set and a validation set in proportion in a ratio of 7:2:1.
[0011] A further technical solution of the present invention is as follows: In step S2, the YOLOv7 network structure is used as the basic framework, and the photovoltaic panel defect electroluminescence image dataset established in step S1 is used for training, verification, and testing to obtain a photovoltaic panel defect detection model.
[0012] S201. Convert images from the photovoltaic panel defect electroluminescence image dataset into RBG three-channel images and feed them into the backbone network of the YOLOv7 model. The first four convolution layers of the network perform initial feature extraction and downsampling. After four convolution layers, the output convolution image enters the ELAN structure for feature information extraction. The simAM attention mechanism is inserted into the transmission path to enhance its feature information extraction capabilities.
[0013] S202. Introducing SPD-Conv into the MP structure of the model, replacing the downsampling branch in the original structure with SPD-Conv, followed by normalization and connecting the Silu activation function, while retaining the pooling branch in the original structure. The extracted feature information is fed into the MP structure reconstructed based on the SPD-Conv module for downsampling and pooling operations. Three sets of ELAN and MP with different gradients are used to extract feature information at different levels. Finally, the multi-scale features are fused in the SPC pooling layer, and feature images of different scales are output to the neck network of the YOLOv7 model. The feature images output by ELAN-SA are downsampled by MP to reduce feature information loss.
[0014] S203. The feature image processed by S202 is transmitted to the ELAN-W module based on distribution shift convolution reconstruction in the neck network and then participates in the subsequent MP downsampling process to quantize the weights and find the optimal distribution shift of fixed integer weights to improve the computational complexity;
[0015] S204. After feature fusion, the feature image is input into the reparameterized convolution inference module in the head network of the YOLOv7 model. During the inference stage, the reparameterized convolution fuses the multiple feature information transmitted by the neck network into a single feature map based on the weights. Finally, the three detection heads of the head network detect large, medium, and small scale defect targets in the electroluminescent image of the photovoltaic panel based on feature maps of different scales.
[0016] A further technical solution of the present invention is as follows: SimAM in step S2 is an attention mechanism based on local self-similarity of feature maps. It dynamically adjusts the weight of each pixel by calculating the similarity between each pixel in the feature map and its surrounding pixels, thereby enhancing important features and suppressing irrelevant features. The formula of the SimAM attention mechanism is as follows:
[0017] (1) Derivation of energy function
[0018] SimAM measures the importance of the t-th neuron by minimizing the following energy function:
[0019]
[0020] Where t is the spatial position of the target neuron; μt is the eigenvalue (original input) of the target neuron t; is the weighted eigenvalue of the target neuron t (the variable to be optimized); is the weighted variance of the target neuron t (the variable to be optimized); λ is the regularization coefficient (a hyperparameter to prevent the denominator from being zero);
[0021] (2) Energy minimization solution
[0022] By minimizing the energy function e t , get the optimal and
[0023]
[0024] Where N is the total number of neurons in the same channel (spatial dimension); x i is the eigenvalue of the i-th neuron;
[0025] (3) Attention weight calculation
[0026] Final attention weight w t By normalizing the energy function, we can get:
[0027]
[0028] Where w t is the attention weight of the t-th neuron (the lower the energy, the higher the weight);
[0029] (4) Feature enhancement formula
[0030] Enhance the input features:
[0031]
[0032] Where x t is the original input feature; The enhanced feature sigmoid is a nonlinear activation function (limiting the weights between 0 and 1).
[0033] The preferred technical solution of the present invention is as follows: in the step S2, the SimAM attention mechanism is added to the ELAN module of the YOLOv7 main network, and the SA module is added after the SiLU function of the CBS module at the end of the ELAN module to form an ELAN-SA module; the introduction of the SA attention mechanism enables the network to perform feature recalibration, allowing the model to pay more attention to the channel features that are important to the information;
[0034] The MP module consists of two main branches, whose main function is to perform downsampling; the upper branch first downsamples through a maximum pooling layer, and then uses a 1x1 convolution layer to adjust the number of channels; the lower branch starts with a 1x1 convolution layer to change the number of channels, and then applies a 3x3 convolution with a stride of 2 for further downsampling; finally, the downsampled results of the two branches are merged together to form the final output.
[0035] A further technical solution of the present invention is as follows: in the step S2, the normalized Gauss-Wasserstein distance is introduced into the regression loss function of the improved YOLOv7 model to measure the similarity between the predicted bounding box and the true target bounding box, and the NWD is used to measure the similarity of the derived Gaussian distribution. For the horizontal bounding box R=(c x ,c y ,w,h), where (c x ,c y ) are the coordinates of the center point, w and h are the width and height respectively, and R is modeled as a Gaussian distribution N(μ,Σ), where:
[0036]
[0037] Where μ and ∑ are the mean vector and covariance matrix of the Gaussian distribution respectively.
[0038] Then use Wasserstein distance to calculate the bounding box R1 = (c x1 ,c y1 ,w1,h1) and bounding box R2=(c x2 ,c y2 ,w2,h2) between two Gaussian distribution distances:
[0039]
[0040] Since this distance metric cannot directly measure the similarity between bounding boxes R1 and R2, it is normalized in exponential form to obtain a new metric L NWD :
[0041]
[0042] Where C is a constant closely related to the dataset.
[0043] The preferred technical solution of the present invention: The ELAN module is an efficient network structure design. ELAN has two branches. In Backbon, the upper branch is a 1x1 convolution, which is responsible for changing the number of channels. The lower branch first performs a 1x1 convolution, which is also used to change the number of channels. It is followed by four 3x3 feature extraction convolutions and finally concats the four feature information to obtain the final result. The difference in Neck is that the lower branch has two more feature outputs than the ELAN in Backbon, and the final result is the concat sum of 6 feature information.
[0044] The present invention preprocesses the collected electroluminescent images of photovoltaic panels to establish a photovoltaic panel defect electroluminescent image dataset, and divides the dataset images into a training set, a test set, and a validation set in proportion; a YOLOv7 network structure is used as the basic structure, and the established photovoltaic panel defect electroluminescent image dataset is used for training to obtain a photovoltaic panel defect detection model; during the training process, the SA attention mechanism is added to the ELAN module of the YOLOv7 main network, and the SA module is added after the SiLU function of the last CBS module of the ELAN module to form an ELAN-SA module; the network is enabled to perform feature recalibration, so that the model pays more attention to the channel features with important information and suppresses those unimportant channel features, thereby increasing the network's recognition of small targets of photovoltaic panel defects and improving the detection accuracy of the model. Then, the SPD-Conv structure is introduced into the MP module. The SPD structure downsamples features, preserving all information in the channel dimension, thereby improving the model's small object detection capabilities. The convolution in the YOLOv7 head detection network is replaced with a distribution shift convolution. The normalized Gauss-Wasserstein distance is then introduced into the detection head's regression loss function, replacing the original model's intersection-over-union (IoU) calculation to enhance small object detection capabilities. Finally, the newly acquired image data is input into the trained photovoltaic panel defect detection model for defect detection, and the detection results are output.
[0045] The present invention embeds the NWD metric into the bounding box regression loss function, replacing the original IoU metric to improve the accuracy of small target detection. The original loss function in YOLOv7 is CIoU. CIoU is very sensitive to the position offset of small targets when measuring the difference between the predicted box and the true box. The slight offset of the target object will affect the IoU. The NWD metric is insensitive to targets of different scales and is not easily affected by the position deviation of small targets. The NWD metric replaces the IoU metric, which not only achieves the goal of high-precision small target detection, but also does not affect the convergence speed of training. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is the main structure flow chart of the improved YOLOv7 in the present invention;
[0047] Figure 2 This is a flow chart of the main structure of the YOLOv7 model in the present invention;
[0048] Figure 3 This is a flow chart of the main module structure of the YOLOv7 model in the present invention;
[0049] Figure 4 The SA structure diagram of the attention mechanism introduced into the ELAN module of this invention;
[0050] Figure 5 The SPD structure diagram is introduced for the MP module in the present invention;
[0051] Figure 6 The distributed shift convolution DS structure diagram is introduced for the ELAN-W module in the present invention;
[0052] Figure 7 This is a flowchart of the data set preprocessing in the present invention;
[0053] Figure 8 This is a flow chart of the photovoltaic panel defect detection model based on improved YOLOv7. DETAILED DESCRIPTION
[0054] The present invention is further described below in conjunction with the embodiments and drawings. The exemplary embodiments of the present invention and their description are used to understand the present invention and do not constitute an improper limitation of the present invention. The technical solutions shown in the accompanying drawings are specific solutions of the embodiments of the present invention and are not intended to limit the scope of the invention claimed for protection. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0055] The embodiment provides a photovoltaic panel defect detection method based on improved YOLOv7, such as Figures 1 to 8 As shown, the specific steps include:
[0056] S1. Establish a dataset of defective electroluminescent images of photovoltaic panels. Preprocess the collected electroluminescent images of photovoltaic panels and classify the defects in the images into seven categories: cracks, black cores, dislocations, failures, thick lines, broken grids, and star-shaped cracks. Then, annotate the defects of the electroluminescent images of photovoltaic panels and divide the dataset images into training, test, and validation sets in proportion.
[0057] The process of preprocessing the photovoltaic panel electroluminescence image set is as follows: first, since some images may be unclear or severely distorted, they are manually screened and eliminated; then, the data set is enhanced. Data enhancement is a common technique for generating more training data by performing various transformations on the original image, aiming to improve the generalization ability of the model; common data enhancement methods include mirror flipping, rotation transformation, brightness adjustment, and contrast adjustment; secondly, the image data is divided into positive and negative samples with defects and without defects in proportion, and then the defective image samples are annotated in detail. When marking defects, key information includes the type and location of the defect. When selecting the target box and performing mirror annotation, there are clear standards for defining the defect boundary and shape. Specifically, the boundary should be defined as close to the target as possible and not too far away to ensure that the annotation is on the tangent of the target or at a point outside the tangent. Regarding shape definition, the annotation box should completely surround the target to ensure that all relevant features are accurately captured. The annotation format required for this dataset is txt text format, and the tool used is LabelImg. LabelImg is a software specifically designed for target detection project annotation. It is simple, efficient, and easy to use. You need to set the image path, label save path, and label format. For defects of the same type, the same label box is used for annotation. Finally, two folders, img and lab, are output, which are used to store images and labels respectively. Finally, after image screening, data augmentation, and defect annotation, the total image dataset is divided into a training set, a test set, and a validation set in a ratio of 7:2:1.
[0058] S2. Using the YOLOv7 network structure as the framework, the model for photovoltaic panel defect detection was trained, validated, and tested using the EL image dataset of photovoltaic panel defects established in step S1. The YOLOv7 network structure consists of three parts: the backbone network (Backbone), the neck network (Neck), and the head network (Head). The backbone network extracts features from the input image. The input image information is initially extracted and downsampled through a four-layer CBS structure. It then passes through three sets of ELAN and MP with different gradients to extract features at different levels and refine semantic information. Finally, the multi-scale features are fused in the SPPCSPC pooling layer. The neck network further processes the fused feature information. Based on the traditional feature pyramid structure, ELAN and MP enhance semantic information from the bottom layer to the top layer, and simultaneously outputs three feature maps at different scales. However, due to the high downsampling ratio, most small object information is lost during the information processing, which is not conducive to small object detection. The head network is the prediction component, primarily consisting of a reparameterized convolution (RepConv) and a detection head (IDetect). During the inference phase, the reparameterized convolution fuses multiple feature information transmitted from the neck network into a single feature map based on weights. The three detection heads then detect large, medium, and small objects, respectively, based on feature maps of different scales.
[0059] The specific process of the above steps is as follows:
[0060] S201. Images from the photovoltaic panel defect electroluminescence image dataset are converted into 640*640 resolution RBG three-channel images and fed into the backbone network. The convolutional bounding box (CBS) layers in the first four layers of the network are responsible for feature extraction and downsampling. After four layers of convolutional bounding box (CBS), the output is a 128-channel 160*160 convolutional image that enters the ELAN structure for feature information extraction. The SimAM attention mechanism is inserted into the transmission path. The SA attention mechanism is added to the ELAN module of the YOLOv7 main network. The SA module is added after the SiLU function of the final CBS module in the ELAN module to form the ELAN-SA module. This reduces the loss of key feature information (ELAN-SA) and enables the network to perform feature recalibration, allowing the model to focus more on information-critical channel features and suppress those that are less important. This improves the network's recognition of small photovoltaic panel defect targets and enhances the model's detection accuracy. SimAM is an attention mechanism based on local self-similarity of feature maps. It dynamically adjusts the weight of each pixel by calculating the similarity between each pixel in the feature map and its surrounding pixels, thereby enhancing important features and suppressing irrelevant features. SimAM's innovation lies in its parameter-free nature, which enables the model to achieve excellent performance while maintaining low complexity. The formula is as follows:
[0061] ① Derivation of energy function
[0062] SimAM measures the importance of the t-th neuron by minimizing the following energy function:
[0063]
[0064] Where t is the spatial position of the target neuron; μ t is the eigenvalue (original input) of the target neuron t; is the weighted eigenvalue of the target neuron t (the variable to be optimized); is the weighted variance of the target neuron t (the variable to be optimized); λ is the regularization coefficient (a hyperparameter to prevent the denominator from being zero).
[0065] ②Energy minimization solution
[0066] By minimizing the energy function e t , get the optimal and
[0067]
[0068]
[0069] Where N is the total number of neurons in the same channel (spatial dimension); x i is the eigenvalue of the i-th neuron.
[0070] ③Attention weight calculation
[0071] Final attention weight w t By normalizing the energy function, we get:
[0072]
[0073] Where w t is the attention weight of the t-th neuron (the lower the energy, the higher the weight)
[0074] ④ Feature enhancement formula
[0075] Enhance the input features:
[0076]
[0077] Where x t is the original input feature; The enhanced feature sigmoid is a nonlinear activation function (limiting the weights between 0 and 1).
[0078] The SA attention mechanism is incorporated into the ELAN module of the YOLOv7 main network. The SA module is added after the SiLU function of the CBS module at the end of the ELAN module to form the ELAN-SA module. The SA attention mechanism enables the network to perform feature recalibration, allowing the model to focus more on information-rich channel features while suppressing less important ones. This facilitates the network's recognition of small defects in photovoltaic panels, thereby improving the model's detection accuracy.
[0079] S202. The SPD-Conv structure is introduced into the MP module, and then it is fed into the MP structure reconstructed based on the SPD-Conv module for downsampling and pooling operations. The ELAN-SA is combined with the MP structure, and the feature image output by the ELAN-SA is downsampled by the MP to reduce the loss of feature information. After four ELAN-SA and three MP operations, three scales of image information are output to the neck network (Neck) of the model, with sizes of 80*80 images with 128 channels, 40*40 images with 256 channels, and 20*20 images with 512 channels.
[0080] Photovoltaic panel defect datasets contain numerous tiny detection targets. The original YOLOv7 model's accuracy and performance degrade rapidly when faced with complex tasks involving small targets. In traditional convolutional neural network (CNN) architectures, the common cross-row convolution and pooling layers often result in the loss of fine-grained information. Furthermore, this architecture results in suboptimal feature representations learned by the model. This design generally exhibits no drawbacks when handling tasks with high resolution and moderately sized targets, as the presence of abundant redundant pixel information allows the model to easily skip cross-row convolution and pooling layers while still effectively learning features. However, in tasks with blurred or small targets, the assumption of abundant redundant information no longer holds true. In these cases, the current design suffers from the loss of fine-grained information and poor feature learning. Therefore, to improve the model's small target detection capabilities, we introduce Space-to-depth-Conv (SPD-Conv). Its simple structure consists of a spatial-to-depth layer and a non-cross-row convolution layer. The SPD structure downsamples features, preserving all information in the channel dimension.
[0081] The present invention introduces SPD-Conv into the MP structure of the model, replaces the downsampling branch in the original structure with SPD-Conv, performs normalization operation afterwards, connects the Silu activation function, and retains the pooling branch in the original structure.
[0082] S203. The processed feature information is transferred to the ELAN-W module based on distribution shift convolution reconstruction and then participates in the subsequent MP downsampling process. This quantizes the weights and finds the optimal distribution shift for fixed integer weights to improve computational complexity, thereby significantly improving the model speed and reducing the number of parameters. Because the output size of the three scale images input to the neck network (Neck) after UP upsampling and MP downsampling in the neck network (Neck) is the same as the size of the information input across the connection, it can be correctly connected.
[0083] S204. The convolution in the YOLOv7 head detection network is replaced with a distribution shift convolution. Next, the normalized Gauss-Wasserstein distance is introduced into the detection head's regression loss function, replacing the original model's intersection-over-union (IoU) calculation to enhance the detection of small objects. After feature fusion, the information is fed into the reparameterized convolution (RepConv) inference module in the head network (Head). Multiple feature information transmitted from the neck network is weighted and fused into a single feature map. Finally, three detection heads (IDetect) detect large, medium, and small-scale defects in the electroluminescent images of photovoltaic panels using feature maps of different scales.
[0084] Distributed Shift Convolution (DSConv) simulates the behavior of convolutional layers by using quantization and distributed shift, decomposing the traditional convolution kernel into two components: variable quantized kernel (VOK) and distributed shift. Distributed Shift Convolution stores integer values through a variable quantized kernel, and then maintains the same output as the original convolution through distributed shifts of the kernel and channels.
[0085] The variable quantization kernel (VQK) is the quantization component of DSConv, which only stores integer values of variable bit length. And VQK is the same as the original convolution tensor (ch0,ch i ,k i ,k j ), where ch0 is the number of channels in the next layer, ch i is the number of channels in the current layer, k i ,k j are the width and height of the kernel, respectively. The parameter values are set to be quantized from the original floating-point model and cannot be changed once set. As the quantization component of DSConv, VQK allows for faster and memory-efficient multiplication.
[0086] The distribution shift component mimics the distribution of the original convolution kernel by changing the distribution of VQK. This is achieved by using two tensors to shift in two domains. “Shift” refers to the scaling and biasing operations. The first tensor is the kernel distribution shifter (KDS), which shifts the distribution in each (1, BLK, 1, 1) slice of VQK, where BLK is a hyperparameter of the block size. The second tensor is the channel distribution shifter (CDS), which shifts the distribution in each channel, which changes the distribution of each (1, ch i ,k i ,k j ) distribution in slices.
[0087] Distributed shift convolutions offer plug-and-play functionality and can be directly applied to any convolutional neural network. This paper primarily improves distributed shift convolutions in the YOLOv7 head network by replacing four convolutional modules in the ELAN-W module with distributed shift convolutions. This results in the new ELAN-W-DS module. Distributed shift convolutions improve computational complexity by quantizing weights and finding the optimal distributed shift for fixed integer weights, significantly improving model speed and reducing the number of parameters.
[0088] Introducing normalized Gauss-Wasserstein distance: In photovoltaic panel defect images, in order to address the problem that traditional IoU metric calculations are sensitive to small target position deviations, it is proposed to introduce normalized Gauss-Wasserstein distance (NWD) in the regression loss function to measure the similarity between the predicted bounding box and the true target bounding box. The original IoU metric calculates the degree of overlap. Since the defects of photovoltaic panels are small and some defects are blurry, a slight position deviation will lead to a significant drop in the IoU value, and the use of NWD can reduce this sensitivity. In order to better describe the weights of different pixels in the bounding box, the paper models the bounding box as a two-dimensional Gaussian distribution and uses NWD to measure the similarity of the derived Gaussian distribution. For the horizontal bounding box R = (c x ,c y ,w,h), where (c x ,c y ) are the coordinates of the center point, w and h are the width and height respectively, and R is modeled as a Gaussian distribution N(μ,Σ), where:
[0089]
[0090] Where μ and ∑ are the mean vector and covariance matrix of the Gaussian distribution respectively.
[0091] Then use Wasserstein distance to calculate the bounding box R1 = (c x1 ,c y1 ,w1,h1) and bounding box R2=(c x2 ,c y2,w2,h2) between two Gaussian distribution distances:
[0092]
[0093] Since this distance metric cannot directly measure the similarity between bounding boxes R1 and R2, it is normalized in exponential form to obtain a new metric L NWD :
[0094]
[0095] Where C is a constant closely related to the dataset.
[0096] In this paper, the NWD metric is embedded in the bounding box regression loss function, replacing the original IoU metric to improve the accuracy of small object detection. The original loss function in YOLOv7 is CIoU. CIoU measures the difference between the predicted and true bounding boxes and is very sensitive to the positional offset of small objects. Even small offsets of the target object will affect the IoU. The NWD metric, on the other hand, is insensitive to objects of different scales and is not easily affected by the positional deviations of small objects. Replacing the IoU metric with the NWD metric not only achieves the goal of high-precision small object detection, but also does not affect the convergence speed of training.
[0097] S3. Input the newly acquired image data into the trained model for defect detection and output the detection results.
[0098] The composition and functions of each module in the implementation are as follows:
[0099] like Figure 3 As shown, the CBS module consists of a three-layer network structure: a convolutional (Conv) layer, a batch normalization (BN) layer, and a SiLU activation function layer. The different suffixes in the CBS modules in the figure represent different strides and kernel sizes. The CBS-Y module uses 1x1 convolutions with a stride of 1, primarily for changing the number of channels; the CBS-B module uses 2x2 convolutions with a stride of 1, primarily for feature extraction; and the CBS-P module uses 3x3 convolutions with a stride of 2, primarily for downsampling. This layered design allows the module to efficiently perform different functions and operations while maintaining a compact structure.
[0100] The structure of the CBM module is roughly the same as that of CBS. Compared with CBS, only the activation function has changed. It consists of a Conv convolution layer, a BN layer, and a sigmoid layer. It is a 1X1 convolution with a stride of 1.
[0101] The REP module has different module compositions in the train and deploy modes. Figure 3As shown in the figure, in training mode, the REP module contains three branches. The top layer consists of a 3x3 convolution and a BN layer, which is mainly used for feature extraction; the middle layer consists of a 1x1 convolution and a BN layer, which is used to smooth features; the bottom layer is an identity, which does not perform convolution operation and is directly moved over. Finally, the three branches are added together.
[0102] In inference mode, the REP structure consists of a 3x3 convolution with a stride of 1 and a batch normalization layer. This structure is obtained from the training module through reparameterization. In training mode, the original structure includes a 3x3 convolution in the upper layer, a 1x1 convolution in the middle layer, and an identity (identity mapping) in the lower layer. During the model reparameterization process, the 1x1 convolution in the middle layer and the identity in the lower layer are converted to 3x3 convolutions and matrix fusion is performed. By superimposing the weights of these layers, a new 3x3 convolution is finally formed, which is actually the result of the addition of three branches. This design optimizes the computational efficiency of the model and simplifies the model structure, making it more suitable for inference tasks.
[0103] like Figure 3 As shown in the figure, the MP module consists of two main branches, whose primary function is to perform downsampling. The upper branch first downsamples through a maxpool layer, followed by a 1x1 convolution layer to adjust the number of channels. The lower branch begins with a 1x1 convolution layer to change the number of channels, followed by a 3x3 convolution with a stride of 2 for further downsampling. Finally, the downsampled results of the two branches are concatenated to form the final output. This structure allows the module to effectively integrate features from different branches while reducing the data dimension.
[0104] The ELAN module is an efficient network structure design that aims to enhance learning ability and robustness by controlling the shortest and longest gradient paths in the network; this design allows the network to capture richer features. Figure 3 As shown in Figure 1, ELAN modules have different structural configurations in different parts of the network, namely the Backbone and Neck. This differentiated design allows the modules to optimize performance and functionality based on the network stage they are in.
[0105] ELAN has two branches. In Backbon, the upper branch is a 1x1 convolution, which changes the number of channels. The lower branch first performs a 1x1 convolution, also used to change the number of channels, followed by four 3x3 feature extraction convolutions, and finally concatenates the four features to produce the final result. The difference in Neck is the lower branch, which has two additional feature outputs compared to ELAN in Backbon, resulting in a final result that is the concatenation of six features.
[0106] The main function of the SPPCSPC module is to obtain different receptive field sizes through the maxpool operation, so as to expand the overall receptive field of the module and enable the algorithm to adapt to images of different resolutions. Figure 3 As shown in the figure, the upper branch of the module has four maxpool branches, with pooling sizes of 5, 9, 13, and 1. These four different pooling sizes represent four different receptive fields, helping to distinguish large and small objects in the image. The lower branch performs conventional convolution processing. Finally, the outputs of these two branches are combined through the concat operation, which not only reduces computational effort but also helps speed up detection.
[0107] To make the objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the present invention are described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the present invention is not limited to the following embodiments, and specific implementation methods can be determined based on the technical solutions of the present invention and actual conditions. To avoid obscuring the essence of the present invention, well-known methods, processes, and procedures are not described in detail.
[0108] The above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art will appreciate that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention, and such modifications or equivalents shall be encompassed by the claims of the present invention. Any techniques, shapes, and structures not described in detail herein are well known.
Claims
1. A photovoltaic panel defect detection method based on improved YOLOv7, characterized in that: The detection method specifically includes: S1. Filter and enhance the collected EL images of photovoltaic panels to identify defective EL images. The defects in the images are classified into seven categories: cracks, black cores, dislocations, failures, thick lines, broken grids, and star-shaped cracks. Defects are then annotated on the EL images to create a dataset of defective EL images of photovoltaic panels. Finally, the dataset of defective EL images of photovoltaic panels is divided into a training set, a test set, and a validation set in proportion. S2. Using the YOLOv7 network structure as the basic framework, the photovoltaic panel defect electroluminescence image dataset established in step S1 is used for training, verification, and testing to obtain a photovoltaic panel defect detection model. During the training process, the SimAM attention mechanism is introduced into the YOLOv7 model to improve the model's attention to positive samples and enhance its key feature extraction capabilities. The MP structure is optimized, and Space-to-depth-Conv is introduced to improve the small target detection capability of the YOLOv7 model. The head network of the YOLOv7 model is improved through distributed shift convolution. The normalized Gauss-Wasserstein distance is introduced into the regression loss function of the improved YOLOv7 model to measure the similarity between the predicted bounding box and the true target bounding box. S3. Input the newly collected photovoltaic panel electroluminescence image data into the photovoltaic panel defect detection model to perform defect detection and output the detection results.
2. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In step S1, unclear or severely distorted electroluminescent images of photovoltaic panels are screened and eliminated, and then the screened original images are mirror-flipped, rotated, brightness-adjusted, and contrast-adjusted, and the image data is divided into positive and negative samples with defects and without defects in proportion.
3. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In the step S1, LabelImg software is used to label the defective electroluminescent images of photovoltaic panels after screening and data enhancement processing; the photovoltaic panel defective electroluminescent image dataset is divided into a training set, a test set, and a validation set in a ratio of 7:2:
1.
4. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In step S2, the YOLOv7 network structure is used as the basic framework, and the photovoltaic panel defect electroluminescence image dataset established in step S1 is used for training, verification, and testing to obtain the photovoltaic panel defect detection model. The specific process is as follows: S201. Convert images from the photovoltaic panel defect electroluminescence image dataset into RBG three-channel images and feed them into the backbone network of the YOLOv7 model. The first four convolution layers of the network perform initial feature extraction and downsampling. After four convolution layers, the output convolution image enters the ELAN structure for feature information extraction. The simAM attention mechanism is inserted into the transmission path to enhance its feature information extraction capabilities. S202. Introducing SPD-Conv into the MP structure of the model, replacing the downsampling branch in the original structure with SPD-Conv, followed by normalization and connecting the Silu activation function, while retaining the pooling branch in the original structure. The extracted feature information is fed into the MP structure reconstructed based on the SPD-Conv module for downsampling and pooling operations. Three sets of ELAN and MP with different gradients are used to extract feature information at different levels. Finally, the multi-scale features are fused in the SPC pooling layer, and feature images of different scales are output to the neck network of the YOLOv7 model. The feature images output by ELAN-SA are downsampled by MP to reduce feature information loss. S203. The feature image processed by S202 is transmitted to the ELAN-W module based on distribution shift convolution reconstruction in the neck network and then participates in the subsequent MP downsampling process to quantize the weights and find the optimal distribution shift of fixed integer weights to improve the computational complexity; S204. After feature fusion, the feature image is input into the reparameterized convolution inference module in the head network of the YOLOv7 model. During the inference stage, the reparameterized convolution fuses the multiple feature information transmitted by the neck network into a single feature map based on the weights. Finally, the three detection heads of the head network detect large, medium, and small scale defect targets in the electroluminescent image of the photovoltaic panel based on feature maps of different scales.
5. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In the S2 step, SimAM is an attention mechanism based on the local self-similarity of the feature map. It dynamically adjusts the weight of each pixel by calculating the similarity between each pixel in the feature map and its surrounding pixels, thereby enhancing important features and suppressing irrelevant features. The formula of the SimAM attention mechanism is as follows: (1) Derivation of energy function SimAM measures the importance of the t-th neuron by minimizing the following energy function: Where t is the spatial position of the target neuron; μ t is the eigenvalue (original input) of the target neuron t; is the weighted eigenvalue of the target neuron t (the variable to be optimized); is the weighted variance of the target neuron t (the variable to be optimized); λ is the regularization coefficient (a hyperparameter to prevent the denominator from being zero); (2) Energy minimization solution By minimizing the energy function e t , get the optimal and Where N is the total number of neurons in the same channel (spatial dimension); x i is the eigenvalue of the i-th neuron; (3) Attention weight calculation Final attention weight w t By normalizing the energy function, we can get: Where w t is the attention weight of the t-th neuron (the lower the energy, the higher the weight); (4) Feature enhancement formula Enhance the input features: Where x t is the original input feature; The enhanced feature sigmoid is a nonlinear activation function (limiting the weights between 0 and 1).
6. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In step S2, the SimAM attention mechanism is added to the ELAN module of the YOLOv7 main network, and the SA module is added after the SiLU function of the CBS module at the end of the ELAN module to form the ELAN-SA module. The introduction of the SA attention mechanism enables the network to perform feature recalibration, allowing the model to pay more attention to the channel features that are important to the information. The MP module consists of two main branches, whose main function is to perform downsampling; the upper branch first downsamples through a maximum pooling layer, and then uses a 1x1 convolution layer to adjust the number of channels; the lower branch starts with a 1x1 convolution layer to change the number of channels, and then applies a 3x3 convolution with a stride of 2 for further downsampling; finally, the downsampled results of the two branches are merged together to form the final output.
7. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 1, characterized in that: In the step S2, the normalized Gauss-Wasserstein distance is introduced into the regression loss function of the improved YOLOv7 model to measure the similarity between the predicted bounding box and the true target bounding box. NWD is used to measure the similarity of the derived Gaussian distribution. For the horizontal bounding box R=(c x ,c y ,w,h), where (c x ,c y ) are the coordinates of the center point, w and h are the width and height respectively, and R is modeled as a Gaussian distribution N(μ,Σ), where: Where μ and ∑ are the mean vector and covariance matrix of the Gaussian distribution respectively. Then use Wasserstein distance to calculate the bounding box R1 = (c x1 ,c y1 ,w1,h1) and bounding box R2=(c x2 ,c y2 ,w2,h2) between two Gaussian distribution distances: Since this distance metric cannot directly measure the similarity between bounding boxes R1 and R2, it is normalized in exponential form to obtain a new metric L NWD : Where C is a constant closely related to the dataset.
8. The photovoltaic panel defect detection method based on improved YOLOv7 according to claim 4, characterized in that: The ELAN module is an efficient network structure design. ELAN has two branches. In Backbon, the upper branch is a 1x1 convolution, which is responsible for changing the number of channels. The lower branch first performs a 1x1 convolution, which is also used to change the number of channels. It is followed by four 3x3 feature extraction convolutions and finally concats the four feature information to obtain the final result. The difference in Neck is that the lower branch has two more feature outputs than the ELAN in Backbon, and the final result is the concat sum of 6 feature information.
Citation Information
Patent Citations
Method and device for detecting defects of photovoltaic cell based on improved Yolov7 network model
CN116797582A
Photovoltaic EL assembly defect detection method based on unmanned aerial vehicle intelligent inspection and acquisition
CN117372895A
Indoor article detection method based on YOLOv7
CN117789012A
Cited By
Photovoltaic panel surface defect detection method and system based on improved D-FINE algorithm
CN121190852A