A YOLO-PV model for detecting defects of a photovoltaic cell assembly and a construction method thereof
Patent Information
- Application Number
- CN202311505439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-11-13
AI Technical Summary
[0014]本发明的目的之一就是提供一种检测光伏电池组件缺陷的YOLO-PV模型,以解决模型对小目标的检测能力变弱、无法根据输入图像的特点进行自适应特征提取以及模型因参数量较大而导致检测速度较慢的问题
[0045]S2.获取有缺陷的光伏电池的图片样本,对有缺陷的光伏电池的图片样本进行预处理;
Smart Images

Figure CN117592518B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a target detection model, specifically a YOLO-PV model for detecting defects in photovoltaic cell modules and its construction method. Background Technology
[0002] In recent years, under the strategic backdrop of carbon peaking and carbon neutrality, my country's photovoltaic (PV) power generation technology has developed rapidly. Statistics show that my country's installed PV capacity is expected to reach 350GW by 2030. Photovoltaic modules are the core components that convert solar energy into electrical energy. Besides inherent material defects, various defects can arise during the manufacturing, transportation, and use of the modules. Multiple processing steps on automated production lines also increase the damage rate of the cells, leading to defects such as cracks, broken grids, black cores, thick lines, and misalignment. These defects reduce the photoelectric conversion efficiency, lifespan, and reliability of the modules to varying degrees, thus affecting the quality of the finished product and the reputation of the manufacturing company. Therefore, defect detection and maintenance of photovoltaic cells are crucial to maximizing benefits. Many defects in photovoltaic modules are often invisible to the naked eye and require appropriate detection technologies for rapid and accurate detection.
[0003] Object detection, or object detection, aims to identify all objects of interest in an image and determine their category and location. It is one of the core problems in computer vision. Due to the different appearances, shapes, and poses of various objects, coupled with interference from factors such as lighting and occlusion during imaging, object detection has always been one of the most challenging problems in computer vision. Defect detection is a specific application of object detection and an important problem in industry. Early methods for detecting surface defects on solar cells, based on deep learning, were proposed by Wang Xianbao, Li Jie, and others. Deep learning has attracted widespread attention due to its powerful feature extraction capabilities from input sample data. This method involves constructing a multi-layer neural network to extract image features, using the backpropagation (BP) algorithm to optimize parameters, and then using the comparison between the reconstructed image and the defect image to achieve defect detection on the test sample.
[0004] With the development of deep learning, the YOLO (You Only Look Once) series, first proposed by Joseph Redmon et al., mainly treats object detection as a regression problem, and can simultaneously predict the location and category of an object in a single neural network. The YOLO series of single-stage detection algorithms extracts features based on region regression, offering higher detection speeds compared to traditional two-stage detection algorithms, making it more suitable for real-time detection tasks.
[0005] YOLOv1 treats detection as a regression problem, using only a single neural network to simultaneously predict the location and class of bounding boxes, making it very fast. Because it doesn't require region proposal extraction but performs detection directly on the entire image, YOLOv1 can incorporate contextual information and features, reducing errors such as detecting background as objects. The YOLOv1 network structure (e.g., ...) Figure 1 The design is quite simple, borrowing from GoogLeNet, and consists of 24 convolutional layers and 2 fully connected layers. For the input image, since the network requires two fully connected layers at the end, and these layers need a fixed-size input, the input size needs to be fixed. The network structure primarily uses 1×1 convolutions for dimensionality reduction, followed by 3×3 convolutions. Leaky ReLU activation is used for the convolutional and fully connected layers, and a linear activation function is used for the final layer.
[0006] Although YOLOv1 has a simple structure and is based on a single regression across the entire image, making it much faster than existing object detectors and achieving good real-time performance, its ability to predict nearby objects is limited because it can only detect a maximum of two objects of the same class in a grid cell. It struggles to predict objects with aspect ratios not present in the training data. Furthermore, due to its design containing only downsampling layers, YOLOv1 can only learn from coarse object features, resulting in significant localization errors.
[0007] To address this, Joseph Redmon et al. proposed the YOLOv2 network structure (such as...). Figure 2 This update incorporates several improvements over the original YOLOv1, maintaining the same speed while offering enhanced detection capabilities, enabling the detection of 9000 categories. Specific improvements include: adding batch normalization to all convolutional layers to improve convergence; pre-training the classification model on 224×224 images followed by 10 fine-tuning iterations on 448×448 images to improve performance on high-resolution inputs; selecting better prior anchor boxes based on K-means clustering, these anchor boxes having predefined shapes to match the prototype shapes of objects; employing a direct location prediction method, where the network predicts five bounding boxes for each detection region, each with five values: tx, ty, tw, th, and to. Here, tx represents the offset coordinate of the detection box's center point relative to the left side of the image, ty represents the offset coordinate of the detection box's center point relative to the top side of the image, tw represents the ratio of the detection box's width to the image width, th represents the ratio of the detection box's height to the image height, and to represents the confidence level; and one pooling layer was removed to capture more fine-grained features.
[0008] Although YOLOv2 improved detection accuracy, it was still insufficient for subsequent industrial applications. Therefore, Joseph Redmon et al. proposed YOLOv3 (such as...). Figure 3 YOLOv3 has a larger feature extractor, Darknet-53, consisting of 53 convolutional layers with Res residual connections, which increases the network depth and improves the network's spatial representation capabilities. Similar to the Feature Pyramid Network, YOLOv3 predicts three bounding boxes at three different scales. Starting with YOLOv3, the detector's structure is described as three parts: Backbone, Neck, and Head.
[0009] To investigate whether using different structures in the three-part object detector first proposed in YOLOv3 could result in faster training speed and higher accuracy, Alexey Bochkovskiy et al. proposed YOLOv4. They replaced the backbone with CSPDarknet-53 (adding multiple residual connections); replaced the Neck with PANet (performing two upsampling operations followed by two downsampling operations, with feature fusion along the way); and employed Mosaic data augmentation (stitching four images together as training samples). While YOLOv5 lacks a scientific paper, it can be understood as an improvement and implementation of the YOLOv4 network. YOLOv5 is the fifth generation of one-stage object detection algorithms, and can be divided into five versions based on the number of parameters: YOLOv5n (nanoscale), YOLOv5s (small), YOLOv5m (medium), YOLOv5l (large), and YOLOv5x (extra-large).
[0010] YOLOv5 model (e.g.) Figure 4The system mainly consists of three parts: a backbone network, a neck network, and a head network. The backbone network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer. Each of the second, third, fourth, and fifth convolutional layers includes a two-dimensional convolutional module and a C3 module. The backbone network is used to extract features from the input image. The neck network includes an upsampling module and a downsampling module. The upsampling module includes, in sequence, a convolutional module, a first upsampling layer, a first stitching layer, a first C3 module, a convolutional module, a second upsampling layer, a second stitching layer, and a second C3 module. The downsampling module includes, in sequence, a convolutional module, a third stitching layer, a third C3 module, a convolutional module, a fourth stitching layer, and a fourth C3 module. The neck network receives the feature maps extracted by the backbone network, performs multi-scale feature fusion on the feature maps, and passes these features to the prediction head of the head network. The prediction head is used to detect targets of different sizes and determine the detection results. During the training process of the model, the classification loss, target loss and localization loss are optimized simultaneously by constructing an objective function, so that the predicted bounding box is closer and closer to the ground truth bounding box.
[0011] YOLOv5 is the fifth-generation one-stage object detection algorithm. This algorithm has a small model size and excellent performance, achieving the best balance between detection accuracy and speed in many application scenarios. It has been used in many research fields and has achieved good experimental results.
[0012] Currently, defect detection in photovoltaic cells faces challenges such as the presence of small, medium, and large defects in the datasets being tested, and the inability of existing models to fully meet the practical requirements for defect detection accuracy and efficiency. For example, small targets such as broken grids and thick lines account for approximately 60% of the defect image datasets for photovoltaic cells, while the YOLOv5 detection head selects features from the last three layers, with feature map sizes of 80×80 pixels, 40×40 pixels, and 20×20 pixels, making it difficult to accurately extract features from small targets.
[0013] Furthermore, the Mosaic data augmentation used in the YOLOv5 model stitches four images together into one, then randomly scales and crops them as training samples to enrich the training set. This makes already small targets appear even smaller, weakening the model's ability to detect small targets. Since the convolutional kernels in ordinary convolutional neural networks are static, all kernel parameters are fixed after training. For any input, all kernels are processed equally, failing to adaptively extract features based on the characteristics of the input image, thus limiting the model's generalization ability. As the model becomes more complex, the C3 module in the YOLOv5 network structure, while effectively helping the model extract features, also introduces a large number of parameters, resulting in slower detection speeds and limited applications. In some real-world applications, such as mobile or embedded devices, such a large and complex model is difficult to apply. Summary of the Invention
[0014] One of the objectives of this invention is to provide a YOLO-PV model for detecting defects in photovoltaic cell modules, in order to solve the problems of the model's weakened ability to detect small targets, its inability to perform adaptive feature extraction based on the characteristics of the input image, and the slow detection speed caused by the large number of parameters.
[0015] The second objective of this invention is to provide a method for constructing a YOLO-PV model for detecting defects in photovoltaic cell modules, so as to establish an effective defect detection model for photovoltaic cell modules.
[0016] One of the objectives of this invention is achieved as follows:
[0017] A YOLO-PV model for detecting defects in photovoltaic cell modules includes:
[0018] The backbone network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer. Each of the second to fifth convolutional layers includes two full-dimensional dynamic convolutional modules. Adjacent full-dimensional dynamic convolutional modules in the backbone network are connected through residuals. The backbone network is used to extract features from the input image.
[0019] The neck network includes an upsampling module and a downsampling module; the upsampling module includes a first C3 Gaussian module, a second C3 Gaussian module, and a third C3 Gaussian module, used to upsample and concatenate the feature maps output by the backbone network; the downsampling module includes a fourth C3 Gaussian module, a fifth C3 Gaussian module, a sixth C3 Gaussian module, and a seventh C3 Gaussian module, used to downsample and concatenate the feature maps output by the upsampling module; and
[0020] The head network, connected to the neck network, is used to detect the feature maps output by the neck network.
[0021] The head network includes:
[0022] The first prediction head is used to detect the feature map output by the fourth C3 Gaussian module of the neck network;
[0023] The second prediction head is used to detect the feature map output by the sixth C3 Gaussian module of the neck network; and
[0024] The third prediction head is used to detect the feature map output by the seventh C3 Gaussian module of the neck network.
[0025] Furthermore, the backbone network includes five convolutional layers, namely the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer. The second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer each include two full-dimensional dynamic convolutional modules.
[0026] This invention significantly improves the effectiveness and accuracy of feature extraction by combining residual networks and full-dimensional dynamic convolutions and applying attention weights to the convolution kernels.
[0027] Furthermore, the upsampling module also includes three ConvModule_2 layers, a first upsampling layer, a second upsampling layer, a third upsampling layer, a first stitching layer, a second stitching layer, and a third stitching layer;
[0028] The first C3 Gaussian module is used to extract features from the feature map output by the fifth convolutional layer of the backbone network, and the first concatenation layer is used to concatenate the feature map output by the fourth convolutional layer and the feature map that has passed through the first C3 Gaussian module, ConvModule_2 and the first upsampling layer in sequence.
[0029] The second C3 Gaussian module is used to extract features from the feature map output by the first concatenation layer. The second concatenation layer is used to concatenate the feature map output by the third convolutional layer and the feature map that has passed through the second C3 Gaussian module, ConvModule_2 and the second upsampling layer in sequence.
[0030] The third C3 Gaussian module is used to extract features from the feature map output by the second concatenation layer. The third concatenation layer is used to concatenate the feature map output by the second convolutional layer and the feature map that has passed through the second C3 Gaussian module, ConvModule_2 and the third upsampling layer in sequence.
[0031] Furthermore, the downsampling module also includes three ConvModule_2 layers, a fourth stitching layer, a fifth stitching layer, and a sixth stitching layer;
[0032] The fourth C3 Gaussian module is used to extract features from the feature map output by the third concatenation layer. The fourth concatenation layer is used to concatenate the feature map output by the third convolutional layer, the feature map that has passed through the third C3 Gaussian module and ConvModule_2 in sequence, and the feature map that has passed through the fourth C3 Gaussian module and ConvModule_2 in sequence.
[0033] The fifth C3 Gaussian module is used to extract features from the feature map output by the fourth concatenation layer. The fifth concatenation layer is used to concatenate the feature map output by the fourth convolutional layer, the feature map that has passed through the second C3 Gaussian module and ConvModule_2 in sequence, and the feature map that has passed through the fifth C3 Gaussian module and ConvModule_2 in sequence.
[0034] The sixth C3 Gaussian module is used to extract features from the feature map output by the fifth stitching layer. The sixth stitching layer is used to stitch together the feature maps that have passed through the first C3 Gaussian module and ConvModule_2 in sequence and the feature maps that have passed through the sixth C3 Gaussian module and ConvModule_2 in sequence. The seventh C3 Gaussian module is used to extract features from the feature map output by the sixth stitching layer.
[0035] This invention performs multi-scale bidirectional feature fusion from four layers of feature maps, thereby enhancing the model's ability to identify small targets and improving detection performance in photovoltaic cell defect datasets.
[0036] Furthermore, each C3 Gaussian module includes: two ConvModule_2s, a seventh concatenation layer, and a Gaussian neck network; the input feature map of each C3 Gaussian module flows to two branches, one branch passes through ConvModule_2, and the other branch passes through ConvModule_2 and the Gaussian neck network in sequence. The seventh concatenation layer is used to concatenate the feature maps output by the two branches.
[0037] The C3 Gaussian module reduces the number of parameters while retaining useful feature map information from redundant information, thereby achieving a lightweight network model that balances speed and accuracy.
[0038] Furthermore, the Gaussian neck network includes: a depthwise separable convolutional module, a Gaussian convolutional module, and a stacking layer. The input feature maps of the Gaussian neck network all flow to two branches. One branch passes through the depthwise separable convolutional module and then through the Gaussian convolutional module, and the other branch passes through the Gaussian convolutional module, then through the depthwise separable convolutional module, and finally through the Gaussian convolutional module. The stacking layer is used to stack the feature maps output from the two branches.
[0039] The second objective of this invention is achieved as follows:
[0040] A method for constructing a YOLO-PV model for detecting defects in photovoltaic cell modules includes the following steps:
[0041] S1. Replace the two-dimensional convolutional modules and C3 module of the second, third, fourth and fifth convolutional layers in the YOLOv5 model with full-dimensional dynamic convolutional modules, and perform residual connections on adjacent full-dimensional dynamic convolutional modules;
[0042] In the upsampling module of the YOLOv5 model, a two-dimensional convolutional module, a third upsampling layer, and a fifth concatenation layer are added above the second C3 module in sequence. A C3 Gaussian module is added at the connection between the two-dimensional convolutional module and the spatial pyramid pooling module in the YOLOv5 model to connect the second convolutional layer and the fifth concatenation layer. The connection between the second C3 module and the two-dimensional convolutional module is removed.
[0043] In the downsampling module of the YOLOv5 model, add a C3 Gaussian module connected to the fifth stitching layer, a two-dimensional convolutional module connected to the C3 Gaussian module, a sixth stitching layer connected to the two-dimensional convolutional module, and a C3 Gaussian module connected to the sixth stitching layer. Connect the third convolutional layer and the sixth stitching layer. In the upsampling module, add a two-dimensional convolutional module and connect it to the sixth stitching layer. Connect the fourth convolutional layer and the third stitching layer.
[0044] Replace the C3 module of the neck network in the YOLOv5 model with a C3 Gaussian module; this yields the YOLO-PV model.
[0045] S2. Obtain image samples of defective photovoltaic cells and preprocess the image samples of defective photovoltaic cells;
[0046] S3. Input the preprocessed image samples into the YOLO-PV model for training to obtain the trained YOLO-PV model;
[0047] S4. Validate and evaluate the trained YOLO-PV model.
[0048] This invention constructs a network structure, Res-ODCnet, with adaptive feature extraction capabilities by combining residual connections and a full-dimensional dynamic convolution module. This structure serves as the backbone of the model. The full-dimensional dynamic convolution adaptively calculates different convolution kernel parameters for different inputs, and adaptively extracts features based on the input image. This avoids the problem of limited generalization ability caused by static convolution kernels processing all inputs equally. Furthermore, the residual connections sum the outputs of shallow and deep layers as the input for the next stage, thus avoiding gradient vanishing and gradient exploding. By combining full-dimensional dynamic convolution and residual connections and applying attention weights to the convolution kernels, the feature extraction effect is improved.
[0049] This invention constructs a feature fusion architecture FB-FPN with accurate small target detection capabilities and uses this architecture to perform multi-scale bidirectional feature fusion. For small, medium, and large targets, detection heads are designed by fusing information from four layers of feature maps. Features are extracted and stitched together from 160×160 pixel feature maps in the backbone and neck networks. This improves the detection capability of small targets while maintaining the detection performance of medium and large targets, thus solving the problem of the model's weakened detection capability for small targets.
[0050] This invention designs a structure by combining a Gaussian convolution module with residual connections and a depthwise separable convolution module, and integrates it with the C3 module of the YOLOv5 model to form a lightweight C3 Gaussian module that retains useful information from redundant information while reducing computational cost. This reduces the number of parameters, thereby retaining useful feature map information from redundant information to achieve a lightweight network model, balancing speed and accuracy, and ultimately improving its robustness and generalization ability during training. Attached Figure Description
[0051] Figure 1 This is a structural diagram of the YOLOv1 model.
[0052] Figure 2 This is a structural diagram of the YOLOv2 model.
[0053] Figure 3 This is a structural diagram of the YOLO3 model.
[0054] Figure 4 This is a structural diagram of the YOLOv5 model.
[0055] Figure 5 This is a structural diagram of the YOLO-PV model used to detect defects in photovoltaic cell modules.
[0056] Figure 6 This is a diagram of the process of full-dimensional dynamic convolution.
[0057] Figure 7 This is a diagram of the residual connection process.
[0058] Figure 8 This is a network structure diagram of Res-ODCnet.
[0059] Figure 9 This is the structure diagram of the C3 Gaussian module.
[0060] Figure 10 This is a flowchart of the Gaussian convolution module.
[0061] Figure 11 This is a diagram of the FB-FPN architecture.
[0062] Figure 12 This is a comparison chart of heatmaps of the YOLO-PV model and the YOLOv5 model.
[0063] Figure 13 This is a curve comparing the average accuracy of the YOLO-PV model and the YOLOv5 model.
[0064] Figure 14 This is a comparison chart of the YOLO-PV model with other models in terms of parameter count, computational cost, and accuracy.
[0065] Figure 15 This is a bar chart showing the average detection accuracy of the YOLO-PV model for each specific type of defect.
[0066] Figure 16 This is a screenshot of the defect detection results from the YOLO-PV model. Detailed Implementation
[0067] The technical solution of the present invention will be clearly and completely described below with reference to examples. The described embodiments do not limit the present invention in any way.
[0068] The YOLO-PV model for detecting defects in photovoltaic cell modules and its construction method provided by this invention specifically include the following steps:
[0069] S1. Improve the YOLOv5 model to obtain the YOLO-PV model.
[0070] After training, the YOLOv5 model suffers from fixed convolutional kernel parameters, preventing static convolution from adaptively extracting features based on changes in input data. Therefore, the 2D convolutional modules and C3 module in the second to fifth convolutional layers of the YOLOv5 model are replaced with full-dimensional dynamic convolutional modules. In this embodiment, full-dimensional dynamic convolution applies attention weighting to the input data. The convolutional kernel parameters are no longer fixed values but variables determined by the input. For different inputs, different convolutional kernel parameters are adaptively calculated, significantly improving the effectiveness and accuracy of feature extraction without a significant increase in computational cost or parameter count.
[0071] like Figure 5As shown, the YOLO-PV model includes a backbone network, a neck network, and a head network. The backbone network includes five convolutional layers, namely the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer. The first convolutional layer includes one convolutional module_1 (ConvModule_1) and one max pooling layer. The second, third, and fourth convolutional layers each include two full-dimensional dynamic convolutional modules (ODConv). The fifth convolutional layer includes two full-dimensional dynamic convolutional modules and one spatial pyramid pooling layer.
[0072] like Figure 6 As shown, full-dimensional dynamic convolution refers to learning four complementary attention mechanisms of the convolutional kernel in parallel along all four dimensions of the kernel space (the spatial size of each kernel, the number of input channels, the number of output channels, and the number of kernels) at any convolutional layer. The computation process of full-dimensional dynamic convolution can be represented as follows:
[0073] y=(α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1+…+α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n )*x
[0074] For the input x, the feature vector is reduced in dimensionality using a global average pooling (GAP) layer, a fully connected (FC) layer, and the ReLU activation function. Then, attention weights (calculate weights) are applied to the four dimensions using sigmoid and softmax through four parallel fully connected layers, where W... i Denotes the convolution kernel, α si α represents the weights of the convolution kernel's spatial dimension. ci α represents the input channel dimension weights. fi α represents the weights of the output channel dimensions. wi The number of convolution kernels and their weights are represented by the weights. Finally, the weighted result is summed with the input x to obtain the output y, thereby enhancing the feature extraction capability of the convolution operation.
[0075] As deep neural networks become deeper, in addition to a dramatic increase in computational cost and the number of parameters, problems such as vanishing and exploding gradients also arise, making it impossible to update the parameters of shallower networks. For example... Figure 7As shown, residual connection refers to using the sum of the outputs of the shallow layer and the deep layer as the input of the next stage. The result is that the weights of this layer originally needed to learn a mapping from x to f(x). After using residual connection, the mapping that the weights need to learn changes from x to f(x) + x. In this way, during backpropagation, the gradient with small loss is more likely to reach the shallow neurons, thereby avoiding the problems of gradient vanishing and gradient explosion.
[0076] like Figure 8 As shown, the concepts of dynamic convolution and residual connections are combined, and an image with a resolution of 1 / 32 of the original input image is obtained through stacking. The eight full-dimensional dynamic convolutional modules in the backbone network of the YOLO-PV model are designated as the first, second, third, fourth, fifth, sixth, seventh, and eighth full-dimensional dynamic convolutional modules, and residual connections are used across these eight modules. The first full-dimensional dynamic convolutional module extracts features from the feature maps that have passed through ConvModule_1 and the max pooling layer. The input and output feature maps of the first full-dimensional dynamic convolutional module serve as the input feature maps of the second full-dimensional dynamic convolutional module. The input and output feature maps of each full-dimensional dynamic convolutional module serve as the input feature maps of the next full-dimensional dynamic convolutional module, and so on, up to the eighth full-dimensional dynamic convolutional module. The spatial pyramid pooling layer pools the input and output feature maps of the eighth full-dimensional dynamic convolutional module. Therefore, the backbone network of the YOLO-PV model is a network structure with adaptive feature extraction capabilities (Residual OMNIDIMENSIONAL DYNAMIC CONVOLUTION Network, Res-ODCnet).
[0077] For example, an input image of size 640×640 pixels is processed by the first convolutional layer to obtain a feature map of 320×320 pixels. This feature map is then processed by the second convolutional layer to obtain a feature map of 160×160 pixels. This feature map is then processed by the third convolutional layer to obtain a feature map of 80×80 pixels. This feature map is then processed by the fourth convolutional layer to obtain a feature map of 40×40 pixels. This feature map is then processed by the fifth convolutional layer to obtain a feature map of 20×20 pixels.
[0078] ConvModule_1 includes a regular 2D convolutional layer, a batch normalization layer, and a ReLU activation function. The spatial pyramid pooling layer includes two convolutional modules_2 (ConvModule_2), three max pooling layers, and a concatenation layer. ConvModule_2 includes a regular 2D convolutional layer, a batch normalization layer, and a SiLU activation function.
[0079] To enhance feature extraction for small targets, a 2D convolutional module, an upsampling layer (the third upsampling layer of the YOLOv5 model), and a concatenation layer (the fifth concatenation layer of the YOLOv5 model) are added sequentially above the second C3 module in the upsampling module of the YOLOv5 model. A C3 Gaussian module is added at the connection between the 2D convolutional module and the spatial pyramid pooling module in the YOLOv5 model. The concatenation layer added to the second convolutional layer and the upsampling module is then connected. The second C3 module and... The connection of the 2D convolution module; add a C3 Gaussian module, a 2D convolution module, a concatenation layer (the sixth concatenation layer of the YOLOv5 model) to the downsampling module of the YOLOv5 model, and a C3 Gaussian module connected to the concatenation layer added in the downsampling module; connect the third convolution layer and the concatenation layer added in the downsampling module; connect the 2D convolution module added in the upsampling module and the concatenation layer added in the downsampling module; and connect the fourth convolution layer and the third concatenation layer of the YOLOv5 model.
[0080] The neck network of the YOLO-PV model includes an upsampling module and a downsampling module. The upsampling module includes a C3 Gaussian module, ConvModule_2, an upsampling layer, and a concatenation layer. The concatenation layer of the upsampling module is used to concatenate the feature maps output by the convolutional layers of the backbone network and the feature maps output by the upsampling layer. The downsampling module includes a C3 Gaussian module, ConvModule_2, and a concatenation layer. The concatenation layer of the downsampling module is used to concatenate the feature maps output by the C3 Gaussian module in the upsampling module and the feature maps output by the ConvModule_2 in the downsampling module.
[0081] Assuming the input image size is 640×640 pixels, and small targets are typically less than or equal to 5% of the input image, this invention proposes an architecture that uses four layers of feature maps with different resolutions for multi-scale feature fusion to improve the accuracy of small target detection.
[0082] The upsampling module includes three C3 Gaussian modules (C3BG): the first C3 Gaussian module, the second C3 Gaussian module, and the third C3 Gaussian module. The upsampling module includes a first upsampling layer, a second upsampling layer, a third upsampling layer, a first stitching layer, a second stitching layer, and a third stitching layer. The downsampling module includes a fourth C3 Gaussian module, a fifth C3 Gaussian module, a sixth C3 Gaussian module, a seventh C3 Gaussian module, a fourth stitching layer, a fifth stitching layer, and a sixth stitching layer.
[0083] The YOLOv5 model does not further utilize the 160×160 pixel feature map. By adding an upsampling module, a 160×160 pixel feature map is obtained. Similarly, by adding a downsampling module, a 160×160 pixel feature map is obtained. Furthermore, an improved backbone network and an improved original downsampling module are connected. The detection head in the original head network that detects the 80×80 pixel feature map in the original upsampling module is replaced with a detection head that detects the 160×160 pixel feature map in the improved downsampling module, making the detected feature map richer in features.
[0084] After the neck network receives the 20×20 pixel feature map obtained from the backbone network through 5 convolutional layers, it passes through a first C3 Gaussian module, a ConvModule_2, and a first upsampling layer to obtain a 40×40 pixel feature map. This feature map is then concatenated with the 40×40 pixel feature map output from the fourth convolutional layer in the first concatenation layer. The concatenated feature map then passes through a second C3 Gaussian module, a ConvModule_2, and a second upsampling layer to obtain an 80×80 pixel feature map. This feature map is then concatenated with the 80×80 pixel feature map output from the third convolutional layer in the second concatenation layer. The concatenated feature map then passes through a third C3 Gaussian module, a ConvModule_2, and a third upsampling layer to obtain a 160×160 pixel feature map. This feature map is then concatenated with the 160×160 pixel feature map output from the second convolutional layer in the third concatenation layer. The concatenated feature map then passes through a fourth C3 Gaussian module and a ConvModule_2. e_2, resulting in an 80×80 pixel feature map; this feature map, along with the 80×80 pixel feature map output from the third convolutional layer and the 80×80 pixel feature map that has passed through the third C3 Gaussian module and ConvModule_2 in sequence, are concatenated in the fourth concatenation layer. The concatenated feature map then passes through the fifth C3 Gaussian module and one ConvModule_2 in sequence to obtain a 40×40 pixel feature map; this feature map, along with the 40×40 pixel feature map output from the fourth convolutional layer and the 40×40 pixel feature map that has passed through the second C3 Gaussian module and ConvModule_2 in sequence, are concatenated in the fifth concatenation layer. The concatenated feature map then passes through the sixth C3 Gaussian module and one ConvModule_2 in sequence to obtain a 20×20 pixel feature map. This feature map, along with the 20×20 pixel feature map output from the first C3 Gaussian module and ConvModule_2 in sequence, are concatenated in the sixth concatenation layer. The concatenated feature map then enters the seventh C3 Gaussian module.
[0085] To ensure fast response times, developing small and efficient models is crucial. This invention addresses this issue by combining Gaussian convolution modules with residual connections and depthwise separable convolution modules, and fusing them with the C3 module to propose a lightweight C3 Gaussian module (C3BG module), such as... Figure 9 As shown, each C3 Gaussian module includes two ConvModule_2s, a Gaussian neck network, and a seventh concatenation layer. The input feature map of each C3 Gaussian module flows to two branches. One branch passes through ConvModule_2, and the other branch passes through ConvModule_2 and the Gaussian neck network in sequence. The seventh concatenation layer is used to concatenate the feature maps output from the two branches. The C3 Gaussian module retains key feature map information while reducing the number of parameters, thus achieving a lightweight network model and balancing speed and accuracy.
[0086] The Gaussian neck network consists of a Gaussian convolutional module, a depthwise separable convolutional module, and a stacking layer. The input feature maps of the Gaussian neck network flow to two branches. One branch passes through the depthwise separable convolutional module and then through the Gaussian convolutional module. The other branch passes through the Gaussian convolutional module, then through the depthwise separable convolutional module, and finally through the Gaussian convolutional module. The stacking layer is used to stack the feature maps output from the two branches.
[0087] The Gaussian convolution module consists of two ConvModule_2 modules and a stacking layer. The input feature maps of the Gaussian convolution module pass through the first ConvModule_2 and the second ConvModule_2 in sequence. The feature maps that have passed through the first ConvModule_2 are stacked with the feature maps that have passed through the two ConvModule_2 modules in sequence in the stacking layer.
[0088] The depthwise separable convolutional module consists of two ConvModule_2 modules. The input image of the depthwise separable convolutional module passes through the two ConvModule_2 modules in sequence.
[0089] While the C3 module in the YOLOv5 network architecture effectively helps the model extract features, it also introduces a large number of parameters, resulting in slower detection speeds and limited applications. Therefore, the C3 module is replaced with a C3 Gaussian module. In the Gaussian (Ghost) convolution module, to avoid redundant information in the feature maps generated by multiple convolution operations in traditional convolution, a simple linear transformation is applied to this part of the feature map to simulate the effect of conventional convolution. For example... Figure 10 As shown, the implementation process of the Gaussian convolution module first compresses the channel dimension of the input feature map through a 1×1 convolution, and then performs some simple linear transformations on each feature map, such as... The corresponding feature map is obtained, and then the original convolutional feature map is fused with the linearly transformed feature map.
[0090] The head network consists of at least three detection heads, which detect feature maps transmitted by the neck network.
[0091] The head network includes three prediction heads: the first prediction head, the second prediction head, and the third prediction head. The first prediction head detects and judges the feature map in the neck network after passing through the fourth C3 Gaussian module. That is, the first prediction head detects the 160×160 pixel feature map and is used for the detection of small targets. The second prediction head detects and judges the feature map in the neck network after passing through the sixth C3 Gaussian module. That is, the second prediction head detects the 40×40 pixel feature map and is used for the detection of medium targets. The third prediction head detects and judges the feature map in the neck network after passing through the seventh C3 Gaussian module. That is, the third prediction head detects the 20×20 pixel feature map and is used for the detection of large targets.
[0092] The prediction method of this prediction head is more accurate than the method of the head network which includes four prediction heads and detects the feature maps output by the fourth, fifth, sixth and seventh C3 Gaussian modules respectively.
[0093] Due to the diversity of features in the concentrated samples of photovoltaic cell defects and the fact that the sample categories simultaneously include small, medium, and large targets to be detected, it is necessary to perform bidirectional feature fusion by utilizing information preserved in the low-level high-dimensional feature distribution space and information in the deep low-dimensional feature space.
[0094] like Figure 11 As shown, a layer is added to the original YOLOv5 feature fusion structure, namely the second convolutional layer of the backbone network, the third C3 Gaussian module, ConvModule_2, the third upsampling layer, and the third concatenation layer in the neck network, and the fourth C3 Gaussian module of the downsampling module, resulting in multi-scale 160×160 resolution feature maps. This large-scale feature map is responsible for detecting small targets. The YOLO-PV model performs dynamic convolution, upsampling, and upsampling on the four-layer resolution feature maps, making the feature maps detected by the prediction head richer in features. That is, the feature map simultaneously utilizes the information preserved in the low-layer high-dimensional feature distribution space and the information preserved in the deep high-dimensional feature distribution space to perform bidirectional feature fusion, resulting in a feature fusion architecture (Four-Layer Bidirectional Feature Pyramid Network, FB-FPN) with accurate small target detection capabilities.
[0095] S2. Obtain image samples of defective photovoltaic cells and preprocess the image samples of defective photovoltaic cells.
[0096] Image samples of defective photovoltaic cells are preprocessed, and four images are stitched together by random scaling and cropping using Mosaic data augmentation to achieve a resolution of 640×640, thereby improving the accuracy and robustness of the YOLO-PV model.
[0097] S3. Input the preprocessed image samples into the YOLO-PV model for training to obtain the trained YOLO-PV model.
[0098] The YOLO-PV model was trained using image samples of photovoltaic cell defects, and the parameters of the YOLO-PV model were optimized to obtain the trained YOLO-PV model.
[0099] S4. Validation and evaluation of the trained YOLO-PV model.
[0100] To verify the performance of the YOLO-PV model in photovoltaic cell defect detection, we conducted comparative experiments with other versions of YOLOv5 and other target detection models on laboratory equipment, and evaluated it from different perspectives, including heatmap analysis, quantitative analysis, and visualization analysis.
[0101] Heatmaps can determine the strength of the correlation between variables based on the correlation coefficients corresponding to different colors; generally, darker colors indicate a stronger correlation, and lighter colors indicate a weaker correlation. For example... Figure 12 The image shows a visualization comparing the heatmaps of the YOLO-PV model and the YOLOv5 model. Clearly, by comparing the heatmaps of each model with the original and labeled images, it can be seen that the YOLO-PV model can more accurately focus on the target region in the image, and its feature extraction performance is superior to other models.
[0102] Furthermore, such as Figure 13 As shown in the figure, the average accuracy curves during the training process of the YOLO-PV model and the YOLOv5 model are compared. The YOLO-PV model achieves better results than the YOLOv5 model and has better practicality in the field of photovoltaic cell module defect detection.
[0103] This invention compares the detection accuracy of the YOLO-PV model with that of the YOLOv5 model with different parameters on small, medium and large targets, and uses the average accuracy obtained by averaging the IOU threshold from 0.5 to 0.95 as the evaluation index. Table 1 shows the detection performance of the proposed YOLO-PV model and the comparison model on the photovoltaic cell defect detection data test set.
[0104] Table 1: Comparison of average accuracy for small, medium, and large target detection
[0105]
[0106] According to the information in Table 1, the YOLO-PV model has higher detection accuracy than the original YOLOv5 model for small, medium and large defects in photovoltaic cell modules. The detection accuracy of the YOLO-PV model on targets at all scales is improved compared to the original YOLOv5 model.
[0107] The following are some of the evaluation indicators used in the quantitative analysis of this invention:
[0108] Precision: The accuracy rate of correctly predicting a single class of samples, also known as precision. Higher precision means better recognition performance. Recall: The ratio of correctly identified samples to the total number of samples in the test set, also known as recall. Higher recall means better recognition performance. Map (Average Precision): In multi-class object detection, a curve can be plotted for each class based on recall and precision; AP is the area under this curve. mAP is the average AP across multiple classes, also known as average precision. Specifically, it can be divided into Map@0.5 (Map value when IOU threshold = 0.5) and Map@0.5:0.95 (average MAP value over IOU thresholds from 0.5 to 0.95 with a step size of 0.05). Higher Map value indicates better recognition performance. Parameters: The number of parameters describes the size of the model. FLOPs (Flatwork Flows): The amount of computation measures the complexity of an algorithm or model.
[0109] Table 2: Comparative experimental results of each model on the photovoltaic cell defect dataset.
[0110]
[0111] Table 2 shows the performance of the proposed YOLO-PV model and the comparative models on the photovoltaic cell defect detection data test set. The proposed YOLO-PV model significantly outperforms other models in detecting photovoltaic cell defects, with substantial improvements in accuracy, recall, and mean precision, resulting in superior detection performance. Furthermore, the YOLO-PV model achieves higher accuracy with fewer parameters compared to other models.
[0112] like Figure 14As shown, the specific performance of the YOLO-PV model in defect detection is further visualized. The experimental results of each model based on the photovoltaic cell defect dataset are visualized as follows: the horizontal axis represents computational cost, the vertical axis represents accuracy, and the size of the circle represents the number of parameters of the model (1M = 10). 6 ).according to Figure 14 It can be observed that the YOLO-PV model achieves high recognition accuracy with relatively small parameter and computational costs.
[0113] like Figure 15 As shown, the YOLO-PV model has a high detection accuracy for each specific type of defect. The YOLO-PV model can accurately detect defects in photovoltaic modules.
[0114] To further verify the defect detection performance of the YOLO-PV model, such as Figure 16 As shown, images selected from a photovoltaic cell defect dataset are used to demonstrate the detection performance of the original YOLOv5 model and the improved model, highlighting the model's detection effectiveness. Clearly, the YOLO-PV model proposed in this invention can accurately detect defects in photovoltaic cell modules.
[0115] In summary, the YOLO-PV model proposed in this invention can be used in the defect detection of photovoltaic cell modules, and has high detection accuracy for common defects in photovoltaic cell modules.
Claims
1. A YOLO-PV model for detecting defects in photovoltaic cell modules, characterized in that, include: The backbone network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer. Each of the second to fifth convolutional layers includes two full-dimensional dynamic convolutional modules. Adjacent full-dimensional dynamic convolutional modules in the backbone network are connected through residuals. The backbone network is used to extract features from the input image. The neck network includes an upsampling module and a downsampling module; the upsampling module includes a first C3 Gaussian module, a second C3 Gaussian module, and a third C3 Gaussian module, used to upsample and stitch together the feature maps output by the backbone network; the downsampling module includes a fourth C3 Gaussian module, a fifth C3 Gaussian module, a sixth C3 Gaussian module, and a seventh C3 Gaussian module, used to downsample and stitch together the feature maps output by the upsampling module; as well as The head network, connected to the neck network, is used to detect the feature maps output by the neck network. The head network includes: The first prediction head is used to detect the feature map output by the fourth C3 Gaussian module of the neck network; The second prediction head is used to detect the feature map output by the sixth C3 Gaussian module of the neck network; and The third prediction head is used to detect the feature map output by the seventh C3 Gaussian module of the neck network; Each C3 Gaussian module includes two ConvModule_2 modules, a Gaussian neck network, and a seventh concatenation layer. The input feature map of each C3 Gaussian module flows to two branches. One branch passes through ConvModule_2, and the other branch passes through ConvModule_2 and the Gaussian neck network in sequence. The seventh concatenation layer is used to concatenate the feature maps output by the two branches. The Gaussian neck network includes a Gaussian convolution module, a depthwise separable convolution module, and a stacking layer. The input feature map of the Gaussian neck network flows to two branches. One branch passes through the depthwise separable convolution module and then through the Gaussian convolution module. The other branch passes through the Gaussian convolution module, then through the depthwise separable convolution module, and finally through the Gaussian convolution module. The stacking layer is used to stack the feature maps output by the two branches.
2. The YOLO-PV model for detecting defects in photovoltaic cell modules according to claim 1, characterized in that, The upsampling module also includes three ConvModule_2 layers, a first upsampling layer, a second upsampling layer, a third upsampling layer, a first stitching layer, a second stitching layer, and a third stitching layer; The first C3 Gaussian module is used to extract features from the feature map output by the fifth convolutional layer of the backbone network, and the first concatenation layer is used to concatenate the feature map output by the fourth convolutional layer and the feature map that has passed through the first C3 Gaussian module, ConvModule_2 and the first upsampling layer in sequence. The second C3 Gaussian module is used to extract features from the feature map output by the first concatenation layer. The second concatenation layer is used to concatenate the feature map output by the third convolutional layer and the feature map that has passed through the second C3 Gaussian module, ConvModule_2 and the second upsampling layer in sequence. The third C3 Gaussian module is used to extract features from the feature map output by the second concatenation layer. The third concatenation layer is used to concatenate the feature map output by the second convolutional layer and the feature map output by the third C3 Gaussian module, ConvModule_2 and the third upsampling layer in sequence.
3. The YOLO-PV model for detecting defects in photovoltaic cell modules according to claim 2, characterized in that, The downsampling module also includes three ConvModule_2 layers, a fourth splicing layer, a fifth splicing layer, and a sixth splicing layer; The fourth C3 Gaussian module is used to extract features from the feature map output by the third concatenation layer. The fourth concatenation layer is used to concatenate the feature map output by the third convolutional layer, the feature map that has passed through the third C3 Gaussian module and ConvModule_2 in sequence, and the feature map that has passed through the fourth C3 Gaussian module and ConvModule_2 in sequence. The fifth C3 Gaussian module is used to extract features from the feature map output by the fourth concatenation layer. The fifth concatenation layer is used to concatenate the feature map output by the fourth convolutional layer, the feature map that has passed through the second C3 Gaussian module and ConvModule_2 in sequence, and the feature map that has passed through the fifth C3 Gaussian module and ConvModule_2 in sequence. The sixth C3 Gaussian module is used to extract features from the feature map output by the fifth stitching layer. The sixth stitching layer is used to stitch together the feature maps that have passed through the first C3 Gaussian module and ConvModule_2 in sequence and the feature maps that have passed through the sixth C3 Gaussian module and ConvModule_2 in sequence. The seventh C3 Gaussian module is used to extract features from the feature map output by the sixth stitching layer.
4. A method for constructing a YOLO-PV model for detecting defects in photovoltaic cell modules, characterized in that, Includes the following steps: S1. Replace the two-dimensional convolutional modules and C3 module of the second, third, fourth and fifth convolutional layers in the YOLOv5 model with full-dimensional dynamic convolutional modules, and perform residual connections on adjacent full-dimensional dynamic convolutional modules; In the upsampling module of the YOLOv5 model, a two-dimensional convolutional module, a third upsampling layer, and a fifth concatenation layer are added above the second C3 module in sequence. A C3 Gaussian module is added at the connection between the two-dimensional convolutional module and the spatial pyramid pooling module in the YOLOv5 model to connect the second convolutional layer and the fifth concatenation layer. The connection between the second C3 module and the two-dimensional convolutional module is removed. In the downsampling module of the YOLOv5 model, add a C3 Gaussian module connected to the fifth stitching layer, a 2D convolutional module connected to the C3 Gaussian module, a sixth stitching layer connected to the 2D convolutional module, and a C3 Gaussian module connected to the sixth stitching layer. Connect the third convolutional layer and the sixth stitching layer. In the upsampling module, add a 2D convolutional module connected to the sixth stitching layer. Connect the fourth convolutional layer and the third stitching layer. Replace the C3 module of the neck network in the YOLOv5 model with a C3 Gaussian module; The YOLO-PV model is obtained; S2. Obtain image samples of defective photovoltaic cells and preprocess the image samples of defective photovoltaic cells; S3. Input the preprocessed image samples into the YOLO-PV model for training to obtain the trained YOLO-PV model; S4. Validate and evaluate the trained YOLO-PV model; Each C3 Gaussian module includes two ConvModule_2 modules, a Gaussian neck network, and a seventh concatenation layer. The input feature map of each C3 Gaussian module flows to two branches. One branch passes through ConvModule_2, and the other branch passes through ConvModule_2 and the Gaussian neck network in sequence. The seventh concatenation layer is used to concatenate the feature maps output by the two branches. The Gaussian neck network includes a Gaussian convolution module, a depthwise separable convolution module, and a stacking layer. The input feature map of the Gaussian neck network flows to two branches. One branch passes through the depthwise separable convolution module and then through the Gaussian convolution module. The other branch passes through the Gaussian convolution module, then through the depthwise separable convolution module, and finally through the Gaussian convolution module. The stacking layer is used to stack the feature maps output by the two branches.
Citation Information
Patent Citations
Multi-modal feature target detection method based on dynamic convolution and attention mechanism
CN116452937A
Power transmission line intelligent defect detection method based on improved YOLOv5 network
CN116843649A