A method, system, device and medium for detecting defects in a photovoltaic panel

By combining CLIP and YOLOv11 models, the YOLOv11 layer is initialized using CLIP's multimodal understanding ability and fine-tuning training is performed under a small amount of data, the training difficulty of photovoltaic panel defect detection model is solved when a large amount of labeled data is missing, and efficient and accurate defect detection is achieved.

CN119851046BActive Publication Date: 2025-06-24SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510322309.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The existing photovoltaic panel defect detection model requires a large number of labeled data sets for training, and performs poorly in different real scenarios, especially in remote areas, which cannot obtain a large amount of data, resulting in the model being unable to achieve large-scale and meticulous detection.

Method used

By initializing the YOLOv11 layer with the multimodal understanding ability of CLIP and fine-tuning training on a small amount of target domain data, the problem of model training difficulties caused by the lack of a large amount of labeled data is solved.

Benefits of technology

It improves the detection performance and training efficiency of the photovoltaic defect detection model, can achieve high-precision defect detection under a small amount of data, and is suitable for photovoltaic system maintenance and monitoring in remote areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851046B_ABST
    Figure CN119851046B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for detecting defects in a photovoltaic panel, relating to the technical field of photovoltaic defect detection. The method includes: acquiring an infrared image and a visible light image of a photovoltaic module; extracting its visual features and fusing them; annotating the defective regions in the fused image dataset and saving the annotation information of the defective regions to construct a training dataset; inputting the training dataset into a CLIP model for processing to generate an embedding representation; using the embedding representation to initialize the layer weights of a YOLOv11 model and fine-tuning the YOLOv11 model; inputting the image data of the photovoltaic module to be tested into the trained photovoltaic defect detection model to detect the defective regions in the image data. This method solves the problem that it is difficult to train a model due to the lack of a large amount of data for photovoltaic detection in remote areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic defect detection, and particularly relates to a method, a system, a device and a medium for detecting photovoltaic panel defects. Background Art

[0002] Photovoltaic (PV) systems play a crucial role in the global transition to renewable energy and are a sustainable and environmentally friendly source of electricity. Current advances in deep learning and computer vision have shown great promise in automating the detection of photovoltaic (PV) panel defects. Object recognition algorithms are specifically designed to accurately identify and classify defects in PVs.

[0003] The currently adopted detection method is YOLO object detection: collecting and annotating photovoltaic panel image data, setting up a training environment, configuring the YOLO model and training it, evaluating the model performance through a validation set and optimizing it, testing the model to ensure its accuracy and then deploying it to an actual detection system, and finally continuously monitoring the model performance and providing necessary technical support and maintenance.

[0004] However, PV defect detection must effectively handle a wide range of detections, identify various types of defects, and provide high-precision real-time processing; but the currently adopted models require a large amount of labeled data sets for training, are usually tailored for specific tasks, and cannot exhibit satisfactory performance in different real-world scenarios; especially in remote areas, it becomes more difficult to detect environmental changes and a large amount of data for photovoltaic detection cannot be obtained, resulting in the inability of the trained model to perform large-scale and detailed detections. Summary of the Invention

[0005] Aiming at the problem that a large amount of data for photovoltaic detection cannot be obtained in the prior art, resulting in the inability of the trained model to perform large-scale and detailed detections, the present invention proposes a method, a system, a device and a medium for detecting photovoltaic panel defects, thereby initializing the layers of YOLOv11 by utilizing the multi-modal understanding ability of CLIP and performing fine-tuning training with a small amount of target domain data, solving the problem of difficult model training due to the lack of a large amount of labeled data.

[0006] A method for detecting photovoltaic panel defects includes the following steps:

[0007] Obtaining image data including different illumination conditions, different defect types and different photovoltaic component types; the image data includes a plurality of infrared images and a plurality of visible light images;

[0008] Extract the visual features of each infrared image and visible light image, and fuse all the visual features to obtain a fused image dataset; label the defective areas in the fused image dataset, and save the text information of the defective areas after labeling; construct a training dataset through the fused image dataset and the text information of the defective areas after labeling;

[0009] Input the training dataset into the CLIP model, and use a text editor to process the text information of the defective areas after labeling to generate encoded vectors ; Use an image encoder to process the fused image dataset to generate encoded vectors ; According to the encoded vector and the encoded vector generate an embedded representation; use this embedded representation to initialize the layer weights of the YOLOv11 model, and adjust the first layer of the YOLOv11 model through the initialized layer weights to obtain a photovoltaic defect detection model;

[0010] Fuse the visual features of the infrared image and visible light image of the photovoltaic component to be measured, and input the fused image data into the photovoltaic defect detection model to detect the defective areas in the image data.

[0011] Further, before extracting the visual features of each infrared image and visible light image, the infrared image and visible light image are aligned using the SIFT algorithm, which specifically includes the following steps:

[0012] Continuously perform parameter transformation on the scale of each infrared image and visible light image using the Gaussian function; the Gaussian function G ( x , y , σ ) is expressed as:

[0013] ;

[0014] where ([[]] x , y ) represents pixel coordinates; σ is the scale parameter, which determines the degree of image blurring;

[0015] Convolve the original image I ( x , y ) with Gaussian functions of different σ values, and the formula for the convolution operation is:

[0016] ;

[0017] Among them, L ( x ,y , σ ) represents the image representation at scale σ , where * represents the convolution operation;

[0018] By changing the value of σ , a series of images with different degrees of blurring are obtained, thus forming a multi-scale spatial sequence;

[0019] Detect local extreme points of the multi-scale spatial sequence by the Difference of Gaussian (DOG) method. Specifically, for each pair of adjacent images in the multi-scale spatial sequence, calculate the difference between them to obtain the DOG response value image :

[0020] ;

[0021] Among them, k is a constant of the adjacent scale space multiple;

[0022] In the DOG response value image, detect local extreme points by comparing each image pixel with the pixel values in its neighborhood;

[0023] Locate the detected local extreme points, and remove key points with contrast lower than the set value and unstable edge response points by fitting a three-dimensional quadratic function; and use the gradient of the pixels in the neighborhood of the feature points to determine their direction parameters;

[0024] Find the stable direction of the local structure of the key points by statistically analyzing the gradient histogram; based on the position and direction of the image gradient within the neighborhood of each key point, construct a 128-dimensional feature descriptor;

[0025] Align the infrared image and the visible light image according to this feature descriptor.

[0026] Furthermore, a dual encoder is used to process each infrared image and visible light image in parallel to obtain the visual features in the two images; a depth-interpretable iterative algorithm is used to fuse all the visual features to obtain a fused image dataset.

[0027] Furthermore, it also includes optimizing the fused image dataset by adopting the loss of the image fusion module and the loss of the semantic segmentation module before constructing the training dataset; among them, the loss of the image fusion module L F is expressed as:

[0028] ;

[0029] Among them, H and W represent the height and width of the image respectively, Target( ir ,vi ) represents the fused target image,

[0030] The cross - entropy loss is used as the loss of the semantic segmentation module, and its formula is:

[0031] ;

[0032] where, ( h , w ) and ( h , w ) respectively represent the predicted label and the ground - truth label of class h , w ) at the position ([[]] h , w ); c is the set of all classes in the semantic segmentation task; is an index variable, and its value range is from 0 to " - 1"; classes represents the cross - entropy loss function;

[0033] Minimize the obtained loss of the image fusion module and the loss of the semantic segmentation module to optimize the fused image dataset; the minimization process is specifically expressed as:

[0034] ;

[0035] where, arg min represents finding the optimal model parameters by minimizing the loss function; L represents the L1 norm; HW represents the product of the image height H and width W, representing the total number of pixels of the image; L F represents the total loss of image fusion, which is composed of the intensity loss and the detail loss weighted; Φ( ir , vi ) represents the fusion model, which takes the infrared image ir and the visible - light image vi as inputs and outputs the fused image f - max ( ir , vi ); , are the normalization weight parameters of the fusion loss, used to adjust the loss magnitude; G(⋅) is the gradient extraction operator; α represents the coefficient for balancing the intensity loss and the detail loss; L S represents the semantic segmentation loss, using the cross - entropy loss function; Ψ( f ) represents the semantic segmentation model, which takes the fused image f, output the semantic segmentation result t f ; λ is the coefficient to balance the fusion loss and the semantic segmentation loss; L cross-entropy is the cross-entropy loss function; c ∈[0, classes ) indicates traversing all semantic categories, with a total of classes categories; f =Φ( ir , vi ) indicates that the fused image f must be generated by the fusion model Φ, t f =Ψ( f ) indicates that the semantic segmentation result t f must be generated by the segmentation model Ψ.

[0036] Furthermore, the layer weights of the YOLOv11 model are initialized using this embedding representation, and the first layer of the YOLOv11 model is adjusted using the initialized layer weights, which specifically includes the following steps:

[0037] Add the output by the text encoder and the output by the image encoder to obtain e = + ;

[0038] e After the weight re-adjustment operation, N weights are generated; where the dimension D of the re-adjusted weights is D = h * w * c, h is the height, w is the width, and c is the number of channels; specifically e = + =CLIP_Embed(X i , Y i ); X i represents the image information input into the CLIP model, and Y i represents the text information input into the CLIP model;

[0039] According to e generate the final weight W1, and adjust the first layer of YOLOv11 using W1.

[0040] The present invention also provides a photovoltaic panel defect detection system, including:

[0041] An acquisition module for acquiring image data including different lighting conditions, different defect types, and different photovoltaic module types; the image data includes a plurality of infrared images and a plurality of visible light images;

[0042] A training set construction module, which is used to extract the visual features of each infrared image and visible light image, fuse all the visual features to obtain a fused image dataset; label the defective areas in the fused image dataset, and save the text information of the defective areas after labeling; construct a training dataset through the fused image dataset and the text information of the defective areas after labeling.

[0043] A training module, which is used to input the training dataset into the CLIP model, process the text information of the defective areas after labeling by using a text editor to generate an encoded vector ; process the fused image dataset by using an image encoder to generate an encoded vector ; generate an embedding representation according to the encoded vector and the encoded vector ; initialize the layer weights of the YOLOv11 model with the embedding representation, and adjust the first layer of the YOLOv11 model through the initialized layer weights to obtain a photovoltaic defect detection model.

[0044] A detection module, which is used to fuse the visual features of the infrared image and visible light image of the photovoltaic module to be tested, input the fused image data into the photovoltaic defect detection model, and detect the defective areas in the image data.

[0045] The present invention also provides a computer device for photovoltaic panel defect detection, including: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the photovoltaic panel defect detection method are implemented.

[0046] The present invention also provides a readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the photovoltaic panel defect detection method.

[0047] The present invention provides a method for photovoltaic panel defect detection, which has the following beneficial effects:

[0048] The present invention collects infrared light images and visible light images under various lighting conditions, defect types, and different photovoltaic module types, extracts visual features for efficient fusion, generates high-quality images with rich details and high information content, thereby improving the accuracy of the trained photovoltaic defect detection model; at the same time, considering the problem of limited labeled data in photovoltaic defect detection, an architecture combining the CLIP model and the YOLOv11 model is proposed. By using CLIP to process the text information of the fused image dataset and the labeled defect regions to generate embedding representations for initializing the YOLOv11 layer weights, and fine-tuning and training with a small amount of target domain data, the problem of difficult model training caused by the inability to obtain a large amount of data for photovoltaic detection in remote areas is solved. This method not only improves the detection performance of the model but also greatly shortens the training time, providing a more efficient and reliable solution for the maintenance and monitoring of photovoltaic systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of the photovoltaic panel defect detection method in the embodiment of the present invention;

[0050] Figure 2 It is a schematic diagram of the architecture of the multi-modal data fusion module in the embodiment of the present invention;

[0051] Figure 3 It is a flowchart of the initialization of YOLOv11 in the embodiment of the present invention;

[0052] Figure 4 It is a schematic diagram of the fine-tuning process of YOLOv11 in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0054] The present invention proposes a method for detecting photovoltaic panel defects, which uses the method of CLIP + YOLO + multi-modal data fusion to detect photovoltaic defects; this method ensures the consistency of infrared and visible light images in terms of resolution and size through image alignment processing, and uses the SIFT algorithm to eliminate shooting angle and perspective differences; then through multi-modal data fusion technology, it receives images from different sensors, extracts features using a dual private encoder, and performs feature fusion using a depth interpretable iterative algorithm, and finally optimizes the reconstructed image through a specific loss function. In the dataset annotation and division process, the labelme tool is used to annotate the fused dataset, and the annotation information is converted into a format acceptable to YOLO, and the training set, validation set, and test set are divided according to the ratio of 8:1:1. In terms of model establishment, combining the fast object detection ability of YOLO and the multi-modal understanding ability of CLIP, the layer weights of YOLOv11 are initialized by CLIP and fine-tuned to improve the performance in the small sample learning scenario; the method specifically includes the following steps:

[0055] S1. Obtain infrared images and visible light images with consistent parameters such as image resolution and size in the same scene. These images cover various lighting conditions, defect types, and different photovoltaic module types.

[0056] S2. Image alignment processing: Use the SIFT (Scale-Invariant Feature Transform) feature matching algorithm to align the infrared and visible light images to eliminate image misalignment caused by differences such as shooting angle and perspective.

[0057] SIFT is a descriptor used in the field of image processing, which has scale invariance and good stability for scale changes, rigid body transformations, light intensity, and occlusion of objects. It can detect key points in an image and is a local feature descriptor. The process of processing an image by the SIFT algorithm includes the following steps:

[0058] Use the Gaussian blur function to continuously perform parameter transformation on the scales of the infrared image and the visible light image to obtain a multi-scale space sequence. The Gaussian function G ( x , y , σ ) is used to construct the scale space, where ( x , y ) represents pixel coordinates, σ is the scale parameter, which determines the degree of image blur. The specific formula of the Gaussian function is:

[0059] (1);

[0060] Formula (1) describes how to σCalculate the weighted value of each pixel in the value calculation image, thereby realizing the blurring process of the image.

[0061] To obtain a multi-scale space sequence, the original image I ( x , y ) needs to be convolved with Gaussian functions of different σ values. The formula for the convolution operation is:

[0062] (2);

[0063] Among them, L ( x , y , σ ) represents the image representation at scale σ , and * represents the convolution operation. By changing the σ value, a series of images with different blurring degrees can be obtained, thereby constituting a multi-scale space sequence.

[0064] In these scale spaces, local extreme points are detected by the Difference of Gaussian (DOG) method. These extreme points are potential feature points. For each pair of adjacent images in the multi-scale space sequence, calculate the difference between them to obtain the DOG response value image, where each pixel value reflects the difference degree between images at adjacent scales.

[0065] DOG response value image:

[0066] (3);

[0067] Among them, k is a constant of the multiple of adjacent scale spaces, usually taking the value of .

[0068] In the DOG response value image, local extreme points are detected by comparing each pixel with the pixel values in its neighborhood. Local extreme points include local maxima and local minima, which are usually regarded as potential feature points.

[0069] Precisely locate the detected extreme points, and remove low-contrast key points and unstable edge response points by fitting a three-dimensional quadratic function. Use the gradient of the pixels in the neighborhood of the feature point to determine its direction parameter. By statistically analyzing the gradient histogram, find the stable direction of the local structure of the key point, which helps to achieve the rotational invariance of the image. For each key point, construct a 128-dimensional feature descriptor. This descriptor is based on the position and direction of the image gradient in the neighborhood of the key point and has rotational invariance and high uniqueness.

[0070] S3, Multimodal data fusion, such asFigure 2 As shown, its core objective is to efficiently fuse infrared images (IR) and visible light images (VI) to generate a high-quality image that contains both rich details and high information content. This module includes a dual encoder for feature extraction, a fusion strategy for promoting feature fusion, and a decoder for reconstructing the result.

[0071] S3.1. In the data input stage, infrared images and visible light images from different sensors are received; each of these two types of images has its own characteristics. Infrared images are good at capturing the thermal radiation information of objects and are particularly effective for object recognition in night or low-light conditions; while visible light images are known for their rich color and detail information and can clearly show the texture and structural features of objects. To ensure the smooth progress of the fusion process, these two types of images are image-aligned (as shown in S2) to eliminate the adverse effects caused by shooting conditions or device differences, laying a solid foundation for subsequent feature extraction and fusion.

[0072] S3.2. Feature extraction module, which adopts a dual private encoder architecture. This design aims to make full use of the unique information of infrared and visible light images respectively. Each encoder is specially trained to accurately capture the visual features in the image; the visual features include high-frequency and low-frequency features. High-frequency features mainly include small information such as edges and details, while low-frequency features cover more macroscopic information such as contours and backgrounds. Through the parallel processing of the dual encoder, the system can simultaneously obtain the important visual features in both types of images, providing comprehensive and detailed information support for subsequent feature fusion.

[0073] S3.3. Feature fusion stage, introducing a deep interpretable iterative algorithm. This algorithm realizes the deep integration of the features of the two modalities through the cascaded use of multiple small convolutional filters (such as 3x3 convolutional kernels). The activation function uses the Sigmoid function. This design not only reduces the spatial dimension of the parameters and the complexity of the model, but also enables the network to more efficiently capture the key information in the image during the training process. Through iterative optimization, the algorithm can gradually optimize the fusion result to ensure that the output image not only retains details but also takes into account the integrity of the overall information.

[0074] S3.4. Generate the final fused image through the reconstruction step. A specific loss function is adopted to guide the reconstruction process to ensure that the output image is visually close to the optimal combination of the infrared and visible light images. This loss function combines the loss of the fusion module and the loss of the segmentation module. The loss function of the segmentation module (cross-entropy loss) is introduced to improve the performance of the fused image in high-level semantic tasks. By introducing the semantic segmentation task, the model can pay more attention to the semantic structure in the image, thereby improving the performance of the fused image in subsequent high-level semantic tasks. This design aims to ensure that the finally generated image not only meets the visual quality requirements but also supports the execution of high-level semantic tasks. Through the combination of adaptive learning and regularization constraints, the depth interpretable iterative algorithm demonstrates excellent performance in the image fusion process, bringing new breakthroughs to the field of image fusion.

[0075] Loss of the image fusion module L F Focus on improving the visual effect of the fused image, ensuring that the fused image not only retains the thermal radiation information of the infrared image but also incorporates the rich color and details of the visible light image.

[0076] The loss function can be designed as:

[0077] (4);

[0078] Where H and W represent the height and width of the image respectively, Target ( ir , vi ) represents the fused target image. Here, the maximum value in the input images is adopted as the fused target, that is, max( ir , vi ). This design aims to ensure that the fused image is close to the optimal value in the input images in terms of brightness while retaining important detail information, thereby optimizing the overall visual effect.

[0079] Loss of the semantic segmentation module L S Then focus on the retention and improvement of high-level semantic information. By introducing the semantic segmentation task, the model can pay more attention to the semantic structure in the image during the fusion process, thereby improving the performance of the fused image in subsequent high-level semantic tasks.

[0080] The present invention adopts the cross-entropy loss as the loss of the semantic segmentation module, and its formula is:

[0081] (5);

[0082] Where ( h , w ) and ( h , w ) represent the predicted label and the ground truth label of class c at the location ( h , w ), respectively; is the set of all classes in the semantic segmentation task; is an index variable whose value ranges from 0 to " classes - 1"; represents the cross - entropy loss function, which, in classification tasks such as semantic segmentation, is used to measure the difference between the probability distribution of the class labels predicted by the model and the probability distribution of the ground truth class labels. By minimizing this loss function, the performance of the model can be optimized to make its predictions closer to the real situation; the cross - entropy loss evaluates the performance of the model in the semantic segmentation task by comparing the differences between the predicted labels and the ground truth labels. By minimizing this loss, the model can learn more accurate semantic information, thus better retaining these high - level information during the image fusion process.

[0083] (6);

[0084] where, arg min represents finding the optimal model parameters by minimizing the loss function; represents the L1 norm (sum of absolute values), which is used to calculate the difference between the predicted value and the reference value; HW represents the product of the image height (H) and width (W), indicating the total number of pixels in the image; L F represents the total loss of image fusion, which is composed of the intensity loss and the detail loss weighted; Φ( ir , vi ) represents the fusion model, which takes the infrared image ir and the visible - light image vi as inputs and outputs the fused image f - max ( ir , vi ); , are the normalization weight parameters of the fusion loss, which are used to adjust the loss magnitude; G(⋅) is the gradient extraction operator, which is used to calculate the gradient details (edges, textures) of the image; α represents the coefficient for balancing the intensity loss and the detail loss; L S represents the semantic segmentation loss, which adopts the cross - entropy loss function; Ψ( f ) represents the semantic segmentation model, which takes the fused image f as input and outputs the semantic segmentation result t f ; λ is the coefficient for balancing the fusion loss and the semantic segmentation loss;L cross-entropy is the cross-entropy loss function. c ∈[0, classes): Traverse all semantic classes (such as roads, vehicles, pedestrians, etc.), a total of classes classes; f = Φ( ir , vi ) represents the fused image f must be generated by the fusion model Φ, t f = Ψ( f ) represents the semantic segmentation result t f must be generated by the segmentation model Ψ.

[0085] Balance coefficient λ plays a crucial role in the loss function. It is used to regulate the loss of the image fusion module L F and the loss of the semantic segmentation module L S between the relative importance. By adjusting λ value, the trade-off between visual effects and semantic information retention can be balanced, so as to obtain a fused image that better meets the actual application requirements. By optimizing the details and intensity distribution, the overall quality of the fused image can be further improved to make it more in line with the actual application requirements.

[0086] In summary, Figure 2 The image fusion module shown realizes the high-quality fusion of infrared and visible light images through carefully designed preprocessing, feature extraction, feature fusion, and reconstruction steps. This module not only improves the quality of the fused image and the integrity of semantic information, but also provides more reliable and rich information support for subsequent visual tasks. This innovative design not only improves the efficiency and accuracy of image fusion technology, but also opens up a new path for the development of the image fusion field and shows broad application prospects.

[0087] S4. Dataset annotation.

[0088] After fusing a large amount of data according to S3, the diversity and comprehensiveness of the dataset are improved, enhancing the generalization ability of the model. Label the fused dataset. Dataset labeling is the process of assigning labels to each object or region in the dataset so that the machine learning model can learn these labels and make predictions. Labelme is an open-source image annotation tool that supports various annotation shapes such as polygons, rectangles, and circles. Since this invention is used for object detection, rectangular boxes are used for annotation. The types of photovoltaic panel defects include diode, open circuit, and hot spot defects. Use the tools provided by Labelme to draw annotation regions on the image as required and assign labels to each defect. After completing the annotation, save the annotation information as a JSON format file, providing the necessary supervision information for model training.

[0089] Use Labelme for annotation. The annotated file (the location information and types of the annotated photovoltaic defects) is saved as a JSON format file. Use open-source code to convert it into a txt file acceptable to YOLO, and divide the training set, validation set, and test set according to the ratio of 8:1:1. YOLO (You Only Look Once) is a popular object detection algorithm that accepts txt files in a specific format. These files contain the category, location, and size information of the objects in the image. Write scripts using programming languages such as Python to read the annotation information in the JSON file and convert it into the txt format acceptable to YOLO. Such conversion tools or libraries already exist in the open-source community. Dividing the converted txt and the fused images into the training set, validation set, and test set according to the ratio of 8:1:1 means that 80% of the data is used to train the model, 10% of the data is used to validate the model (adjust parameters during training), and the other 10% of the data is used to test the model (evaluate performance after training). The data used for training is the aligned dataset.

[0090] S4. Establish the CLIP and YOLOv11 models.

[0091] YOLO is well-known for its excellent speed and effective object detection. The model adopts a framework consisting of many convolutional layers, which are followed by fully connected layers. The task of each grid cell in the model is to detect objects within its designated area. CLIP is a multimodal model that effectively combines text and visual inputs to generate embeddings capable of performing zero-shot learning. The architecture includes two main components: a text encoder and an image encoder. The text encoder processes written descriptions, while the image encoder processes visual information. Both encoders generate embeddings in a shared latent space. CLIP has the ability to analyze and understand written and visual data in this collaborative environment simultaneously. The model has the ability to generate embeddings for any possible image and text pair, enabling it to complete various tasks without the need to provide specific training data for each individual task. The zero-shot ability is particularly beneficial for applications lacking labeled data, providing a general solution for various AI tasks.

[0092] The method of combining CLIP with YOLOv11 needs to utilize the multimodal capabilities of CLIP to improve the startup and training processes of YOLO. Initialize the layers of YOLOv11 using the embeddings generated by CLIP for the provided representative dataset, thus providing the model with a comprehensive understanding of the context from the very beginning, as Figure 3 shown. This strategy allows the use of more informative weights, thereby enhancing the effectiveness of the training process.

[0093] As Figure 3 shown, X i represents the input image data, which is a feature vector representing a specific image, and these images are extracted from the photovoltaic defect detection dataset (aligned dataset). Y i represents the bounding box information related to the image. These bounding boxes are used to identify the location and size of defects in the image, usually represented in coordinate form, indicating the specific area of the defect in the image. CLIP (Contrastive Language-Image PreTraining) is a multimodal model capable of processing image and text data. It generates rich embedding representations by learning the relationship between images and corresponding texts. By calling, the CLIP model takes the image feature vector and its corresponding bounding box information as input and generates an embedding representation e . This embedding representation e contains the visual features of the image and the location information related to the bounding box, and can provide context information and feature initialization for the subsequent YOLOv11n model. The generated embedding eThe layer weights used to initialize the YOLOv11n model are enhanced to improve the model's understanding and detection ability of image features. This process can improve the performance of the model in the few-shot learning scenario, especially when the data is limited. By using CLIP embeddings to initialize the initial layer, YOLOv11 is modified to include the most important features. In addition, the utilization of CLIP embeddings improves the efficiency of YOLOv11 in few-shot learning, enabling the model to quickly adapt to new tasks with less data( Figure 4 ). This minimizes the need for generating a large amount of data sets and reduces the training time required.

[0094] Figure 3 The process of generating weights based on CLIP and initializing the first layer of YOLOv11 specifically includes the following steps:

[0095] (1)CLIP embedding extraction:

[0096] Input part: The bounding box is input into the text encoder for processing, and the fused image is sent to the image encoder for processing.

[0097] Encoding process: The text encoder receives the bounding box as input and outputs an encoded vector after processing ; The image encoder encodes the fused image to generate an encoded vector .

[0098] (2)Comprehensive embedding calculation: Add the output by the text encoder and the output by the image encoder to obtain + , and further obtain the comprehensive embedding e .

[0099] (3)Weight initialization: Initialize the weights of the first convolutional layer of the YOLOv11 model to the comprehensive embedding e ; Assign the value of e directly to the convolutional kernel parameters of the first layer of the model. This initialization method breaks the traditional random initialization mode and enables the model to start learning from the pre-trained multi-modal feature space. e After the weight adjustment operation, the dimension D of the re-adjusted weights is D = h * w * c (h is the height, w is the width, and c is the number of channels), and N weights are generated ( Figure 3 is indicated by 1, 2,..., N in e = + = CLIP_Embed(X i , Y i ); Xi Represents the image information input into the CLIP model, Y i Represents the text information input into the CLIP model. The finally generated weight W1 (W1 = e ) is used to initialize the first layer of YOLOv11. The whole process aims to fuse the information of images and bounding boxes, and generate appropriate weights with the help of CLIP to initialize the first layer of the object detection model YOLOv11, so as to improve the performance of the model in subsequent detection tasks.

[0100] (4) Fine-tuning and optimization: After initialization, the standard backpropagation algorithm is used to fine-tune the entire model network:

[0101] The model is trained using a selected representative dataset. All network weights are iteratively updated by minimizing the detection loss function. The Adam optimizer is used, and appropriate learning rates and batch sizes are selected. The weights are updated through backpropagation, and the detection loss function used for minimization is: , by optimizing the weight matrix W , so that the prediction error (loss function x i , y i ) of the model for all data points ( L ) is minimized. Among them, min represents the minimization operation; Σ is the summation symbol, indicating the accumulation of all data points; L ( x i , y i ; W ) is the loss function, calculating the loss value of a single data point ( x i , y i ) under W ; x i is the input feature; y i is the ground truth label, which are the bounding box and class label in object detection.

[0102] The steps of using the Adam optimizer include:

[0103] Initialization: The weight W The initial value is the CLIP comprehensive embedding e .

[0104] Forward propagation: The input data x i passes through the model with weights W , and the prediction results (bounding boxes and class probabilities) are output.

[0105] Calculate Loss: Calculate the loss for each sample L ( x i , y i ; W ) and sum them up to get the total loss Σ L 。

[0106] Backpropagation: Calculate the gradient ▽ W with respect to the total loss W ·∑ L to determine the weight update direction

[0107] Update weights using the Adam optimizer: W1 := W- η * ▽W1·∑ L where η is the learning rate, and repeat for multiple iterations until convergence

[0108] Repeat the above steps until the loss function no longer decreases significantly or the preset number of iterations is reached, generating the updated weights

[0109] The fine - tuning process mainly includes two scenarios:

[0110] (1) Training from scratch: Train the YOLOv11n model from scratch using representative images

[0111] (2) Pre - trained YOLOv11: Fine - tune the YOLOv11n model using the pre - trained weights as a starting point

[0112] Training from scratch requires an adaptive strategy to learn discriminative information from images, especially when the data is less than that of the pre - trained model. The combination of CLIP embeddings and YOLOv11n has produced a significant enhancement. First, the rich context information in CLIP embeddings should help the model perform well when there are fewer training samples. Leveraging the multi - modal understanding of CLIP, YOLOv11n will be more adaptable to new tasks. The combination of CLIP embeddings and YOLOv11n has greatly improved the model's performance, making it a reliable and efficient solution for PV defect detection. This method minimizes the requirements for large amounts of data and training time, facilitating wider applications in the field of PV energy system maintenance and monitoring. Extending the model to include real - time defect detection of operating PV systems will be an important milestone for practical implementation. Ultimately, expanding the model's capabilities to identify a wider range of defects and irregularities in other categories of PV energy systems will increase its utility and efficiency in the sustainable energy field

[0113] The present invention provides a method for detecting photovoltaic panel defects, which has the following beneficial effects: 1) The SIFT algorithm is used for image alignment processing, ensuring the consistency of infrared and visible light images in terms of resolution and size, eliminating misalignment caused by shooting angle and perspective differences, and improving the accuracy and stability of image processing. 2) Through a dual private encoder architecture and a depth interpretable iterative algorithm, efficient fusion of infrared and visible light images is achieved, extracting and integrating high-frequency and low-frequency features, generating high-quality images containing rich details and high information content, providing comprehensive feature support for subsequent defect detection. The efficiency and comprehensiveness of multi-modal data fusion are improved. 3) Considering the problem of limited labeled data in photovoltaic defect detection, an architecture combining CLIP and YOLOv11n is proposed. By using the multi-modal understanding ability of CLIP to initialize the layers of YOLOv11n and fine-tuning the training on a small amount of target domain data, the problem of difficult model training due to the lack of a large amount of labeled data is solved. This method not only improves the detection performance of the model but also greatly shortens the training time, providing a more efficient and reliable solution for the maintenance and monitoring of photovoltaic systems.

[0114] Based on the same inventive concept, the present invention also proposes a photovoltaic panel defect detection system, including:

[0115] An acquisition module for acquiring image data including different lighting conditions, different defect types, and different photovoltaic component types; the image data includes a plurality of infrared images and a plurality of visible light images.

[0116] A training set construction module for extracting the visual features of each infrared image and visible light image, fusing all the visual features to obtain a fused image data set; annotating the defective areas in the fused image data set and saving the text information of the annotated defective areas; constructing a training data set through the fused image data set and the text information of the annotated defective areas.

[0117] A training module for inputting the training data set into the CLIP model, processing the text information of the annotated defective areas using a text editor to generate an encoded vector ; processing the fused image data set using an image encoder to generate an encoded vector ; generating an embedding representation according to the encoded vector and the encoded vector ; initializing the layer weights of the YOLOv11 model using the embedding representation, and adjusting the first layer of the YOLOv11 model through the initialized layer weights to obtain a photovoltaic defect detection model.

[0118] A detection module, which is used to fuse the visual features of the infrared image and the visible light image of the photovoltaic module to be tested, and input the fused image data into a photovoltaic defect detection model to detect the defective areas in the image data.

[0119] The present invention also provides a computer device for detecting photovoltaic panel defects, including: a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, the steps of the photovoltaic panel defect detection method are implemented.

[0120] The present invention also provides a readable storage medium storing a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the photovoltaic panel defect detection method.

[0121] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A photovoltaic panel defect detection method, characterized in that: The following steps are involved: Acquire image data containing different lighting conditions, different defect types, and different photovoltaic module types; the image data includes a plurality of infrared images and a plurality of visible light images; Extract visual features from each infrared image and visible light image, and fuse all visual features to obtain a fused image dataset; annotate defective areas in the fused image dataset, and save text information of the annotated defective areas; construct a training dataset using the fused image dataset and text information of the annotated defective areas; Input the training data set into the CLIP model, use a text editor to process the text information of the marked defect area, and generate a coding vector ; Use the image encoder to process the fused image data set and generate the encoding vector ; According to the encoding vector and the encoding vector Generating Embedded Representations e = + ; e After re-adjusting the weights, N weights are generated; The dimension after re-adjusting the weight is D=h*w*c, where h is the height, w is the width, and c is the number of channels; specifically: e = + =CLIP_Embed(X i ,Y i );X i Represents the image information input to the CLIP model, Y i Represents the text information input to the CLIP model; according to the embedding representation e Generate the final weight W1, adjust the first layer of YOLOv11 through W1, and obtain the photovoltaic defect detection model; The visual features of the infrared image and the visible light image of the photovoltaic module to be tested are fused, and the fused image data is input into the photovoltaic defect detection model to detect the defective area in the image data.

2. A photovoltaic panel defect detection method according to claim 1, characterized in that: Before extracting the visual features of each infrared image and visible light image, the infrared image and the visible light image are aligned using the SIFT algorithm, which specifically includes the following steps: The scale of each infrared image and visible light image is continuously transformed using a Gaussian function; the Gaussian function G ( x , y , σ ) is expressed as: ; in( x , y ) represents pixel coordinates; σ is the scale parameter, which determines the blurriness of the image; For the original image I ( x , y ) is different from σ The convolution operation is performed on the Gaussian function of the value. The formula for the convolution operation is: ; in, L ( x , y , σ ) indicates that the scale σ The image below shows that * represents the convolution operation; By changing σ The value of is used to obtain a series of images with different blur levels, thus forming a multi-scale spatial sequence; The local extreme points of the multi-scale space sequence are detected by the Gaussian difference DOG method. Specifically, for each pair of adjacent images in the multi-scale space sequence, the difference between them is calculated to obtain the DOG response value image : ; in, k is a constant that is a multiple of the adjacent scale space; In the DOG response value image, local extreme points are detected by comparing each image pixel with the pixel values ​​in its neighborhood; The detected local extreme points are located, and the key points with contrast lower than the set value and unstable edge response points are removed by fitting a three-dimensional quadratic function; and the gradient of the pixels in the neighborhood of the feature point is used to determine its direction parameter; By statistically analyzing the gradient histogram, the stable direction of the local structure of the key point is found; based on the position and direction of the image gradient in the neighborhood of each key point, a 128-dimensional feature descriptor is constructed; The infrared image and the visible light image are aligned according to the feature descriptor.

3. A photovoltaic panel defect detection method according to claim 1, characterized in that: A dual encoder is used to process each infrared image and visible light image in parallel to obtain the visual features of the two images; All visual features are fused using a deep interpretable iterative algorithm to obtain a fused image dataset.

4. A photovoltaic panel defect detection method according to claim 1, characterized in that: The method also includes optimizing the fused image dataset by using image fusion module loss and semantic segmentation module loss before constructing the training dataset; wherein the image fusion module loss L F It is expressed as: ; in, H and W Represents the height and width of the image respectively, Target( ir , vi ) represents the fusion target image, The cross entropy loss is used as the semantic segmentation module loss, and its formula is: ; in, ( h , w )and ( h , w ) respectively represent the positions ( h , w ) Category c The predicted labels and true labels of It is the set of all categories in the semantic segmentation task; is an index variable whose value ranges from 0 to " classes -1”; represents the cross entropy loss function; The obtained image fusion module loss and semantic segmentation module loss are minimized to optimize the fused image dataset; the minimization process is specifically expressed as: ; Among them, arg min means finding the optimal model parameters by minimizing the loss function; represents the L1 norm; HW represents the product of the image height H and width W, and represents the total number of pixels in the image; L F Represents the total image fusion loss, which is composed of intensity loss and loss of detail Weighted composition; Φ( ir , vi ) represents the fusion model, input infrared image ir and visible light images vi , output fused image f - max ( ir , vi ); , is the normalized weight parameter of the fusion loss, which is used to adjust the loss magnitude; G(⋅) is the gradient extraction operator; α represents the coefficient for balancing the intensity loss and detail loss; L S represents semantic segmentation loss, using the cross entropy loss function; Ψ( f ) represents the semantic segmentation model, and the input fusion image f , output semantic segmentation results t f ; λ The coefficient to balance the fusion loss and semantic segmentation loss; L cross-entropy is the cross entropy loss function; c ∈[0, classes ) means traversing all semantic categories, classes kind; f =Φ( ir , vi ) represents the fused image f must be generated by the fusion model Φ, t f =Ψ( f ) represents the semantic segmentation result t f must be generated by the segmentation model Ψ.

5. A photovoltaic panel defect detection system, characterized in that: include: An acquisition module, used to acquire image data containing different lighting conditions, different defect types and different photovoltaic module types; the image data includes a plurality of infrared images and a plurality of visible light images; The training set construction module is used to extract the visual features of each infrared image and visible light image, and fuse all the visual features to obtain a fused image data set; annotate the defective areas in the fused image data set, and save the text information of the annotated defective areas; and construct a training data set through the fused image data set and the text information of the annotated defective areas; The training module is used to input the training data set into the CLIP model, use a text editor to process the text information of the marked defect area, and generate a coding vector ; Use the image encoder to process the fused image data set and generate the encoding vector ; According to the encoding vector and the encoding vector Generating Embedded Representations e = + ; e After re-adjusting the weights, N weights are generated; The dimension after re-adjusting the weight is D=h*w*c, where h is the height, w is the width, and c is the number of channels; specifically: e = + =CLIP_Embed(X i ,Y i );X i Represents the image information input to the CLIP model, Y i Represents the text information input to the CLIP model; according to the embedding representation e Generate the final weight W1, adjust the first layer of YOLOv11 through W1, and obtain the photovoltaic defect detection model; The detection module is used to fuse the visual features of the infrared image and the visible light image of the photovoltaic module to be tested, input the fused image data into the photovoltaic defect detection model, and detect the defective area in the image data.

6. A photovoltaic panel defect detection computer device, characterized in that: include: A memory, a processor and a computer program stored in the memory, wherein the processor implements the steps of the photovoltaic panel defect detection method according to any one of claims 1 to 4 when executing the computer program.

7. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, they are used to execute the steps of the photovoltaic panel defect detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Target detection method, system and device fusing bimodal features and storage medium

    CN113962246A

  • Power transformer defect diagnosis method based on multi-mode sound image fusion

    CN118779807A