Photovoltaic module infrared image fault detection method based on unmanned aerial vehicle inspection
By using high-frame rate and high-resolution infrared image acquisition and deep learning models in drone inspections, the noise and uneven brightness problems in existing technologies are solved, high-precision photovoltaic module fault detection is achieved, and detection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510696799.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
AI Technical Summary
Existing infrared image acquisition methods for drone inspections lack adaptability to noise and uneven brightness, resulting in weakened image feature information and affecting the accuracy of fault detection. In particular, missed detections and false detections are prone to occur in scenes with complex backgrounds or blurred hot spot edges, and the recognition accuracy of small targets is low.
A DJI M300RTK drone platform equipped with an H20T infrared thermal imager was used. High frame rate and high resolution image acquisition standards were set. Combined with GPS and attitude data, grayscale normalization, two-stage noise reduction, histogram equalization, and edge enhancement processing were performed. A DMR-RTDETR model based on RT-DETR was constructed, integrating the MPCAFF module, RePshuffleRoss module, and RepConv detection head structure to improve image quality and small target detection accuracy.
It achieves high-precision infrared image acquisition and fault detection of photovoltaic modules, significantly improves image clarity and the accuracy of small target detection, reduces false detection and missed detection rates, and improves the efficiency and intelligence level of module-level fault inspection in photovoltaic power stations.
Smart Images

Figure CN120634977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault detection, and in particular to a photovoltaic component infrared image fault detection method based on unmanned aerial vehicle inspection. Background Art
[0002] With the large-scale application of photovoltaic power generation around the world, the importance of photovoltaic power station operation and maintenance has become increasingly prominent. Photovoltaic modules are exposed to complex outdoor environments for a long time, and are prone to problems such as hot spot failures, fragmentation failures, and diode failures, which affect power generation efficiency and pose safety hazards. To ensure the stable operation of photovoltaic power stations, it is urgent to carry out efficient and accurate component fault detection. In recent years, the inspection method of drones equipped with infrared thermal imagers has become the mainstream means to replace manual inspections, with the advantages of high coverage, low labor costs and remote operation.
[0003] Although the existing infrared image acquisition method based on drone inspection has been widely used, the following prominent problems still exist. On the one hand, the traditional image preprocessing process lacks adaptability to infrared image noise and uneven brightness, which can easily lead to the weakening of image feature information, thereby affecting the recognition accuracy of subsequent fault detection models. On the other hand, the existing detection model has low recognition accuracy for smaller fault targets, especially in scenes with complex backgrounds or blurred hot spot edges, which are more prone to missed detections and false detections. Summary of the Invention
[0004] The purpose of the present invention is to address the shortcomings of the existing technology and provide a photovoltaic component infrared image fault detection method based on drone inspection. The method aims to realize high-precision acquisition of large-scale infrared images of photovoltaic components by configuring the DJI M300RTK drone platform and carrying the H20T infrared thermal imager; by setting the acquisition standard of not less than 5 frames / second and image resolution not less than 640×512 pixels, and combining GPS and attitude data, accurate binding of image acquisition and positioning information is achieved; in terms of image processing, a multi-level image enhancement strategy of grayscale normalization, two-stage noise reduction, histogram equalization and edge enhancement is adopted to improve image quality and target boundary clarity; in terms of detection algorithm, a DMR-RTDETR model based on RT-DETR is constructed, which integrates the MPCAFF module, RePshuffleRoss module and RepConv detection head structure, significantly improving the detection accuracy and real-time performance of small targets, and effectively overcoming the problem of poor recognition performance of traditional methods in highly complex scenarios.
[0005] To this end, the present application provides a photovoltaic module infrared image fault detection method based on drone inspection, comprising the following steps:
[0006] Step S100: configure the UAV platform, plan the flight path, set the aerial photography parameters, and set the infrared thermal image acquisition and infrared image data transmission and storage mechanism.
[0007] Step S200 , performing grayscale normalization processing, image noise reduction, histogram equalization, edge enhancement processing, and size standardization processing on the infrared image data, and performing data enhancement operations.
[0008] Step S300: Configure the deep neural network model architecture, perform feature fusion and recalibration operations, integrate the RePshuffleRoss module to optimize features, configure the decoder structure and detection head, configure the training data set and training parameters, and perform evaluation.
[0009] Step S400: Execute reasoning operations and IoU-Aware target confidence screening mechanism, perform target classification and fault type judgment, perform target classification and fault type judgment, and perform data encapsulation and merge duplicate targets.
[0010] Step S500: Configure the performance evaluation indicator system and perform quantitative evaluation, introduce a visual explanation mechanism, generate a focused heat map, perform comparative analysis between different models, and visualize the test results and generate a structured test report.
[0011] In some specific implementations, the step S100 specifically includes:
[0012] Step S100.1: Configure the UAV platform, plan the flight path, and set the aerial photography parameters.
[0013] The DJI M300RTK drone is used as the flight platform, and the H20T infrared thermal imager is selected as the image acquisition payload. It is installed on the drone platform through a dedicated mounting interface, and the flight control system performs stable attitude control.
[0014] A flight mission planning system with a three-dimensional trajectory generation function is configured on the ground control terminal. By importing the geographic information system data of the photovoltaic power station, the photovoltaic module arrangement structure, terrain height difference data, fixed obstacle coordinates and boundary information are loaded.
[0015] According to the scope of the image acquisition mission, set the flight mission parameters, including: track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate not less than 80% and lateral overlap rate not less than 70%.
[0016] A flight mission file is generated based on the track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate and lateral overlap rate, including: waypoint coordinate sequence, flight altitude, speed, and mission number information, and the flight mission file is uploaded to the UAV flight control system as an execution instruction.
[0017] The standard flight altitude range of the drone is set to 40 meters to 110 meters. The flight altitude is dynamically adjusted according to the terrain undulations and the distribution density of photovoltaic modules. The flight speed range is set to 1.5 meters / second to 3.0 meters / second. The flight control system controls the stability of the flight attitude, and the roll angle and pitch angle are set to no more than 5 degrees.
[0018] Step S100.2: Setting the infrared thermal image acquisition and infrared image data transmission and storage mechanism.
[0019] The UAV automatically flies according to the flight mission file. During the flight, the infrared thermal imager collects infrared thermal images at a frame rate of not less than 5 frames per second. The collected images are infrared radiation grayscale images of 640×512 pixels or larger than 640×512 pixels, and the image format is 16-bit grayscale image.
[0020] The fifth-generation mobile communication module integrated in the drone payload transmits infrared thermal image data in real time to the ground control station or edge computing server, and the image data is transmitted using the User Datagram Protocol.
[0021] The receiving end analyzes the image data in real time and archives and stores it according to the image number, drone global positioning system positioning information and flight timestamp.
[0022] During the image acquisition process, the flight attitude data and GPS positioning data output by the drone's inertial measurement unit are collected synchronously, and each frame of image is bound to the image frame number, GPS coordinate information, flight attitude information and image acquisition timestamp metadata fields.
[0023] In some specific implementations, the step S200 specifically includes:
[0024] Step S200.1: normalize the grayscale of the infrared image, and perform image noise reduction, histogram equalization, and edge enhancement.
[0025] Receive infrared radiation grayscale image data transmitted from the ground control station or edge computing server and perform pixel-by-pixel normalization processing. Use the grayscale normalization algorithm to transform the image data. Specifically:
[0026]
[0027] Where: I norm(x, y) is the normalized grayscale value at the pixel position (x, y), I(x, y) is the original grayscale value at the pixel position (x, y), and I min is the minimum grayscale value in the entire infrared image, I max is the maximum grayscale value in the entire infrared image, x is the horizontal pixel coordinate of the image, and y is the vertical pixel coordinate of the image.
[0028] For the thermal noise and patch noise in the original infrared thermal image caused by image sensor noise, flight vibration jitter or electromagnetic interference, median filtering and Gaussian filtering are used in sequence to perform image denoising.
[0029] A median filter with a 3×3 sliding window is used to replace the grayscale value of each pixel with the median of the grayscale of its neighboring pixels, and a two-dimensional Gaussian filter with a standard deviation of σ=1.2 is used to perform convolution smoothing on the image.
[0030] The whole image grayscale histogram equalization operation is performed on the denoised image, and the Laplacian edge enhancement operator is used to perform edge enhancement on the image.
[0031] Step S200.2: Normalize the size of the infrared image and perform data enhancement.
[0032] For images with inconsistent aspect ratios, bilinear interpolation resampling algorithm is used for scaling. If the image boundary is still less than 640×640 pixels, the missing area is filled by edge symmetrical filling without introducing redundant information.
[0033] Data augmentation operations are performed on the standardized images, and diverse training image pseudo samples are generated through random horizontal flipping, random rotation, random brightness perturbation, local occlusion simulation, and Gaussian noise perturbation simulation.
[0034] Generate a label file required for target detection for each processed image. The label file contains: hot spot fault location bounding box, hot spot fault type, image index number and task number.
[0035] In some specific implementations, step S300 specifically includes:
[0036] Step S300.1: Configure the deep neural network model architecture and perform feature fusion and recalibration operations.
[0037] Through the RT-DETR structure, the neural network is configured. The model structure includes: backbone feature extraction network, feature fusion encoding module and decoding detection module.
[0038] ResNet-50 is used as the backbone feature extraction network, which outputs five layers of multi-scale feature maps P1, P2, P3, P4, and P5. Each layer of feature maps corresponds to a different receptive field and resolution, and extracts shallow edge and deep semantic features respectively.
[0039] The two mid-scale feature maps P3 and P4 are selected as input to the multi-path cross-scale attention feature fusion module.
[0040] In the multi-path cross-scale attention feature fusion module, global features and local detail features are extracted respectively through two parallel pathways to enhance the network's sensitivity to small targets of different scales.
[0041] The global pathway introduces a spatial attention mechanism, performs maximum pooling and average pooling on the input feature map in the channel dimension, and obtains a spatial attention map.
[0042] The local path introduces a channel-by-channel weighted channel attention mechanism, combined with a multi-layer perceptron to learn the feature importance of each channel. After global and local fusion, a recalibrated fusion feature map F is obtained. MPCAFF .
[0043] Step S300.2: Integrate the RePshuffleRoss module optimization features and configure the decoder structure and detection head.
[0044] The fused feature map is input into the RePshuffleRoss module to perform multi-path convolution and channel shuffling operations. During the training phase, three-way parallel convolution operations of 1×1, 3×3, and 5×5 are used to extract edge textures and deep patterns, and channel shuffling operations are performed to break the inter-channel dependencies. During the inference phase, the multi-path structure is compressed into an equivalent single-path convolution through reparameterization technology.
[0045] The fused and optimized feature map is input into the Transformer decoder structure, and the IoU-AwareQueryMatching strategy is used to initialize the target query. Two scale feature maps with resolutions of 40×40 and 80×80 are selected and input into the decoder, and the RepConv optimized detection head structure is introduced.
[0046] RepConv merges the multi-layer structure of the training phase into a single layer of efficient convolution. The detection head output includes: target bounding box coordinates (x, y, w, h) and fault category probability, and outputs a set of embedding vectors, which contain: target location and category information.
[0047] Step S300.3: Configure the training data set and training parameters, and perform evaluation.
[0048] The photovoltaic fault detection dataset is used as the training set, which includes hot spot faults, fragmentation faults and diode faults.
[0049] Use the LabelImg tool to annotate the bounding boxes of the objects in the image and generate a COCO format label file.
[0050] After training is completed, export the model weight file θ final , and evaluate the model detection performance. The performance of the model on the validation set is quantitatively evaluated using the average precision, the average precision under the condition of IoU threshold of 0.5, the average precision under the condition of IoU threshold of 0.75, and the average precision of small targets. Specifically:
[0051]
[0052] Where A is the area of the predicted bounding box, B is the area of the true bounding box, A∩B is the area of the intersection of the predicted box and the true box, A∪B is the area of the union of the predicted box and the true box, and IoU is the weight ratio between the predicted result and the true result.
[0053] If the model achieves the performance standards of mAP ≥ 0.87 and APs ≥ 0.44 on the PVFD photovoltaic fault detection dataset, it is considered a trained model that meets deployment requirements.
[0054] In some specific implementations, the step S400 specifically includes:
[0055] Step S400.1: Execute reasoning operations and the IoU-Aware target confidence screening mechanism, and perform target classification and fault type judgment.
[0056] The infrared thermal image after grayscale normalization, noise suppression, enhancement processing and size standardization is input into the trained DMR-RTDETR network model for feature extraction, feature fusion, feature optimization and decoding inference.
[0057] The backbone network ResNet-50 outputs multi-scale feature maps for feature extraction, the MPCAFF module performs global-local attention fusion for feature fusion, the RePshuffleRoss module performs channel shuffling and re-parameterization for feature optimization, and the Transformer decoder and RepConv detection head output detection results for decoding inference.
[0058] The output of the DMR-RTDETR model includes: the predicted bounding box coordinates of the detected target, the category label probability distribution, and the target confidence value.
[0059] The target confidence value output by the IoU-AwareQueryMatching mechanism is used as the screening basis, and the detection result screening threshold is set to 0.5. The target detection box D with a target confidence value greater than or equal to 0.5 is retained. i , remove low-quality, low-overlap or background misdetected areas, and calculate the target confidence of the IoU-Aware mechanism, specifically:
[0060] C i =cls i ×IoU i
[0061] Where: C i is the target detection box D i Confidence, cls i is the target detection box D i Class probability, IoU i is the target detection box D i The intersection-over-union ratio with the ground-truth bounding box, D i is the i-th target detection box output by the model.
[0062] Step S400.2: Target classification and fault type determination are performed, and data encapsulation and merging of duplicate targets are performed.
[0063] The retained target detection box is selected according to the maximum value of the category probability vector to classify the target fault type. The fault types are divided into: hot spot fault, fragmentation fault and diode fault. The final output of the target includes: position bounding box coordinates, fault type label, confidence value and image index number.
[0064] The detection results corresponding to each infrared image are formatted and saved as a structured data file, encapsulated in JSON format. The fields include: image file number, number of detected targets, bounding box coordinates of each target, fault type label, target confidence value and detection timestamp.
[0065] If there are multiple predicted bounding boxes in the same image that are repeatedly annotated near the same position, NMS is used to merge the redundant boxes.
[0066] Sort all candidate boxes from large to small according to the confidence value, and select the box with the maximum confidence D from the sorted list max , add it to the final output set, and combine it with D max Other candidate boxes whose IoU is greater than 0.5 are eliminated, and the steps are repeated 2-3 times.
[0067] In some specific implementations, the step S500 specifically includes:
[0068] Step S500.1: Configure a performance evaluation indicator system, perform quantitative evaluation, introduce a visual explanation mechanism, and generate a focus heat map.
[0069] A hybrid evaluation index system is used to comprehensively evaluate the model detection results. The evaluation indicators include: precision, recall rate, F1 score, average precision and small target average precision. The F1 score is calculated as follows:
[0070]
[0071] Where: F1 is the F1 score, P is the precision, which indicates the proportion of positive samples detected as positive, R is the recall, which indicates the proportion of positive samples that are correctly detected, × is the multiplication operator, which indicates the product of P and R, + is the addition operator, which indicates the addition of P and R. It is the harmonic mean of precision and recall.
[0072] If the model achieves the following results in the validation set: average precision mAP ≥ 0.87, small target average precision APs ≥ 0.44, and F1 score F1 ≥ 0.82, the model is considered to be ready for actual deployment.
[0073] During the model inference process, a visual interpretation mechanism based on gradient class activation mapping is introduced to reversely track the convolutional feature responses of the detection target area and generate a heat map of the attention area.
[0074] The Grad-CAM process is as follows: extract the gradient of a key convolutional layer in the detection head, perform weighted averaging of each channel and multiply it with the feature map to generate a thermal distribution map that is consistent with the input image size. The map is then displayed overlaid on the original image using pseudo-color to highlight the key areas of the model.
[0075] Step S500.2: Perform comparative analysis between different models, visualize the test results, and generate a structured test report.
[0076] The DMR-RTDETR model is compared with existing detection models on the same photovoltaic fault dataset. The comparison dimensions include: inference speed, number of parameters, FLOPs calculation amount, mAP, AP50, AP75 and APs indicators.
[0077] Configure a graphical user interface to display the detection results in graphical form on the device terminal. The image output content includes: original infrared thermal image, detection frame overlay image, fault type statistical histogram, thermal area response map, image frame sequence browsing and target jump function.
[0078] The model inference results are converted into a standard format detection report, and archivable files are generated in task batches. The detection report content includes: inspection task number and timestamp, total number of images and proportion of abnormal images, number and distribution ratio of various fault targets, summary of detection accuracy and confidence, fault image index and visualization screenshot embedding, and model version information and training indicator summary.
[0079] Compared with the prior art, the present invention has the following beneficial effects:
[0080] 1. By combining drone infrared inspection technology with a deep learning target detection model, a complete photovoltaic module infrared image fault detection process was established. This process automates the entire process from image acquisition, image preprocessing, model detection, to result evaluation and visualization. This method offers the advantages of high fault identification accuracy, fast detection speed, and standardized processing procedures, significantly improving the efficiency and intelligence of component-level fault inspections in photovoltaic power plants.
[0081] 2. To address the problems of existing infrared images being subject to large noise interference, non-uniform illumination, and poor image contrast, the present invention introduces grayscale normalization, two-stage filtering noise reduction, histogram equalization, and edge enhancement mechanisms in the image preprocessing stage, effectively improving image clarity and the ability to express target boundary information; at the same time, a DMR-RTDETR deep neural network model is constructed, integrating a multi-path cross-scale attention mechanism with a lightweight detection head structure, thereby improving the detection accuracy of small-sized fault targets and significantly reducing the false detection rate and missed detection rate in traditional methods.
[0082] 3. UAVs support multi-scenario deployment and customized mission planning. Combined with functions such as real-time image transmission, automatic numbering and archiving, visual recognition result display, and structured report output, they have good system scalability and engineering applicability. By connecting data with the operation and maintenance system platform, they can also realize fault warning, trend analysis, and maintenance decision-making assistance, further promoting the development of photovoltaic power station fault inspection towards intelligent and high-precision directions. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 This is an overall flow chart of a photovoltaic module infrared image fault detection method based on drone inspection provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] Please refer to Figure 1 , which shows the process of an embodiment of a photovoltaic module infrared image fault detection method based on drone inspection according to the present disclosure
[0085] like Figure 1 As shown, a photovoltaic module infrared image fault detection method based on drone inspection includes the following steps:
[0086] Step S100: configure the UAV platform, plan the flight path, set the aerial photography parameters, and set the infrared thermal image acquisition and infrared image data transmission and storage mechanism.
[0087] Step S200 , performing grayscale normalization processing, image noise reduction, histogram equalization, edge enhancement processing, and size standardization processing on the infrared image data, and performing data enhancement operations.
[0088] Step S300: Configure the deep neural network model architecture, perform feature fusion and recalibration operations, integrate the RePshuffleRoss module to optimize features, configure the decoder structure and detection head, configure the training data set and training parameters, and perform evaluation.
[0089] Step S400: Execute reasoning operations and IoU-Aware target confidence screening mechanism, perform target classification and fault type judgment, perform target classification and fault type judgment, and perform data encapsulation and merge duplicate targets.
[0090] Step S500: Configure the performance evaluation indicator system and perform quantitative evaluation, introduce a visual explanation mechanism, generate a focused heat map, perform comparative analysis between different models, and visualize the test results and generate a structured test report.
[0091] In some specific implementations, the step S100 specifically includes:
[0092] Step S100.1: Configure the UAV platform, plan the flight path, and set the aerial photography parameters.
[0093] The DJI M300RTK drone is used as the flight platform, and the H20T infrared thermal imager is selected as the image acquisition payload. It is installed on the drone platform through a dedicated mounting interface, and the flight control system performs stable attitude control.
[0094] A flight mission planning system with a three-dimensional trajectory generation function is configured on the ground control terminal. By importing the geographic information system data of the photovoltaic power station, the photovoltaic module arrangement structure, terrain height difference data, fixed obstacle coordinates and boundary information are loaded.
[0095] According to the scope of the image acquisition mission, set the flight mission parameters, including: track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate not less than 80% and lateral overlap rate not less than 70%.
[0096] A flight mission file is generated based on the track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate and lateral overlap rate, including: waypoint coordinate sequence, flight altitude, speed, and mission number information, and the flight mission file is uploaded to the UAV flight control system as an execution instruction.
[0097] The standard flight altitude range of the drone is set to 40 meters to 110 meters. The flight altitude is dynamically adjusted according to the terrain undulations and the distribution density of photovoltaic modules. The flight speed range is set to 1.5 meters / second to 3.0 meters / second. The flight control system controls the stability of the flight attitude, and the roll angle and pitch angle are set to no more than 5 degrees.
[0098] Step S100.2: Setting the infrared thermal image acquisition and infrared image data transmission and storage mechanism.
[0099] The UAV automatically flies according to the flight mission file. During the flight, the infrared thermal imager collects infrared thermal images at a frame rate of not less than 5 frames per second. The collected images are infrared radiation grayscale images of 640×512 pixels or larger than 640×512 pixels, and the image format is 16-bit grayscale image.
[0100] The fifth-generation mobile communication module integrated in the drone payload transmits infrared thermal image data in real time to the ground control station or edge computing server, and the image data is transmitted using the User Datagram Protocol.
[0101] The receiving end analyzes the image data in real time and archives and stores it according to the image number, drone global positioning system positioning information and flight timestamp.
[0102] During the image acquisition process, the flight attitude data and GPS positioning data output by the drone's inertial measurement unit are collected synchronously, and each frame of image is bound to the image frame number, GPS coordinate information, flight attitude information and image acquisition timestamp metadata fields.
[0103] In some specific implementations, the step S200 specifically includes:
[0104] Step S200.1: normalize the grayscale of the infrared image, and perform image noise reduction, histogram equalization, and edge enhancement.
[0105] Receive infrared radiation grayscale image data transmitted from the ground control station or edge computing server and perform pixel-by-pixel normalization processing. Use the grayscale normalization algorithm to transform the image data. Specifically:
[0106]
[0107] Where: I norm(x, y) is the normalized grayscale value at the pixel position (x, y), I(x, y) is the original grayscale value at the pixel position (x, y), and I min is the minimum grayscale value in the entire infrared image, I max is the maximum grayscale value in the entire infrared image, x is the horizontal pixel coordinate of the image, and y is the vertical pixel coordinate of the image.
[0108] For the thermal noise and patch noise in the original infrared thermal image caused by image sensor noise, flight vibration jitter or electromagnetic interference, median filtering and Gaussian filtering are used in sequence to perform image denoising.
[0109] A median filter with a 3×3 sliding window is used to replace the grayscale value of each pixel with the median of the grayscale of its neighboring pixels, and a two-dimensional Gaussian filter with a standard deviation of σ=1.2 is used to perform convolution smoothing on the image.
[0110] The whole image grayscale histogram equalization operation is performed on the denoised image, and the Laplacian edge enhancement operator is used to perform edge enhancement on the image.
[0111] Step S200.2: Normalize the size of the infrared image and perform data enhancement.
[0112] For images with inconsistent aspect ratios, bilinear interpolation resampling algorithm is used for scaling. If the image boundary is still less than 640×640 pixels, the missing area is filled by edge symmetrical filling without introducing redundant information.
[0113] Data augmentation operations are performed on the standardized images, and diverse training image pseudo samples are generated through random horizontal flipping, random rotation, random brightness perturbation, local occlusion simulation, and Gaussian noise perturbation simulation.
[0114] The probability of random horizontal flipping is 0.5, the range of random rotation angle is ±15 degrees, the range of random brightness perturbation adjustment is ±10%, local occlusion simulation is to randomly select image areas for filling interference, and the variance of Gaussian noise perturbation simulation is 0.01.
[0115] Generate a label file required for target detection for each processed image. The label file contains: hot spot fault location bounding box, hot spot fault type, image index number and task number.
[0116] In some specific implementations, step S300 specifically includes:
[0117] Step S300.1: Configure the deep neural network model architecture and perform feature fusion and recalibration operations.
[0118] Through the RT-DETR structure, the neural network is configured. The model structure includes: backbone feature extraction network, feature fusion encoding module and decoding detection module.
[0119] ResNet-50 is used as the backbone feature extraction network, which outputs five layers of multi-scale feature maps P1, P2, P3, P4, and P5. Each layer of feature maps corresponds to a different receptive field and resolution, and extracts shallow edge and deep semantic features respectively.
[0120] The two mid-scale feature maps P3 and P4 are selected as input to the multi-path cross-scale attention feature fusion module.
[0121] In the multi-path cross-scale attention feature fusion module, global features and local detail features are extracted respectively through two parallel pathways to enhance the network's sensitivity to small targets of different scales.
[0122] The global pathway introduces a spatial attention mechanism, performs maximum pooling and average pooling on the input feature map in the channel dimension, and obtains a spatial attention map.
[0123] The local path introduces a channel-by-channel weighted channel attention mechanism, combined with a multi-layer perceptron to learn the feature importance of each channel. After global and local fusion, a recalibrated fusion feature map F is obtained. MPCAFF .
[0124] Step S300.2: Integrate the RePshuffleRoss module optimization features and configure the decoder structure and detection head.
[0125] The fused feature map is input into the RePshuffleRoss module to perform multi-path convolution and channel shuffling operations. During the training phase, three-way parallel convolution operations of 1×1, 3×3, and 5×5 are used to extract edge textures and deep patterns, and channel shuffling operations are performed to break the inter-channel dependencies. During the inference phase, the multi-path structure is compressed into an equivalent single-path convolution through reparameterization technology.
[0126] The fused and optimized feature map is input into the Transformer decoder structure, and the IoU-AwareQueryMatching strategy is used to initialize the target query. Two scale feature maps with resolutions of 40×40 and 80×80 are selected and input into the decoder, and the RepConv optimized detection head structure is introduced.
[0127] RepConv merges the multi-layer structure of the training phase into a single layer of efficient convolution. The detection head output includes: target bounding box coordinates (x, y, w, h) and fault category probability, and outputs a set of embedding vectors, which contain: target location and category information.
[0128] Step S300.3: Configure the training data set and training parameters, and perform evaluation.
[0129] The photovoltaic fault detection dataset is used as the training set, which includes hot spot faults, fragmentation faults and diode faults.
[0130] Use the LabelImg tool to annotate the bounding boxes of the objects in the image and generate a COCO format label file.
[0131] The training parameters include: 150 training rounds, 1×10 learning rate -4 , optimizer AdamW, loss function combining IoU loss function with FocalLoss, batch size 4 and image input size 640×640 pixels.
[0132] After training is completed, export the model weight file θ final , and evaluate the model detection performance. The performance of the model on the validation set is quantitatively evaluated using the average precision, the average precision under the condition of IoU threshold of 0.5, the average precision under the condition of IoU threshold of 0.75, and the average precision of small targets. Specifically:
[0133]
[0134] Where A is the area of the predicted bounding box, B is the area of the true bounding box, A∩B is the area of the intersection of the predicted box and the true box, A∪B is the area of the union of the predicted box and the true box, and IoU is the weight ratio between the predicted result and the true result.
[0135] If the model achieves the performance standards of mAP ≥ 0.87 and APs ≥ 0.44 on the PVFD photovoltaic fault detection dataset, it is considered a trained model that meets deployment requirements.
[0136] In some specific implementations, the step S400 specifically includes:
[0137] Step S400.1: Execute reasoning operations and the IoU-Aware target confidence screening mechanism, and perform target classification and fault type judgment.
[0138] The infrared thermal image after grayscale normalization, noise suppression, enhancement processing and size standardization is input into the trained DMR-RTDETR network model for feature extraction, feature fusion, feature optimization and decoding inference.
[0139] The backbone network ResNet-50 outputs multi-scale feature maps for feature extraction, the MPCAFF module performs global-local attention fusion for feature fusion, the RePshuffleRoss module performs channel shuffling and re-parameterization for feature optimization, and the Transformer decoder and RepConv detection head output detection results for decoding inference.
[0140] The output of the DMR-RTDETR model includes: the predicted bounding box coordinates of the detected target, the category label probability distribution, and the target confidence value.
[0141] The target confidence value output by the IoU-AwareQueryMatching mechanism is used as the screening basis, and the detection result screening threshold is set to 0.5. The target detection box D with a target confidence value greater than or equal to 0.5 is retained. i , remove low-quality, low-overlap or background misdetected areas, and calculate the target confidence of the IoU-Aware mechanism, specifically:
[0142] C i =cls i ×IoU i
[0143] Where: C i is the target detection box D i Confidence, cls i is the target detection box D i Class probability, IoU i is the target detection box D i The intersection-over-union ratio with the ground-truth bounding box, D i is the i-th target detection box output by the model.
[0144] Step S400.2: Target classification and fault type determination are performed, and data encapsulation and merging of duplicate targets are performed.
[0145] The retained target detection box is selected according to the maximum value of the category probability vector to classify the target fault type. The fault types are divided into: hot spot fault, fragmentation fault and diode fault. The final output of the target includes: position bounding box coordinates, fault type label, confidence value and image index number.
[0146] The detection results corresponding to each infrared image are formatted and saved as a structured data file, encapsulated in JSON format. The fields include: image file number, number of detected targets, bounding box coordinates of each target, fault type label, target confidence value and detection timestamp.
[0147] If there are multiple predicted bounding boxes in the same image that are repeatedly annotated near the same position, NMS is used to merge the redundant boxes.
[0148] Sort all candidate boxes from large to small according to the confidence value, and select the box with the maximum confidence D from the sorted list mas , add it to the final output set, and combine it with D mas Other candidate boxes whose IoU is greater than 0.5 are eliminated, and the steps are repeated 2-3 times.
[0149] The final inspection result file is stored in a local disk directory or uploaded to a cloud edge computing server using the task number and image sequence number as the naming method. A data docking interface with the photovoltaic power station maintenance system platform is established, supporting the following: fault image echo, inspection coordinate heat map display, fault classification statistical map generation, data export, and alarm notification triggering.
[0150] In some specific implementations, the step S500 specifically includes:
[0151] Step S500.1: Configure a performance evaluation indicator system, perform quantitative evaluation, introduce a visual explanation mechanism, and generate a focus heat map.
[0152] A hybrid evaluation index system is used to comprehensively evaluate the model detection results. The evaluation indicators include: precision, recall rate, F1 score, average precision and small target average precision. The F1 score is calculated as follows:
[0153]
[0154] Where: F1 is the F1 score, P is the precision, which indicates the proportion of positive samples detected as positive, R is the recall, which indicates the proportion of positive samples that are correctly detected, × is the multiplication operator, which indicates the product of P and R, + is the addition operator, which indicates the addition of P and R. It is the harmonic mean of precision and recall.
[0155] If the model achieves the following results in the validation set: average precision mAP ≥ 0.87, small target average precision APs ≥ 0.44, and F1 score F1 ≥ 0.82, the model is considered to be ready for actual deployment.
[0156] During the model inference process, a visual interpretation mechanism based on gradient class activation mapping is introduced to reversely track the convolutional feature responses of the detection target area and generate a heat map of the attention area.
[0157] The Grad-CAM process is as follows: extract the gradient of a key convolutional layer in the detection head, perform weighted averaging of each channel and multiply it with the feature map to generate a thermal distribution map that is consistent with the input image size. The map is then displayed overlaid on the original image using pseudo-color to highlight the key areas of the model.
[0158] Step S500.2: Perform comparative analysis between different models, visualize the test results, and generate a structured test report.
[0159] The DMR-RTDETR model is compared with existing detection models on the same photovoltaic fault dataset. The comparison dimensions include: inference speed, number of parameters, FLOPs calculation amount, mAP, AP50, AP75 and APs indicators.
[0160] The results show that the DMR-RTDETR model maintains mAP ≥ 0.87 and APs ≥ 0.44, with an inference time of less than 40 milliseconds and FLOPs below 48G, which is suitable for edge computing node deployment requirements.
[0161] Configure a graphical user interface to display the detection results in graphical form on the device terminal. The image output content includes: original infrared thermal image, detection frame overlay image, fault type statistical histogram, thermal area response map, image frame sequence browsing and target jump function.
[0162] The model inference results are converted into a standard format detection report, and archivable files are generated in task batches. The detection report content includes: inspection task number and timestamp, total number of images and proportion of abnormal images, number and distribution ratio of various fault targets, summary of detection accuracy and confidence, fault image index and visualization screenshot embedding, and model version information and training indicator summary.
[0163] The inspection report is output in PDF format and uploaded to the cloud platform storage directory. It can also be pushed to the power plant operation and maintenance management system through the RESTful API interface to achieve automatic alarm, risk classification and maintenance scheduling triggering.
[0164] In the above content, in actual application, first, the photovoltaic module infrared image fault detection system uses an industrial-grade multi-rotor drone platform to perform inspection tasks. The flight platform uses the DJI M300RTK drone, and the image acquisition payload uses the H20T infrared thermal imager, which is fixed to the lower part of the drone platform through a dedicated mounting interface. The flight attitude is adjusted in real time by the flight control system to ensure stability during the cruise. The flight mission planning system runs on the ground control terminal, loads the geographic information system data of the photovoltaic power station, and sets the track coverage range, flight strip spacing, flight altitude, route direction, heading overlap rate of not less than 80%, and lateral overlap rate of not less than 70% based on the layout structure of the photovoltaic modules, terrain height difference, obstacle coordinates and boundary information. The planning system generates a track mission file based on the set parameters, and contains the waypoint coordinate sequence, flight altitude, flight speed and mission number. The mission file is transmitted by the control terminal to the drone flight control system for execution.
[0165] Next, the UAV performs an automatic flight mission according to the preset track. The infrared thermal imager collects infrared thermal images at a frame rate of not less than five frames per second during the flight. The resolution of the collected images is not less than 640 by 512 pixels, and the image format is a 16-bit grayscale image. The image data is transmitted in real time to the ground control station or edge computing server via the fifth-generation mobile communication module installed on the UAV using the user datagram protocol. The image receiving end parses the image data and numbers and archives it according to the image frame number, global positioning system location information and acquisition timestamp, and simultaneously records the flight attitude information to achieve a complete mapping of the image and flight status.
[0166] Subsequently, the image data processing process includes normalization, image noise reduction, contrast enhancement and size standardization. Image normalization is used to eliminate grayscale deviations caused by changes in shooting time and ambient temperature. Noise reduction uses a combination of sliding window median filtering and Gaussian filtering to deal with isolated noise points and background interference respectively. Image enhancement includes histogram equalization and edge structure enhancement to enhance the contrast between hot spot areas and the background. The image size is uniformly adjusted to 640 by 640 pixels, using bilinear interpolation scaling, and processing insufficient size areas through edge symmetric filling to ensure consistency of image input.
[0167] Then, the deep detection model is constructed based on the real-time detection transformer structure. The backbone network adopts a residual neural network structure to output multi-layer feature maps. The mid-scale feature maps are selected and input into the multi-path cross-scale attention fusion module to perform channel attention and spatial attention operations to improve the recognition ability of small targets. The fused feature maps are further input into the feature optimization module to perform multi-scale convolution, channel shuffling and re-parameterization processing to optimize the feature expression effect. The decoder structure adopts an attention query matching mechanism to combine the feature maps of two scales to predict the target bounding box and fault category. The detection head adopts a lightweight convolution structure to output the target position coordinates, category probability and embedding vector. The training data set contains three types of samples: hot spot fault, fragmentation fault and diode fault. The image annotation uses a unified format file. The training process sets the number of training rounds to 150 times, and uses a specific optimizer and loss function combination for iterative training. The image input resolution is set to 640 by 640 pixels, and finally the trained model weights are exported.
[0168] Finally, the detection process loads the trained model to perform image inference, and outputs the bounding box coordinates, classification probability and confidence score of each target. The system sets the confidence screening threshold to 50%, and filters the detection results below this threshold. The identified targets are classified into three types of fault types based on the probability vector: hot spot fault, fragmentation fault and diode fault. The detection results are saved in a structured manner as standard data files. The file content includes image number, number of targets, location coordinates, fault type, confidence score and acquisition time. For possible duplicate bounding boxes in the same target area, the system performs non-maximum suppression processing and retains the highest confidence target. All detection results are named by task number and image sequence number, stored in the local data directory or cloud edge server, and uploaded to the operation and maintenance platform through the system interface, supporting image display, detection statistics, alarm triggering and automatic generation of inspection reports.
Claims
1. A photovoltaic module infrared image fault detection method based on drone inspection, characterized in that: The steps include: S100: Configure the UAV platform, plan the flight path, set the aerial photography parameters, and set the infrared thermal image acquisition and infrared image data transmission and storage mechanism; S200, performing grayscale normalization processing, image noise reduction, histogram equalization, edge enhancement processing, and size standardization processing on the infrared image data, and performing data enhancement operations; S300: Configure the deep neural network model architecture, perform feature fusion and recalibration operations, integrate the RePshuffleRoss module to optimize features, configure the decoder structure and detection head, configure the training data set and training parameters, and perform evaluation; S400: Execute reasoning operations and IoU-Aware target confidence screening mechanism, perform target classification and fault type judgment, perform target classification and fault type judgment, and perform data encapsulation and merge duplicate targets; S500: Configure a performance evaluation indicator system and perform quantitative evaluation. Introduce a visual explanation mechanism and generate focused heat maps. Perform comparative analysis between different models, visualize the test results, and generate a structured test report.
2. The photovoltaic module infrared image fault detection method based on drone inspection according to claim 1 is characterized in that: The S100 specifically includes: S100.
1. Configure the UAV platform, plan the flight path, and set the aerial photography parameters; The DJI M300RTK drone is used as the flight platform, and the H20T infrared thermal imager is selected as the image acquisition payload. It is installed on the drone platform through a dedicated mounting interface, and the flight control system performs stable attitude control. A flight mission planning system with a three-dimensional trajectory generation function is configured on the ground control terminal. By importing the geographic information system data of the photovoltaic power station, the photovoltaic module arrangement structure, terrain height difference data, fixed obstacle coordinates and boundary information are loaded; According to the image acquisition mission scope, set the flight mission parameters, including: track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate not less than 80% and lateral overlap rate not less than 70%; Generate a flight mission file based on the track coverage boundary, flight strip spacing, altitude setting, route direction, heading overlap rate, and side overlap rate, including: waypoint coordinate sequence, flight altitude, speed, and mission number information, and upload the flight mission file to the UAV flight control system as an execution instruction; The standard flight altitude range of the drone is set to 40 meters to 110 meters. The flight altitude is dynamically adjusted according to the terrain conditions and the distribution density of photovoltaic modules. The flight speed range is set to 1.5 meters per second to 3.0 meters per second. The flight control system controls the flight attitude stability, and the roll angle and pitch angle are set to no more than 5 degrees. S100.
2. Set up the mechanism for infrared thermal image acquisition and infrared image data transmission and storage; The UAV will automatically fly according to the flight mission file. During the flight, the infrared thermal imager will capture infrared thermal images at a frame rate of not less than 5 frames per second. The captured images are infrared radiation grayscale images of 640×512 pixels or larger, and the image format is 16-bit grayscale images; The fifth-generation mobile communication module integrated in the UAV payload transmits infrared thermal image data in real time to the ground control station or edge computing server. The image data transmission adopts the User Datagram Protocol. The receiving end analyzes the image data in real time and archives and stores it according to the image number, drone global positioning system location information and flight timestamp; During the image acquisition process, the flight attitude data and GPS positioning data output by the drone's inertial measurement unit are collected synchronously, and each frame of image is bound to the image frame number, GPS coordinate information, flight attitude information and image acquisition timestamp metadata fields.
3. The photovoltaic module infrared image fault detection method based on drone inspection according to claim 1 is characterized in that: The S200 specifically includes: S200.
1. Grayscale normalization of infrared images, image noise reduction, histogram equalization, and edge enhancement are performed; Receive infrared radiation grayscale image data transmitted from the ground control station or edge computing server and perform pixel-by-pixel normalization processing. Use the grayscale normalization algorithm to transform the image data. Specifically: Where: I norm (x, y) is the normalized grayscale value at the pixel position (x, y), I(x, y) is the original grayscale value at the pixel position (x, y), and I min is the minimum gray value in the entire infrared image, I max is the maximum grayscale value in the entire infrared image, x is the horizontal pixel coordinate of the image, and y is the vertical pixel coordinate of the image; For the thermal noise and patch noise in the original infrared thermal image caused by image sensor noise, flight vibration jitter or electromagnetic interference, median filtering and Gaussian filtering are used in sequence to perform image noise reduction. A median filter with a 3×3 sliding window is used to replace the grayscale value of each pixel with the median grayscale value of its neighboring pixels, and a two-dimensional Gaussian filter with a standard deviation of σ = 1.2 is used to perform convolution smoothing on the image. Perform full image grayscale histogram equalization on the image after noise reduction, and use Laplace edge enhancement operator to perform edge enhancement on the image; S200.
2. Normalize the size of the infrared image and perform data enhancement. For images with inconsistent aspect ratios, bilinear interpolation resampling algorithm is used for scaling. If the image boundary is still less than 640×640 pixels, the missing area is filled by edge symmetric filling without introducing redundant information. Perform data augmentation on the standardized images and generate diverse training image pseudo samples by performing random horizontal flipping, random rotation, random brightness perturbation, local occlusion simulation, and Gaussian noise perturbation simulation. Generate a label file required for target detection for each processed image. The label file contains: hot spot fault location bounding box, hot spot fault type, image index number and task number.
4. The photovoltaic module infrared image fault detection method based on drone inspection according to claim 1 is characterized in that: The S300 specifically includes: S300.
1. Configure the deep neural network model architecture and perform feature fusion and recalibration operations; Through the RT-DETR structure, the neural network is configured. The model structure includes: backbone feature extraction network, feature fusion encoding module and decoding detection module; ResNet-50 is used as the backbone feature extraction network, which outputs five layers of multi-scale feature maps P1, P2, P3, P4, and P5. Each layer of feature maps corresponds to a different receptive field and resolution, respectively extracting shallow edge features and deep semantic features. The two mid-scale feature maps P3 and P4 are selected as input to the multi-path cross-scale attention feature fusion module; In the multi-path cross-scale attention feature fusion module, two parallel pathways are used to extract global features and local detail features respectively, enhancing the network's sensitivity to small objects of different scales. The global pathway introduces a spatial attention mechanism, which performs maximum pooling and average pooling on the input feature map in the channel dimension to obtain a spatial attention map. The local path introduces a channel-by-channel weighted channel attention mechanism, combined with a multi-layer perceptron to learn the feature importance of each channel. After global and local fusion, a recalibrated fusion feature map F is obtained. MPCAFF ; S300.2, integrates the RePshuffleRoss module optimization features and configures the decoder structure and detection head; The fused feature map is input into the RePshuffleRoss module to perform multi-path convolution and channel shuffling operations. During the training phase, three-way parallel convolution operations of 1×1, 3×3, and 5×5 are used to extract edge textures and deep patterns. Channel shuffling operations are performed to break inter-channel dependencies. During the inference phase, the multi-path structure is compressed into an equivalent single-path convolution through reparameterization technology. The fused and optimized feature map is input into the Transformer decoder structure, and the IoU-AwareQueryMatching strategy is used to initialize the target query. Two scale feature maps with resolutions of 40×40 and 80×80 are selected and input into the decoder, and the RepConv optimization detection head structure is introduced. RepConv combines the multi-layer structure of the training phase into a single layer of efficient convolution. The detection head output includes: target bounding box coordinates (x, y, w, h) and fault category probability, and outputs a set of embedding vectors containing: target location and category information; S300.
3. Configure the training data set and training parameters, and perform evaluation; The photovoltaic fault detection dataset is used as the training set, which includes hot spot faults, fragmentation faults and diode faults; Use the LabelImg tool to annotate the bounding boxes of the objects in the image and generate a COCO format label file. After training is completed, export the model weight file θ final , and evaluate the model detection performance. The performance of the model on the validation set is quantitatively evaluated using the average precision, the average precision under the condition of IoU threshold of 0.5, the average precision under the condition of IoU threshold of 0.75, and the average precision of small targets. Specifically: Where: A is the area of the predicted bounding box, B is the area of the true bounding box, A∩B is the area of the intersection of the predicted box and the true box, A∪B is the area of the union of the predicted box and the true box, and IoU is the weight ratio between the predicted result and the true result; If the model achieves the performance standards of mAP ≥ 0.87 and APs ≥ 0.44 on the PVFD photovoltaic fault detection dataset, it is considered a trained model that meets deployment requirements.
5. The photovoltaic module infrared image fault detection method based on drone inspection according to claim 1 is characterized in that: The S400 specifically includes: S400.
1. Execute reasoning operations and IoU-Aware target confidence screening mechanism, and perform target classification and fault type judgment; The infrared thermal image after grayscale normalization, noise suppression, enhancement processing and size standardization is input into the trained DMR-RTDETR network model for feature extraction, feature fusion, feature optimization and decoding inference; The backbone network ResNet-50 outputs multi-scale feature maps for feature extraction, the MPCAFF module performs global-local attention fusion for feature fusion, the RePshuffleRoss module performs channel shuffling and reparameterization for feature optimization, and the Transformer decoder and RepConv detection head output detection results for decoding inference. The output of the DMR-RTDETR model includes: the predicted bounding box coordinates of the detected target, the category label probability distribution, and the target confidence value; The target confidence value output by the IoU-AwareQueryMatching mechanism is used as the screening basis, and the detection result screening threshold is set to 0.
5. The target detection box D with a target confidence value greater than or equal to 0.5 is retained. i , remove low-quality, low-overlap or background misdetected areas, and calculate the target confidence of the IoU-Aware mechanism, specifically: C i =cls i ×IoU i Where: C i is the target detection box D i Confidence, cls i is the target detection box D i Class probability, IoU i is the target detection box D i The intersection-over-union ratio with the ground-truth bounding box, D i is the i-th target detection box output by the model; S400.
2. Execute target classification and fault type determination, and perform data encapsulation and merge duplicate targets; The retained target detection box is selected according to the maximum value of the category probability vector to classify the target fault type. The fault types are divided into: hot spot fault, fragmentation fault and diode fault. The final output of the target includes: position bounding box coordinates, fault type label, confidence value and image index number; The detection results corresponding to each infrared image are formatted and saved as a structured data file, encapsulated in JSON format. The fields include: image file number, number of detected targets, bounding box coordinates of each target, fault type label, target confidence value and detection timestamp; If there are multiple predicted bounding boxes in the same image that are repeatedly annotated near the same position, NMS is used to merge the redundant boxes; Sort all candidate boxes from large to small according to the confidence value, and select the box with the maximum confidence D from the sorted list mas , add it to the final output set, and combine it with D mas Other candidate boxes whose IoU is greater than 0.5 are eliminated, and the steps are repeated 2-3 times.
6. The photovoltaic module infrared image fault detection method based on drone inspection according to claim 1 is characterized in that: The S500 specifically includes: S500.
1. Configure a performance evaluation indicator system, perform quantitative evaluation, introduce a visual explanation mechanism, and generate a focused heat map. A hybrid evaluation index system is used to comprehensively evaluate the model detection results. The evaluation indicators include: precision, recall rate, F1 score, average precision and small target average precision; During the model inference process, a visual interpretation mechanism based on gradient class activation mapping is introduced to reversely track the convolutional feature response of the detection target area and generate a heat map of the attention area. Grad-CAM processing involves extracting the gradient of a key convolutional layer in the detection head, performing a weighted average of each channel and multiplying it with the feature map to generate a heat map with the same size as the input image. This is then overlaid on the original image using pseudo-color to highlight the key areas of the model. S500.
2. Perform comparative analysis between different models, visualize the test results, and generate a structured test report. The DMR-RTDETR model was compared with existing detection models on the same photovoltaic fault dataset. The comparison dimensions included inference speed, number of parameters, FLOPs computation, mAP, AP50, AP75, and APs indicators. Configure a graphical user interface to display the detection results in graphical form on the device terminal. The image output content includes: original infrared thermal image, detection frame overlay image, fault type statistical histogram, thermal area response map, image frame sequence browsing and target jump function; The model inference results are converted into a standard format detection report, and archivable files are generated in task batches. The detection report content includes: inspection task number and timestamp, total number of images and proportion of abnormal images, number and distribution ratio of various fault targets, summary of detection accuracy and confidence, fault image index and visualization screenshot embedding, and model version information and training indicator summary.
Citation Information
Cited By
Inspection method and device of self-inspection infrared unmanned aerial vehicle
CN120831362A
Unmanned aerial vehicle target detection and tracking method
CN121170655A
Insulator image segmentation method based on infrared enhanced image of unmanned aerial vehicle
CN121190504A
Visual guidance posture correction and safety control method for photovoltaic cleaning robot
CN122239770A