A method and system for detecting infrared defects in a hybrid network photovoltaic module
Patent Information
- Application Number
- CN202611001556.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
这类方法的处理流程复杂,通常需要设计多个独立的模块,导致整体模型体积大、计算量大
[0025]本发明提出了通道剪枝和INT8量化的组合式轻量化改造方案,用于将训练好的高精度模型部署到极端边缘设备,通道剪枝基于L1范数评估通道重要性,裁剪冗余通道,INT8量化将权重量化为8位整型,为保护小目标检测精度,检测头的回归和分类分支保留FP32精度。
Smart Images

Figure CN122824104A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of solar photovoltaic panel defect detection technology, and more specifically to an infrared defect detection method and system for hybrid network photovoltaic modules. Background Technology
[0002] As the core component of a photovoltaic power plant, photovoltaic modules often experience localized heating and infrared defects such as hot spots during long-term operation due to factors such as performance degradation, external shading, bypass diode failure, or terminal faults. These defects severely impact power generation efficiency and can even lead to safety accidents. Therefore, efficient and accurate defect detection of photovoltaic modules is crucial for the operation and maintenance of power plants.
[0003] Currently, the main methods for detecting infrared defects in photovoltaic modules include the following categories: Traditional machine vision methods: These methods extract features from photovoltaic modules and identify defects by preprocessing infrared images, such as threshold segmentation and edge detection. However, these methods are highly dependent on ambient light, temperature, and shooting angle, and are prone to misjudgment and missed detection in complex backgrounds.
[0004] Improved target recognition algorithms based on deep learning: These methods are based on general target detection algorithms (such as YOLO, SSD, Faster R-CNN, etc.) and adapt to the defect recognition of photovoltaic modules by adjusting the network structure or parameters. However, these methods are generally limited by the original design intent of the algorithm, and their robustness and recognition accuracy are poor when facing small targets (such as tiny hot spots), large scale variations, and complex background interference in photovoltaic defects.
[0005] The segmentation-then-classification method: This method first segments the photovoltaic module from the infrared image, and then classifies the segmented module regions to determine the presence and type of defects. This method has a complex processing flow, typically requiring multiple independent modules, resulting in a large overall model size and high computational load. In industrial applications, it places high demands on the computing power and storage of hardware, leading to high deployment and operating costs, and making it difficult to guarantee real-time performance.
[0006] Therefore, how to provide a photovoltaic module infrared defect detection method that can take into account high accuracy, strong robustness, small target sensitivity, compact model and high computational efficiency is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a method and system for infrared defect detection of photovoltaic modules using hybrid networks, in order to solve the problems in the background art.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: On one hand, this invention discloses a method for infrared defect detection of photovoltaic modules using hybrid networks, comprising the following specific steps: Infrared images of photovoltaic modules were acquired, and a defect dataset containing diode faults, hot spot faults, and terminal faults was constructed. A hybrid network model is constructed, comprising, in sequence, an infrared image mapping layer, an infrared image fast backbone network, a channel reconstruction neck network, and a small target no-anchor-frame detection head. The infrared image mapping layer adaptively maps the input three-channel pseudo-color infrared image into a single-channel temperature feature map using a learnable 1×1 convolutional kernel. The infrared image fast backbone network, based on a structure integrating Vision Transformer and convolutional neural networks, extracts and outputs multi-scale feature maps from the single-channel temperature feature maps. The channel reconstruction neck network upsamples, stitches, reduces dimensionality, and downsamples the multi-scale feature maps, extracts channel information through a channel separation processing block (CSB), and outputs the reconstructed multi-scale feature maps. The small target no-anchor-frame detection head uses a detection head without preset anchor frames and employs Normalized Wasserstein distance (NWD) as part of the loss function to predict the location and category of defects based on the reconstructed multi-scale feature maps. The hybrid network model is trained using the defect dataset. The infrared image of the photovoltaic module to be inspected is input into the trained hybrid network model, which outputs the category and location of the defect.
[0009] This invention replaces the traditional manual coefficient-based pseudo-color to single-channel conversion method with learnable 1×1 convolutional kernel adaptive mapping, enabling it to adapt to different pseudo-color modes (white-hot, iron-red, lava), effectively removing pseudo-color noise while retaining key temperature amplitude information, thus improving the algorithm's robustness to different infrared imaging devices. By fusing the VisionTransformer and convolutional neural network backbone, it simultaneously captures local temperature gradient features and global contextual dependencies, solving the problem of small target defects being easily lost in complex backgrounds. Through the channel reconstruction neck network and CSB module, parameters and computational load are reduced during multi-scale feature fusion, enhancing the ability to capture key defect information. By using an anchor-free detection head combined with NWD loss, the limitations of preset anchor frames on defect scale changes are eliminated, and the problem of traditional IoU being overly sensitive to small target position deviations is overcome.
[0010] Preferably, in the above-mentioned hybrid network method for infrared defect detection of photovoltaic modules, the infrared image fast backbone network is designed based on the EfficientFormer structure, and the specific processing steps are as follows: The single-channel temperature feature map is processed through multiple stages in sequence. In the early stage, a convolutional neural network is used to extract local features, and in the later stage, a Vision Transformer is used to capture global contextual dependencies. The multi-scale feature map includes outputs at least three different scales, which are fed into the channel reconstruction neck network for feature fusion processing.
[0011] This invention designs the backbone network as a progressive structure of early CNN + later ViT. It leverages the low computational cost of CNN to quickly extract local edge and temperature gradient features, while utilizing the global self-attention mechanism of ViT to capture the long-range dependency between defects and the background. This avoids the problems of insufficient receptive field of a single CNN or excessive computational cost of a single ViT. Simultaneously, it outputs feature maps at three different scales, enabling the network to simultaneously detect large-sized terminal block faults (20×20 scale) and small-sized hot spot faults (80×80 scale), improving the detection accuracy of multi-scale targets.
[0012] Preferably, in the above-mentioned hybrid network photovoltaic module infrared defect detection method, the specific processing steps of the channel reconstruction neck network are as follows: The smaller-scale feature maps output by the backbone network are upsampled to the same size as the largest-scale feature maps and then stitched together along the channel dimension. Channel dimensionality reduction is performed on the stitched feature map using a 1×1 convolutional layer; Multi-scale downsampling is performed on the dimensionality-reduced feature map to generate multiple feature maps of different scales; The feature map at each scale is fed into the Channel Separation Processing Block (CSB) for processing. The CSB processing flow includes: splitting the input feature map into two parts in the channel dimension, one part being processed by the Spatial Reconstruction Unit (SRU) and the Channel Reconstruction Unit (CRU) in sequence, and then concatenating it with the other part before outputting.
[0013] This invention expands the three-scale features of the backbone network into five-scale features through cascaded operations of upsampling concatenation, 1×1 dimensionality reduction, and multi-scale downsampling. This enables the network to obtain a rich receptive field from details to semantics, adapting to extreme scale variations in photovoltaic defects. A channel splitting block (CSB) is introduced to process the feature map from a spatial perspective. Through operations such as convolution, channel segmentation, spatial reconstruction, and channel reconstruction, the ability to capture key information is enhanced while reducing parameters and computational load. Especially when dealing with small target defects, it effectively highlights the features of small targets and improves the recognizability of small target defects in complex backgrounds.
[0014] Preferably, in the above-mentioned hybrid network photovoltaic module infrared defect detection method, the processing of the spatial reconstruction unit includes sequentially performing Depthwise convolution, batch normalization, and Pointwise convolution; the processing of the channel reconstruction unit includes sequentially performing global average pooling, dimensionality-reduced 1×1 convolution, ReLU activation, dimensionality-upgraded 1×1 convolution, and Sigmoid activation.
[0015] This invention clarifies the specific operational sequences of the Spatial Reconstruction Unit (SRU) and the Channel Reconstruction Unit (CRU). The SRU, through a combination of depthwise and pointwise convolutions, captures the local spatial dependencies of defect edges and temperature gradients with extremely low computational overhead (approximately 1 / 9 the number of parameters of a standard convolution). The CRU, through a channel attention mechanism using global average pooling and two layers of 1×1 convolutions, automatically learns the importance weights of each feature channel, suppressing pseudo-color residue and background noise channels, and enhancing feature channels related to defects such as diodes and hot spots. When used in series, and in conjunction with channel dimension segmentation, both achieve dual feature recalibration in both spatial and channel dimensions.
[0016] Preferably, in the above-mentioned hybrid network photovoltaic module infrared defect detection method, the small target no-anchor-frame detection head is a NASFCOS detection head obtained based on neural network search technology, which includes a regression branch, a classification branch and a centrality branch. The regression branch and the classification branch adopt deformable convolution. The normalized Wasserstein distance (NWD) is used instead of the traditional intersection-over-union ratio (IoU) to measure the error between the predicted target bounding box and the real target bounding box, and it is used to construct the total loss function.
[0017] This invention employs a NASFCOS detection head obtained through NAS search. Compared to manually designed FCOS, its convolutional kernel combination (3×3 deformable convolution for classification branches and 5×5 deformable convolution for regression branches) is the optimal configuration adapted to the irregular shapes of photovoltaic defects, obtained through extensive searches. The deformable convolution can dynamically adjust the sampling point position according to the actual shape of the defect, solving the problem of insufficient feature extraction capability of traditional convolution for slender or irregular defects (such as hot spot propagation and irregular heating in diode regions). Normalized Wasserstein distance (NWD) is used instead of IoU, modeling the target box as a two-dimensional Gaussian distribution. It considers the overlap area, shape similarity, and center point distance of the boxes, making it insensitive to pixel-level small target position deviations, allowing the loss function to "focus" on small target defects.
[0018] Preferably, in the above-mentioned hybrid network photovoltaic module infrared defect detection method, the total loss function of the small target frameless detection head is a weighted sum of classification loss, centrality loss, and NWD loss: ; NWD loss calculation method: Target bounding box modeling: converting the actual bounding box into a target bounding box. With prediction box Each is modeled as a two-dimensional Gaussian distribution: ; ; Wasserstein distance calculation: ; Normalization and loss: ; ; in, Use Focal Loss or binary cross-entropy loss. Using binary cross-entropy loss, , , These are the weighting coefficients; the NWD is obtained by calculating the normalized Wasserstein distance after modeling the target box as a two-dimensional Gaussian distribution. The maximum value of the Wasserstein distance between all pairs of ground truth bounding boxes and predicted bounding boxes in the training dataset is used as a normalization constant to ensure... The value of is in the range of [0,1].
[0019] This invention presents the specific composition of the total loss function and the calculation method of each component. The classification loss (FocalLoss) effectively solves the problem of extreme imbalance between positive and negative samples (defective samples only account for a small portion of the image); the centrality loss suppresses low-quality predicted boxes far from the target center; and the NWD loss provides robust bounding box regression supervision for small target positional deviations. The three parts are weighted with α=1.0, β=0.5, and γ=2.0, highlighting the importance of the NWD loss for small target detection. This loss function design allows the model to simultaneously optimize classification accuracy, predicted box quality, and positional accuracy during training.
[0020] Preferably, in the above-mentioned hybrid network photovoltaic module infrared defect detection method, when constructing the photovoltaic module infrared defect dataset, an infrared thermal imager mounted on a drone is used to acquire images. The temperature range of the infrared thermal imager is -20℃ to 150℃, the drone's shooting height is 5 to 8 meters, and the shooting angle is vertically downward. The defect categories include diode faults, hot spot faults, and terminal faults, and their judgment criteria are as follows: Diode failure: A blocky temperature rise area appears that matches the installation location of the bypass diode, with a temperature rise difference ≥10℃; Hot spot fault: Spot-like or small-area temperature rise areas appear, with a temperature rise difference ≥8℃; Terminal block failure: The overall temperature rises at the terminal block location, and the overall temperature of the component is ≥5℃ higher than that of a normal component.
[0021] This invention defines data acquisition standards and defect quantification criteria, making the method repeatable and engineering-applicable. The drone is used to capture images from a height of 5-8 meters, at a vertical downward angle, ensuring that the photovoltaic modules occupy ≥70% of the image and that temperature characteristics are clearly defined. The criteria for judging three types of defects (e.g., diode fault temperature rise difference ≥10℃, hot spot fault ≥8℃, and terminal fault overall temperature rise ≥5℃) provide clear quantification thresholds, avoiding subjective judgment differences. The dataset constructed using this standard contains 6096 images, and the distribution of the three types of defect samples closely resembles the proportion of actual power plant faults, providing a high-quality, highly representative data foundation for model training.
[0022] Preferably, in the above-mentioned hybrid network method for infrared defect detection of photovoltaic modules, the infrared image fast backbone network of the hybrid network model is replaced with a lightweight backbone network based on MobileViT. The MobileViT backbone network uses depthwise separable convolution to extract local features and fuses global context information through micro Transformer blocks. The multi-scale feature map includes outputs at least three different scales, which are fed into the channel reconstruction neck network for feature fusion processing.
[0023] This invention provides a lightweight backbone network replacement scheme, which replaces the main scheme's EfficientFormer backbone network with MobileViT. MobileViT combines the depthwise separable convolution (lightweight) of MobileNet with the global modeling capabilities of ViT. This replacement scheme is particularly suitable for edge computing scenarios such as UAV-borne real-time detection and embedded inspection terminals, and can still achieve near real-time defect detection on devices with limited computing power.
[0024] Preferably, in the above-mentioned hybrid network method for infrared defect detection of photovoltaic modules, the trained hybrid network model undergoes a lightweight modification, the lightweight modification including: Channel pruning: Calculate the L1 norm of each channel in the convolutional layer of the network, prune redundant channels below a preset threshold, and fine-tune the pruned model; INT8 quantization: For the pruned model, post-training quantization is performed using a calibration set to convert the weights and activation values in the model from floating-point to integer.
[0025] This invention proposes a lightweight modification scheme combining channel pruning and INT8 quantization to deploy a trained high-precision model to extreme edge devices. Channel pruning evaluates channel importance based on the L1 norm and removes redundant channels. INT8 quantization quantizes the weights into 8-bit integers. To protect the detection accuracy of small targets, the regression and classification branches of the detection head retain FP32 accuracy.
[0026] On the other hand, the present invention discloses a hybrid network photovoltaic module infrared defect detection system, which uses the method described above to perform defect detection on the infrared image of the photovoltaic module.
[0027] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a photovoltaic module infrared defect detection method and system. Through a learnable mapping layer, a backbone network integrating CNN and Vision Transformer, a neck network containing channel separation processing blocks, and an anchorless detection head based on NWD loss, it achieves high-precision identification of diode faults, hot spots, and terminal faults. This method adaptively removes pseudo-color noise, considers both local and global features, and enhances small target detection capabilities through multi-scale feature fusion. The NWD loss effectively overcomes the sensitivity of traditional IoU to small target position deviations. It can be flexibly deployed in high-precision cloud scenarios or edge real-time detection scenarios such as drones and embedded terminals, providing a high-precision, high-efficiency, and practical solution for the intelligent operation and maintenance of photovoltaic power plants. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0029] Figure 1 The network structure for the HSN algorithm; Figure 2 This is a schematic diagram of the infrared image mapping layer; Figure 3 Infrared image fast backbone network structure diagram; Figure 4 Channel separation processing block network structure diagram; Figure 5 Flowchart of the method of this invention; Figure 6 Schematic diagram of the detection head network. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Example 1: This example discloses an infrared defect detection method for photovoltaic modules using a hybrid network, such as... Figure 5 As shown, the details are as follows: I. Construction of a standardized infrared defect dataset for photovoltaic modules; 1.1 Data Acquisition: 1. Data Acquisition Equipment: A DJI M300RTK drone equipped with an FLIRT640 infrared thermal imager was used to take on-site photos of photovoltaic power plants, and to assist in crawling publicly available infrared defect images of photovoltaic power plants from the Internet. The FLIRT640 infrared thermal imager has a temperature range of -20℃ to 150℃, an infrared resolution of 640×480 pixels, and supports three common pseudo-color modes: white-hot, iron red, and lava.
[0032] 2. Shooting Standards: The drone shooting height should be controlled between 5 and 8 meters, the shooting angle should be vertically downward (0° nadir), and the shooting environment should be during the normal operation period of the photovoltaic power station (sunlight intensity ≥ 200W / The operating temperature of the components is 25~60℃, and adverse weather conditions such as rain and fog should be avoided. The number of photovoltaic modules covered by a single image is 4~16, ensuring that the components account for ≥70% of the image.
[0033] 3. Total Data Volume and Sample Composition: The final dataset contains 6096 valid infrared images, including 2132 normal samples (images of defect-free photovoltaic modules) and 3964 defective samples. The defective samples include three types of typical photovoltaic module faults: 1286 diode fault samples, 1895 hot spot fault samples, and 783 terminal fault samples. The criteria for judging these three types of faults are as follows: ① Diode failure: Localized block-shaped temperature rise areas appear in the photovoltaic module, with a temperature rise difference ≥10℃, and the location of the area matches the installation location of the bypass diode; ② Hot spot fault: Point-like / small area-like temperature rise areas appear on the surface of photovoltaic modules, with a temperature rise difference ≥8℃, caused by module shading and performance degradation; ③ Terminal fault: The photovoltaic module terminal temperature rises as a whole, and the overall module temperature is ≥5℃ higher than that of a normal module.
[0034] 1.2 Data Labeling: 1. Labeling tool: Use the LabelImg labeling tool to label defect targets. The labeling format is VOC format, which is compatible with the MMDetection framework.
[0035] 2. Annotation box rules: Use axis-aligned rectangular annotation boxes. The annotation box must completely enclose the temperature rise area of the defect, and the boundary of the box coincides with the edge pixels of the temperature rise area. For point-like small target hot spot defects, the minimum side length of the annotation box should not be less than 8 pixels to avoid the model training failure due to insufficient pixels in the small target annotation box.
[0036] 3. Labeling content: Label the location (coordinates of the label box) of the defect and the defect type (diode / diode failure, hotspot / hot spot failure, terminal / terminal failure) in each image. No labeling is required for images without defects.
[0037] 1.3 Dataset Partitioning: The labeled dataset was randomly divided into training, validation, and test sets in a ratio of 7:2:1, with 4267 images in the training set, 1219 in the validation set, and 610 in the test set. This division ensured that the proportion of the three types of defective samples in each dataset was consistent with the total dataset, thus avoiding sample imbalance.
[0038] 1.4 Processing logic for pseudo-color to single-channel conversion: 1. Before converting to single channel: The acquired three-channel pseudo-color infrared images are preprocessed, including image size normalization (uniformly adjusted to 640×640 pixels) and pixel value normalization (scaling pixel values to the 0~1 range). No other enhancement processing is performed to ensure the integrity of the original temperature features.
[0039] 2. Single-channel conversion process: The infrared image mapping layer (1×1 convolutional kernel layer) of the backbone network input layer realizes the adaptive mapping from three channels to one channel, replacing the traditional method of manually giving coefficients. The mapping process automatically optimizes the convolutional kernel weights through backpropagation of the network, retains the effective information of temperature amplitude, and removes noise interference caused by false colors.
[0040] 3. After conversion to single channel: Gaussian blur denoising is performed on the output single channel tensor (kernel size 3×3, standard deviation 0.5), and then it is sent to the feature extraction module for further processing.
[0041] II. Constructing the HSN (HybridSolarNet) algorithm structure; The HSN algorithm is functionally divided into three parts: a fast backbone network for infrared image processing, a channel reconstruction neck network, and a small target frameless detection head. The overall architecture is a single-stage target detection architecture, with each part working collaboratively to achieve feature extraction, fusion, and accurate detection of infrared defects in photovoltaic modules. The network structure is as follows: Figure 1As shown, the input is a three-channel pseudo-color infrared image → infrared image mapping layer → infrared image fast backbone network (Stage 1~Stage 4) → channel reconstruction neck network → small target no-anchor-frame detection head → output defect location / category / confidence. Input: A three-channel pseudo-color infrared image with dimensions of 640×640×3 (pixel values normalized to 0~1). ; 2.0 Infrared image mapping layer, such as Figure 2 As shown, (single-channel conversion module); Function: Adaptively maps a three-channel pseudo-color infrared image to a single-channel temperature feature map, removes pseudo-color noise, and retains temperature amplitude information; Structure: Consists of a 1×1 convolutional layer + a BatchNorm layer + a ReLU activation layer; Parameter and feature size variations: Input: 640×640×3 Convolution kernel: 1×1×3×1 (kernel size 1×1, input channels 3, output channels 1) Output: 640×640×1 (single-channel temperature feature map) Training logic: The convolutional kernel weights are automatically optimized by backpropagation of the entire network, adapting to different pseudo-color modes such as white-hot / iron-red / lava, without the need for manual setting of mapping coefficients.
[0042] 2.1 Infrared Image Fast Backbone Network; The backbone network is responsible for feature extraction and encoding of infrared images. It is designed to address the false-color characteristics and small-size, dispersed distribution of defects in photovoltaic module infrared images. The structure is as follows: Figure 3 As shown.
[0043] 1. Core Processing Steps: The three-channel color pseudo-color image output from the infrared camera is processed through an infrared image mapping layer (a 1×1 convolutional layer with a single-channel output kernel, structured as follows). Figure 2 As shown, the encoding is a single-channel tensor. This mapping layer adjusts the mapping relationship from three channels to one channel through the backpropagation process of the entire network, adapts to different pseudo-color modes (white heat, iron red, lava), retains the effective information of temperature amplitude, and avoids noise interference caused by pseudo-color. 2. Feature extraction architecture: Based on the EfficientFormer structure design, it combines the advantages of VisionTransformer and Convolutional Neural Network (CNN). The CNN part captures local temperature features in the early stage of the network, while the VisionTransformer part captures the long-distance dependency between defects and the background through global context modeling capabilities, thereby improving the detection accuracy of small target defects. 3. Output features: The network finally outputs feature maps of three scales: 80×80, 40×40, and 20×20, which are fed into the neck network for feature fusion processing.
[0044]
[0045] 2.2 Channel Reconstruction of Neck Network; The neck network is responsible for fusing and reconstructing the multi-scale feature maps output by the backbone network, enhancing the ability to capture key defect information. Its structure is as follows: Figure 4 As shown.
[0046] 1. Feature fusion processing: The two smaller-scale feature maps, 40×40 and 20×20, output by the backbone network are upsampled (bilinear interpolation) to the same scale as the 80×80 large-scale feature map. Then, the three scale feature maps are concatenated along the channel dimension. 2. Channel dimensionality reduction and scale expansion: The concatenated feature maps are reduced in number of channels by 1×1 convolutional layers to reduce computation. Then, they are downsampled in parallel to generate five feature maps of different scales: 80×80, 40×40, 20×20, 10×10, and 5×5. This allows the network to obtain different receptive fields and adapt to defect targets of different sizes. 3. Channel Information Extraction: The feature maps at five scales are fed into the Channel Separation Block (CSB) for processing. The CSB processes the feature maps from a spatial perspective, reducing parameters and computational load while enhancing the ability to capture key information. The CSB processing flow is as follows: the input feature map is convolved by 1×1 and then split into 1 / 4 and 3 / 4 parts in the channel dimension. The 1 / 4 part is processed by the Spatial Reconstruction Unit (SRU) and the Channel Reconstruction Unit (CRU) in sequence. After processing, it is concatenated with the 3 / 4 part in the channel dimension, and then convolved by 1×1 before being output. Finally, the feature maps processed by the five CSBs are sent to the detection head.
[0047] The module connection order is as follows: Backbone outputs P2 (80×80×96), P3 (40×40×192), P4 (20×20×384) → upsampling and concatenation → 1×1 convolutional dimensionality reduction → multi-scale downsampling → CSB processing → output feature maps of five scales from P2 to P6 Upsampling and stitching: P3, after bilinear interpolation and upsampling, is 80×80×192; P4, after two bilinear interpolations and upsampling, is 80×80×384. Concatenating P2 with P2 along the channel dimension yields 80×80×(96+192+384)=80×80×672 1×1 convolution dimensionality reduction: The convolution kernel is 1×1×672×256, and the output is 80×80×256, reducing the number of channels and computational load. Multi-scale downsampling: Performing three 3×3 convolutions (stride 2) on an 80×80×256 array yields: P2: 80×80×256; P3: 40×40×256; P4: 20×20×256; P5: 10×10×256; P6: 5×5×256; CSB (Channel Separation Block) internal process: Input: H×W×C (C=256); Step 1: Channel segmentation → Split along the channel dimension into H×W×64 (1 / 4) and H×W×192 (3 / 4); Step 2: SRU (Spatial Reconfiguration Unit) processes the H×W×64 branch: Operation: DepthwiseConv3×3→BN→ReLU→PointwiseConv1×1→BN; Function: Captures local spatial dependencies and enhances defect edge and temperature gradient features; Step 3: CRU (Channel Reconstruction Unit) processes SRU output: Operation: Global average pooling → 1×1 convolution (dimensionality reduced to C / 16) → ReLU → 1×1 convolution (dimensionality restored to C / 4) → Sigmoid; Functions: Learn channel attention weights, suppress noise channels, and enhance the weights of defect-related channels; Step 4: Concatenation and Output → The CRU output and the H×W×192 branch are concatenated in the channel dimension and restored to H×W×256 by 1×1 convolution; Output: Five scales P2~P6, with dimensions of 80×80×256, 40×40×256, 20×20×256, 10×10×256, and 5×5×256 respectively; 2.3 Small target detection head without anchor frame; The detection head is responsible for mapping the feature map output by the neck network to the location and category of the defect target. It employs an anchor-free design to adapt to the diverse sizes of photovoltaic defects, and its structure is as follows: Figure 6 As shown.
[0048] 1. Detection Head Architecture: The NASFCOS detection head, obtained through Neural Network Search (NAS) technology, is a fully convolutional structure without predefined anchor boxes. This simplifies model design and training by eliminating the need to predefine anchor box parameters. Each detection head contains three branches: regression, classification, and centroid calculation. These three branches consist of variable convolutional layers with different kernel sizes. The centroid calculation and classification branches share the same weights, balancing detection accuracy and computational efficiency.
[0049] 2. Loss Function Design: Normalized Wasserstein distance (NWD) is used instead of the traditional intersection-over-union ratio (IoU) to measure the error between the predicted target bounding box and the labeled target bounding box. NWD models the target box as a two-dimensional Gaussian distribution, comprehensively considering the overlapping area, shape and position information of the target box. By normalizing, the influence of scale factors is eliminated, which has stronger robustness to the positional deviation and scale change of small target defects. It effectively solves the problem that IoU is sensitive to the positional deviation of small target defects, causing the network to ignore small targets.
[0050] The module connection order is as follows: Neck output P2~P6 → Shared convolutional feature extraction → Classification / central branch + regression branch → Decoding prediction results; Infrastructure improvements compared to FCOS: Basic architecture: Inheriting the anchorless design of FCOS, without preset anchor boxes, it directly predicts the offset and class of each point on the feature map; NAS Improvements: Automatic selection of convolution kernel combinations through Neural Architecture Search (NAS): 3×3 deformable convolution is used for classification / central branches, and 5×5 deformable convolution is used for regression branches to adapt to defective and irregular shapes; Shared convolutional layers: The classification branch and the centrality branch share the first two 3×3 convolutional layers, reducing the number of parameters by about 30%; Feature alignment: Add a deformable convolution (DeformableConv) before the regression branch to solve the problem of feature map offset from target position; Branching functionality and output: Classification branch: Input H×W×256 → 4 layers of 3×3 deformable convolution → Output H×W×3 (3 types of defects + background, activated by Sigmoid); Centrality branch: Shares the first two convolutional layers with the classification branch → Outputs H×W×1 (centrality score, suppressing low-quality predictions far from the target center). Regression branch: Input H×W×256 → 4 layers of 5×5 deformable convolutions → Output H×W×4 (predicting [l,t,r,b], i.e., the distance from the current point to the four sides of the target box) NWD loss calculation method: Target bounding box modeling: converting the actual bounding box into a target bounding box. With prediction box Each is modeled as a two-dimensional Gaussian distribution:
[0051]
[0052] Wasserstein distance calculation:
[0053] Normalization and loss:
[0054]
[0055] Total loss:
[0056] in Balanced classification, centrality, and location regression loss III. Model Training and Testing; 3.1 Training environment and parameter settings; 1. Training Framework: The HSN algorithm model is built and trained based on the MMDetection framework; 2. Hardware configuration: Intel Xeon E5-2680 processor, GeForce RTX3090 graphics card, CUDA 11.3 driver; 3. Training parameters: A total of 2000 training epochs were set, with 32 data points trained in each batch; the optimizer used was the SGD optimizer with momentum, and the momentum value was set to 0.9; the learning rate strategy was as follows: for the first 10000 iterations, the learning rate was linearly increased from 0 to 0.0002, and this learning rate was maintained until the 500th epoch. Then, the learning rate was decreased using cosine annealing until the end of training.
[0057] 3.2 Model Detection Process 1. Input: Preprocess the newly acquired infrared images of photovoltaic modules according to the standards of the dataset (size normalized to 640×640, pixel value normalized to 0~1); 2. Feature processing: The preprocessed image is fed into the infrared image fast backbone network of the HSN algorithm to extract features. After channel reconstruction and neck network fusion reconstruction, it is fed into the small target no-anchor-box detection head. 3. Defect Output: The detection head optimizes the prediction results through the NWD loss function and outputs the location (rectangular coordinates), defect category and confidence level of the defect in the image. The confidence level threshold is set to 0.5. Prediction results below the threshold are considered invalid. Finally, the defect location and category identification results of the photovoltaic module are obtained.
[0058] IV. Experimental Verification and Comparison; 1.1 Dataset; To verify the effectiveness of the algorithm proposed in this patent, an infrared imager mounted on a drone was used to photograph photovoltaic power plants, and a dataset of infrared defects in photovoltaic panels was collected using web crawling methods. After collection, images with poor image quality were removed, and the defects were labeled. The dataset was then divided into training, validation, and test sets in a 7:2:1 ratio. The dataset contains a total of 6096 images, with three types of defects: diode faults, hotspots, and terminal faults.
[0059] 1.2 Implementation details; The algorithms in this patent are all implemented based on the MMDetection framework, using a workstation with an Intel Xeon E5-2680 processor, a GeForce RTX 3090 graphics card, and CUDA 11.3 driver. When training the HSN algorithm model, a total of 2000 training epochs were set, with 32 data points trained per batch. The optimizer used was the momentum SGD optimizer, with a momentum value set to 0.9. During training, the learning rate was linearly increased from 0 to 0.0002 in the first 10000 iterations and maintained until the 500th epoch, after which cosine annealing was used to decrease the learning rate until the overall training was completed.
[0060] The purpose of this patent is to propose a precise photovoltaic panel defect diagnosis algorithm. Therefore, this patent mainly uses the mean Average Precision (mAP) to measure the detection accuracy. The mAP calculation method is as follows: for each category, calculate the precision (P) and recall (R) of the predicted boxes that are higher than the intersection-union (IU) threshold according to the following formula, where P is the number of correctly identified boxes. Then, for each category, calculate its average precision (AP), which is the area under the precision-recall curve. Finally, take the average of the AP values of all categories to obtain the mAP value. In this patent, the IU threshold is successively adjusted from 0.5 to 0.95 in steps of 0.05, and the final mAP is the average mAP at each threshold.
[0061]
[0062] The table shows the mAP values of each comparison algorithm and the HSN algorithm of this patent in the test set, as well as the AP values of each type of defect. The HSN algorithm achieved an mAP value of 0.79, the highest among the comparison algorithms, demonstrating the effectiveness of the method proposed in this patent and its ability to improve the accuracy of photovoltaic panel defect identification.
[0063] Comparison of test results
[0064] Among the compared algorithms, the two-stage target recognition algorithm performed better on average. In particular, the Faster-R-CNN and RegNet algorithms effectively identified terminal block faults and also showed excellent accuracy in identifying hot spot faults in small targets. The two-stage algorithm has certain advantages if detection efficiency is not a concern. This patented algorithm has advantages over the two-stage algorithm in identifying small target defects, such as hot spot faults and diode faults. Compared to the superior Faster-R-CNN algorithm, the HSN algorithm has a 6.4% higher accuracy in identifying hot spot faults and an 8.9% higher accuracy in identifying diode faults. This demonstrates the effectiveness of techniques such as Vision Transformer and normalized Wasserstein distance used in this patented algorithm for small target defects. Among the single-stage algorithms, the anchor-frame based algorithm performed worse than the anchor-free algorithm. This is because photovoltaic panel defects are distributed widely and vary in scale, while the anchor-frame method has poor robustness. This also proves the correctness of using an anchor-free detection head in this patent. Compared to the superior YOLOv8 algorithm among the compared single-stage algorithms, the HSN algorithm demonstrates superior accuracy in defect identification across all categories, with a 13.3% higher accuracy rate for diode fault identification, a 5.6% higher accuracy rate for hot spot fault identification, and a 2.5% higher accuracy rate for terminal fault identification. This proves the effectiveness of the patented Vision Transformer and its channel reconstruction and spatial reconstruction technologies, and also demonstrates the potential of the single-stage target detection algorithm structure in photovoltaic defect detection.
[0065] V. Alternative Solutions; 1. Option 1: Replaceable backbone network structure – a lightweight backbone network based on MobileViT; 1.1 Design Background; The main solution, which uses a backbone network based on EfficientFormer fusion of VisionTransformer (ViT) and Convolutional Neural Network (CNN), achieves a balance between detection accuracy and computational efficiency. However, it still suffers from high computational load when deployed on low-computing-power edge devices in photovoltaic power plants (such as UAV-borne terminals and embedded inspection instruments). This alternative solution uses a MobileViT structure to construct a fast backbone network for infrared images. This structure combines the lightweight, depthwise separable convolutions of MobileNet with micro-ViT blocks, further reducing model parameters and computational load while retaining global context modeling and local feature extraction capabilities, making it more suitable for lightweight on-site detection needs.
[0066] 1.2 Specific implementation method; Input layer processing: Consistent with the main scheme, the three-channel pseudo-color infrared image is first input into the infrared image mapping layer, and then adaptively adjusted to a single-channel tensor through a 1×1 convolution kernel. The pseudo-color noise is removed and the effective temperature amplitude information is retained. The output single-channel tensor is used as the input of the MobileViT backbone network.
[0067] MobileViT backbone network architecture construction; Shallow local feature extraction: The initial feature extraction module is constructed using the depthwise separable convolutional layer of MobileNet. The single-channel infrared image is downsampled and local features are extracted. The image is then processed by a 3×3 depthwise separable convolution with a stride of 2, batch normalization, and SiLU activation function to generate a low-dimensional, high-resolution local feature map, which is suitable for capturing the detailed features of photovoltaic module defects.
[0068] Mid-level global feature fusion: Introducing MobileViT's miniature Transformer block, the local feature map extracted by convolution is divided into blocks according to spatial dimensions, converted into sequence features, and then input into a multi-head self-attention layer to capture the long-distance dependency between photovoltaic module defects and the background; at the same time, the global features output by Transformer are fused with the convolutional local features through residual connections, taking into account both global and local information, and solving the problem that features of small target defects (such as small hot spots and diode faults) are easily lost.
[0069] Multi-scale feature output: By stacking three MobileViT modules of different scales, feature maps with downsampling of 8×, 16×, and 32× are output respectively, which are consistent with the output scale of the backbone network of the main scheme. It can be directly connected to the channel reconstruction neck network of the main scheme without any structural adjustment to the neck network, thus achieving seamless connection with the original HSN algorithm.
[0070] Network optimization and adaptation: Based on the temperature feature distribution characteristics of infrared images, the attention layer of MobileViT is improved by reducing the number of heads in the multi-head self-attention from 8 to 4, while reducing the dimension of the feature sequence, further reducing the amount of computation without affecting global feature capture.
[0071] 1.3 Technical Effects and Adapted Scenarios; The MobileViT backbone network constructed by this alternative has approximately 40% fewer parameters than the main solution and the EfficientFormer backbone network, while improving inference speed by approximately 50% and reducing detection accuracy by only 1%-2% on the same dataset. At the same time, it retains the core advantage of the main solution's "local + global" feature extraction, and can still maintain a high level of accuracy in identifying small target defects.
[0072] Suitable scenarios: real-time onboard detection by drones at photovoltaic power plants, edge detection by embedded inspection terminals, and portable detection equipment with limited computing power.
[0073] 2. Alternative Solution Two: Lightweight Edge Deployment Solution – Lightweight Transformation Combining Model Quantization and Channel Pruning; 2.1 Background of the scheme design; The main solution's HSN algorithm can achieve high-precision infrared defect identification of photovoltaic modules on the server side, but the model size is large (approximately 280MB) and the computational load is high, making it impossible to achieve real-time inference directly on low-computing-power edge devices (such as Jetson Nano, RK3588, and microcontroller-based inspection terminals) at the photovoltaic power plant site. This alternative solution uses a combined model compression technique of INT8 quantization and channel pruning to lightweightly modify the HSN algorithm trained by the main solution. While slightly sacrificing detection accuracy, it significantly reduces the model size and computational load, enabling efficient deployment on edge devices. Furthermore, the modified model completely retains the detection logic and core features of the original algorithm.
[0074] 2.2 Specific implementation method; This solution involves post-processing the trained HSN algorithm model without retraining it. The specific steps are as follows: Channel pruning: Trimming redundant feature channels; Feature importance calculation: The channel weights of all convolutional layers in the backbone and neck networks of the main scheme HSN algorithm are analyzed. The feature contribution of each channel is measured by calculating the L1 norm. The larger the L1 norm, the higher the contribution of the channel to the extraction of defect features, and vice versa.
[0075] Adaptive pruning threshold setting: For photovoltaic module defect detection scenarios, the pruning threshold is set to 0.2 (which can be adjusted according to accuracy requirements). All redundant channels with L1 norm below the threshold are pruned, while the core feature channels of the backbone network multi-scale feature output layer and the neck network channel separation processing block (CSB) are retained. The pruning ratio of the backbone network is 30%, and the pruning ratio of the neck network is 20%, to avoid feature loss due to excessive pruning.
[0076] Fine-tuning after pruning: The pruned model is fine-tuned for a small number of rounds (50 rounds) using the validation set of the original dataset. The learning rate is set to 0.00001 to compensate for the accuracy loss caused by pruning and ensure the model's ability to capture defective features.
[0077] INT8 quantization: reduces parameter storage precision; Quantization-aware calibration: 1000 infrared images from the original dataset are selected as the calibration set. Post-training quantization is performed on the pruned model. All 32-bit floating-point (FP32) weights, biases and activation values in the model are converted to 8-bit integers (INT8). At the same time, quantization parameters are calculated through the calibration set to ensure the linear mapping relationship of the quantized features.
[0078] Quantization accuracy compensation: Quantization protection is performed on the regression and classification branches of the detection head to retain their FP32 accuracy and avoid the regression error of small target defect bounding boxes and the decrease in category classification accuracy caused by quantization. Only the backbone network and neck network are quantized with INT8 to balance the degree of lightweighting and detection accuracy.
[0079] Model deployment adaptation: The quantized and pruned lightweight model is converted to ONNX format, and then converted to TensorRT, TFLite or RKNN format according to the hardware architecture of the edge device to adapt to the inference framework of different embedded devices. At the same time, the inference process of the model is optimized, redundant feature processing steps are removed, and real-time inference speed is improved.
[0080] 2.3 Technical Effects and Adapted Scenarios; The modified lightweight HSN model of this alternative solution reduces the model size from 280MB to about 35MB, and the number of parameters is reduced by about 85%. The inference speed on the Jetson Nano edge device is increased from 5 frames / second to 25 frames / second, meeting the requirements for real-time detection. The mAP value on the same test set is slightly reduced from 0.79 to 0.76-0.77, and the recognition accuracy of small target defects (hot spots, diode faults) is reduced by only 2%-3%, which fully meets the defect detection requirements of photovoltaic power plants.
[0081] Suitable scenarios: large-scale on-site inspection of photovoltaic power plants, real-time defect identification on drones, offline detection of embedded inspection terminals, and portable equipment detection in environments without network access.
[0082] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0083] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for infrared defect detection of photovoltaic modules using hybrid networks, characterized in that, The specific steps include the following: Infrared images of photovoltaic modules were acquired, and a defect dataset containing diode faults, hot spot faults, and terminal faults was constructed. A hybrid network model is constructed, comprising, in sequence, an infrared image mapping layer, an infrared image fast backbone network, a channel reconstruction neck network, and a small target no-anchor-frame detection head. The infrared image mapping layer adaptively maps the input three-channel pseudo-color infrared image into a single-channel temperature feature map using a learnable 1×1 convolutional kernel. The infrared image fast backbone network, based on a structure integrating Vision Transformer and convolutional neural networks, extracts and outputs multi-scale feature maps from the single-channel temperature feature maps. The channel reconstruction neck network upsamples, stitches, reduces dimensionality, and downsamples the multi-scale feature maps, extracts channel information through a channel separation processing block (CSB), and outputs the reconstructed multi-scale feature maps. The small target no-anchor-frame detection head uses a detection head without preset anchor frames and employs Normalized Wasserstein distance (NWD) as part of the loss function to predict the location and category of defects based on the reconstructed multi-scale feature maps. The hybrid network model is trained using the defect dataset. The infrared image of the photovoltaic module to be inspected is input into the trained hybrid network model, which outputs the category and location of the defect.
2. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 1, characterized in that, The infrared image fast backbone network is designed based on the EfficientFormer architecture, and the specific processing steps are as follows: The single-channel temperature feature map is processed through multiple stages in sequence. In the early stage, a convolutional neural network is used to extract local features, and in the later stage, a Vision Transformer is used to capture global contextual dependencies. The multi-scale feature map includes outputs at least three different scales, which are fed into the channel reconstruction neck network for feature fusion processing.
3. The method for infrared defect detection of hybrid network photovoltaic modules according to claim 1, characterized in that, The specific steps for channel reconfiguration of the neck network are as follows: The smaller-scale feature maps output by the backbone network are upsampled to the same size as the largest-scale feature maps and then stitched together along the channel dimension. Channel dimensionality reduction is performed on the stitched feature map using a 1×1 convolutional layer; Multi-scale downsampling is performed on the dimensionality-reduced feature map to generate multiple feature maps of different scales; The feature map at each scale is fed into the Channel Separation Processing Block (CSB) for processing. The CSB processing flow includes: splitting the input feature map into two parts in the channel dimension, one part being processed by the Spatial Reconstruction Unit (SRU) and the Channel Reconstruction Unit (CRU) in sequence, and then concatenating it with the other part before outputting.
4. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 3, characterized in that, The spatial reconstruction unit's processing includes sequentially performing Depthwise convolution, batch normalization, and Pointwise convolution; the channel reconstruction unit's processing includes sequentially performing global average pooling, dimensionality-reduced 1×1 convolution, ReLU activation, dimensionality-upgraded 1×1 convolution, and Sigmoid activation.
5. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 1, characterized in that, The small target no-anchor-box detection head is a NASFCOS detection head obtained based on neural network search technology. It includes a regression branch, a classification branch, and a centrality branch. The regression branch and the classification branch adopt deformable convolution. The normalized Wasserstein distance (NWD) is used instead of the traditional intersection-over-union ratio (IoU) to measure the error between the predicted target bounding box and the real target bounding box, and it is used to construct the total loss function.
6. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 5, characterized in that, The total loss function of the small target no-anchor-box detection head is a weighted sum of classification loss, centrality loss, and NWD loss: ; NWD loss calculation method: Target bounding box modeling: converting the actual bounding box into a target bounding box. With prediction box Each is modeled as a two-dimensional Gaussian distribution: ; ; Wasserstein distance calculation: ; Normalization and loss: ; ; in, Use Focal Loss or binary cross-entropy loss. Using binary cross-entropy loss, , , The weighting coefficients are used; the NWD is obtained by calculating the normalized Wasserstein distance after modeling the target box as a two-dimensional Gaussian distribution. The maximum value of the Wasserstein distance between all pairs of ground truth bounding boxes and predicted bounding boxes in the training dataset is used as a normalization constant to ensure... The value of is in the range of [0,1].
7. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 1, characterized in that, When constructing the infrared defect dataset for the photovoltaic modules, images were acquired using a drone equipped with an infrared thermal imager. The temperature range of the infrared thermal imager was -20℃ to 150℃, the drone's shooting altitude was 5 to 8 meters, and the shooting angle was vertically downward. The defect categories included diode faults, hot spot faults, and terminal faults, and their judgment criteria were as follows: Diode failure: A blocky temperature rise area appears that matches the installation location of the bypass diode, with a temperature rise difference ≥10℃; Hot spot fault: Spot-like or small-area temperature rise areas appear, with a temperature rise difference ≥8℃; Terminal block failure: The overall temperature rises at the terminal block location, and the overall temperature of the component is ≥5℃ higher than that of a normal component.
8. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 1, characterized in that, The infrared image fast backbone network of the hybrid network model is replaced with a lightweight backbone network based on MobileViT. The MobileViT backbone network uses depthwise separable convolution to extract local features and fuses global context information through micro Transformer blocks. The multi-scale feature map includes outputs at least three different scales, which are fed into the channel reconstruction neck network for feature fusion processing.
9. The method for infrared defect detection of a hybrid network photovoltaic module according to claim 1, characterized in that, The trained hybrid network model is subjected to lightweight modification, which includes: Channel pruning: Calculate the L1 norm of each channel in the convolutional layer of the network, prune redundant channels below a preset threshold, and fine-tune the pruned model; INT8 quantization: For the pruned model, post-training quantization is performed using a calibration set to convert the weights and activation values in the model from floating-point to integer.
10. A hybrid network photovoltaic module infrared defect detection system, characterized in that, Defect detection of photovoltaic modules using infrared images is performed using any one of claims 1 to 9.