Photovoltaic module defect detection method for enhancing YOLOv11n performance by combining Elmix

The dual-splitting feature interactive sampling module (Elmix) enhances YOLOv11n performance, solves the problem of difficulty in multi-scale feature capture in photovoltaic panel defect detection, and improves the small object detection effect and the recognition accuracy of complex defects.

CN120411022APending Publication Date: 2025-08-01HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500697.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing methods are difficult to effectively capture multi-scale features, resulting in poor detection of small and medium-sized targets for photovoltaic panel defect detection and limited detection accuracy for complex defects.

Method used

The dual-splitting feature interactive sampling module (Elmix) is adopted to extract local details and contextual features of photovoltaic modules through dual-branch architecture and core interaction, combining convolutional operation, pixel recombination, cross-branch attention guidance and pooling processing, and optimize feature representation through dynamic fusion and attention mechanism.

Benefits of technology

It improves the accuracy and recall of defect detection of photovoltaic modules, especially the ability to identify small defects, reduces the interference of background noise, and achieves higher detection accuracy and recall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411022A_ABST
    Figure CN120411022A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic module defect detection method for enhancing YOLOv11n performance in combination with Elmix, and the method comprises the steps: building an EL photovoltaic module picture data set, and inputting the data set into Elmix; according to the Elmix first branch, efficient extraction and dynamic enhancement of local detail features of an image are realized through a three-stage processing flow of convolution operation, pixel recombination and cross-branch attention guidance. In the second branch, feature maps are extracted through maximum pooling and average pooling and then spliced, the feature maps are extracted in an interpolation down-sampling mode, and weight value self-adaptive balance pooling and interpolation are designed; the first branch generates guiding attention to the second branch, and the second branch generates reverse attention to the first branch; the feature maps of the two branches are subjected to attention feature weighting, resolution alignment and splicing; and finally, passing through a YOLOv11n model. Compared with the prior art, the method has the advantages that the Elmix is combined to enhance the YOLOv11n model, so that the performance of detecting the defects of the photovoltaic module by the YOLOv11n is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a photovoltaic module defect detection method that combines a dual-branch feature interaction sampling module (Elmix) to enhance the performance of YOLOv11n. Background Art

[0002] With the transformation of the global energy structure towards cleaner and lower-carbon forms, solar energy, as an inexhaustible and renewable energy source, is gradually becoming an important part of the energy field. Photovoltaic power generation, as the main form of solar energy utilization, the quality of its core component - the photovoltaic panel directly affects the power generation efficiency and service life. However, during the production, transportation, installation, and use of photovoltaic panels, various defects will inevitably occur, such as black spots, scratches, poor sintering, and over-etching. These defects will not only significantly reduce the power generation efficiency of the panel but may also pose safety hazards. Therefore, achieving rapid and accurate detection of photovoltaic panel defects is of great significance for ensuring the safe and stable operation of photovoltaic power stations and improving power generation efficiency.

[0003] In the task of photovoltaic panel defect detection, the existing methods have the following problems:

[0004] 1) Traditional convolutional neural networks are difficult to effectively capture multi-scale features, resulting in poor detection effects for small targets.

[0005] 2) There are various types of photovoltaic panel defects, and the existing methods have limited detection accuracy for complex defects. Summary of the Invention

[0006] Object of the Invention: Aiming at the problems pointed out in the background art, the present invention designs a photovoltaic module defect detection method that combines a dual-branch feature interaction sampling module (Elmix) to enhance the performance of YOLOv11n. Through the dual-branch feature interaction sampling module (Elmix), feature extraction is performed based on deep learning, and through the dual-branch structure and core interaction, tiny defects are prevented from being submerged by background noise.

[0007] Technical Solution: The present invention discloses a photovoltaic module defect detection method that combines Elmix to enhance the performance of YOLOv11n, including the following steps:

[0008] Step 1: Establish an EL photovoltaic module image dataset and preprocess the image dataset;

[0009] Step 2: Construct the Elmix module, and input the dataset into the Elmix module for processing; the Elmix module includes two branches. The first branch: through a three-level processing process of convolution operation + pixel recombination + cross-branch attention guidance, efficiently extract and dynamically enhance the local detail features of the dataset images. The second branch: extract the feature maps by using max pooling and average pooling and then splice them, extract the feature maps by using interpolation downsampling, and design a weight value to adaptively balance the pooling and interpolation processes. The first branch generates guiding attention for the second branch, and the second branch generates reverse attention for the first branch. The feature maps of the two branches are weighted by attention features, and after resolution alignment, they are spliced.

[0010] Step 3: Input the feature maps processed by the Elmix module into the YOLOv11n model for photovoltaic module defect detection.

[0011] Furthermore, the Elmix module inputs the original image where: B is the batch size, C is the number of channels, H×W is the spatial size, the first branch is the detail enhancement branch, with a high-resolution path of 2H×2W; the second branch is the context encoding branch, with a multiple downsampling path of H / 2×W / 2.

[0012] Furthermore, the process of the first branch is specifically as follows:

[0013] (1) Use the initial convolution to extract the basic features and expand the channels by 4 times:

[0014]

[0015] where Conv2D is the convolution operation, X is the input feature map, W c1 is the weight matrix of the initial convolution kernel, is the intermediate feature map of the first branch;

[0016] (2) Use PixelShuffle upsampling to convert the channel dimension into spatial resolution:

[0017]

[0018] where F1 (2) is the feature map after upsampling F1 (1) by PixelShuffle, and its dimension is (B, C, 2H, 2W);

[0019] (3) Use the detail enhancement convolution to refine the local features at high resolution:

[0020]

[0021] where Wc2 is the weight matrix of the detail-enhanced convolution kernel, and Y1 represents the output feature map of the first branch.

[0022] Further, the process of the second branch is specifically as follows:

[0023] (1) Utilize the dual pooling path to fuse the multi-scale features of the two poolings, namely local saliency and global smoothness:

[0024]

[0025]

[0026] F pool = Conv2D(Concat(P max , P avg ); W p )

[0027]

[0028] where MaxPool2D represents performing max pooling with a 2×2 window and a stride of 2, P max represents the feature after max pooling, AvgPool2D represents performing average pooling with a 2×2 window and a stride of 2, P avg represents the feature after average pooling, Concat represents the concatenation operation; F pool represents the fused feature map; W p represents the weight of the Conv2D convolution;

[0029] (2) Interpolation path:

[0030] F interp = Conv2D(BilinearInterp(X, scale_factor = 0.5); W i )

[0031]

[0032] where F interp represents the feature after bilinear interpolation, BilinearInterp represents the bilinear interpolation operation; W i represents the weight of the 3×3 convolution kernel; scale_factor = 0.5 means the image size becomes half of the original;

[0033] (4) Dynamic fusion, adaptively balance the contributions of the pooling and interpolation paths:

[0034] γ = σ(w γ ), Y2 = γ·F pool + (1 - γ)·Finterp

[0035]

[0036] Among them, the weight parameter w γ is a learnable parameter, mapped to [0, 1] through the Sigmoid function (σ), and Y2 is the output feature map of the second branch.

[0037] Furthermore, the first branch and the second branch perform bidirectional attention interaction. The first branch generates guiding attention for the second branch, and the second branch generates reverse attention for the first branch, specifically as follows:

[0038] (1) First branch → Second branch: The first branch performs feature downsampling, aligns the high-resolution features to the resolution of the second branch, then generates an attention feature map to identify the regions in the second branch that need to be enhanced, and uses the high-resolution details to guide the selection of low-resolution semantic features:

[0039]

[0040] Among them, is the feature map after downsampling by 4 times average pooling, W g1 , W g2 represent the convolutional weight parameters for generating the attention map, σ represents the Sigmoid function, compresses the attention values to [0, 1], A 1→2 represents the guiding attention map, identifying the regions in the second branch that need to be enhanced, ⊙ represents element-wise multiplication, represents the features of the second branch after attention weighting, Y1 represents the output feature map of the first branch, and Y2 is the output feature map of the second branch;

[0041] (2) Second branch → First branch: Enhance the detailed features through context information:

[0042]

[0043] Among them, [[ID=...]] (the ellipsis indicates that there are some tags not fully translated in the original, but following the rules, they should be kept as they are) represents the feature map after upsampling 4 times by bilinear interpolation, W r1 , W r2 represent the convolutional weights for generating the reverse attention map, A 2→1 represents the reverse attention map, suppressing irrelevant background noise, represents the features of the first branch after context enhancement.

[0044] Furthermore, the feature maps of the two branches are weighted by attention features and concatenated after resolution alignment, specifically as follows:

[0045] (1) Resolution alignment

[0046]

[0047] Among them, represents the feature after the final alignment of the first branch, represents the feature of the first branch after context enhancement;

[0048] (2) Feature concatenation

[0049]

[0050] Among them, represents the feature of the second branch after attention weighting, Y combined represents the result of double-branch feature concatenation;

[0051] (3) Final output

[0052]

[0053] Among them, W f represents the weight of the fusion convolution, and Output represents the final output feature.

[0054] Furthermore, during the process of processing the dataset using the Elmix module, a joint loss function is also established, as follows:

[0055] (1) Total loss function:

[0056]

[0057] Among them, is the main output loss, is the branch consistency loss, is the attention regularization loss, α ∈ [0, 1] is the main loss weight, λ is the regularization coefficient, and the three parts of the loss are optimized collaboratively to balance the output accuracy and branch consistency;

[0058] (2) Main output loss:

[0059]

[0060] Among them, O is the module output, T is the true label, C out represents the number of channels of the model output. The purpose of the main output loss is to supervise the consistency between the final output of the module and the true label, and directly optimize the pixel-level accuracy of the final output; H' and W' respectively represent the height and width of the output feature map after a certain process. After the operations of the first branch and the second branch of the Elmix module, the size of the output feature map is represented by H′ and W′;

[0061] (3) Branch consistency loss:

[0062]

[0063] Among them, represents the first-branch downsampled feature, represents the second-branch upsampled feature; feature alignment is performed on the two branches, the first-branch feature is downsampled, and the second-branch feature is upsampled. The branch consistency loss aims to constrain the consistency of the dual-branch features, promote cross-resolution feature alignment, and moreover, force the two branches to learn similar feature representations at different resolutions, as well as prevent the first branch from overly focusing on noisy details or the second branch from being overly smoothed;

[0064] (4) Attention regularization loss

[0065]

[0066] Positive attention regularization:

[0067] Negative attention regularization:

[0068] The goal of the attention regularization loss is to prevent the attention map from degrading and maintain the activity of the attention mechanism. It can constrain the average activation value of the attention map, avoid being overly sparse or saturated, and ensure that the attention mechanism can dynamically adjust the attention area;

[0069] (5) Gradient backpropagation

[0070] The gradient of the total loss is the weighted sum of each component:

[0071]

[0072] Among them, Θ is the set of model parameters, including all trainable parameters of the model, and α and λ are hyperparameters that respectively control the weights of the main loss and the regularization loss;

[0073] Key gradient components:

[0074] Dynamic fusion parameter w γ :

[0075] Attention weight W g :

[0076] Among them, * represents the convolution operation, represents the gradient of the loss function with respect to the attention map A 1→2 , reflecting the impact of the attention map on the loss, is the intermediate feature, is the indicator function that only retains the gradients of the activated neurons, represents the gradient propagation of the convolution operation, is the input feature map.

[0077] Beneficial effects:

[0078] (1) Dual-branch collaborative architecture

[0079] Detail enhancement branch: Through PixelShuffle non-interpolation upsampling, the high-frequency information advantage of the fine cracks on the surface of the battery panel is retained, thereby improving the detection rate of pixel-level small defects. Context encoding branch: By fusing dual pooling (MaxPool + AvgPool) and dynamic interpolation, the semantic features of large-scale hidden cracks or hot spots are identified. Compared with traditional methods (such as U-Net) that rely on a single encoding-decoding path, the dual branches clearly separate details and semantics, avoiding small defects being submerged by background noise.

[0080] (2) Bidirectional dynamic attention mechanism

[0081] Branch 1 → Branch 2: The high-resolution features generate a spatial weight map to guide the low-resolution branch to focus on the crack area. Branch 2 → Branch 1: The low-resolution semantics generate a reverse mask to suppress false detections in irrelevant areas such as the battery panel border. Compared with traditional attention (such as CBAM) that only works within a single branch, the bidirectional mechanism achieves cross-resolution closed-loop optimization.

[0082] (3) Dynamic feature fusion strategy

[0083] Learnable weight γ: Adaptively adjusts the contribution ratio of pooling and interpolation. In the detection of battery panels, γ can automatically bias towards the interpolation path (retaining fine cracks) or the pooling path (capturing large-size defects). Description of the drawings

[0084] Figure 1 is the framework diagram of the photovoltaic module defect detection method of the present invention;

[0085] Figure 2 is the framework diagram of the dual-branch feature interaction sampling module (Elmix) of the present invention. Detailed implementation manners

[0086] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0087] The present invention discloses a photovoltaic module defect detection method that combines Elmix to enhance the performance of YOLOv11n, specifically including the following steps:

[0088] Step 1: Establish an EL photovoltaic module picture dataset and preprocess the picture dataset.

[0089] Step 2: Construct the Elmix module and input the dataset into the Elmix module for processing; the Elmix module includes two branches. The first branch (hereinafter referred to as Branch 1): Through a three-level processing process of convolution operation + pixel recombination + cross-branch attention guidance, efficient extraction and dynamic enhancement of local detail features of the dataset images are achieved; The second branch (hereinafter referred to as Branch 2): After using max pooling and average pooling to extract feature maps and then splicing them, feature maps are extracted using interpolation downsampling, and a weight value is designed to adaptively balance the pooling and interpolation processes; The first branch generates guiding attention for the second branch, and the second branch generates reverse attention for the first branch. The feature maps of the two branches are weighted by attention features, and after resolution alignment, they are spliced.

[0090] Step 2.1: Module Input and Overall Structure

[0091] Input the original image: Where: B is the batch size, C is the number of channels, and H×W is the spatial size. Dual-branch structure: Branch 1 (detail enhancement branch): High-resolution path (2H×2W). Branch 2 (context encoding branch): Multiple downsampling paths (H / 2×W / 2). Core interaction: Bidirectional attention mechanism + dynamic feature fusion.

[0092] Step 2.2: Step-by-step Analysis of the Whole Process

[0093] Step 2.2.1 Branch 1 (detail enhancement):

[0094] (1) Initial convolution: Extract basic features and expand the channels by 4 times.

[0095]

[0096] Among them, Conv2D is the convolution operation, X is the input feature map, W c1 is the weight matrix of the initial convolution kernel, is the intermediate feature map of the first branch.

[0097] (2) PixelShuffle upsampling: Convert the channel dimension to spatial resolution

[0098]

[0099] Among them, F1 (2) is the feature map after upsampling F1 (1) through PixelShuffle, and its dimension is (B, C, 2H, 2W).

[0100] (3) Detail enhancement convolution: Refine local features at high resolution

[0101]

[0102] Among them, W c2 is the weight matrix of the detail-enhanced convolutional kernel, and the dimension of Y1 is (B, C, 2H, 2W).

[0103] Step 2.2.2 Branch 2 (context encoding branch):

[0104] Dual pooling path: fuse the multi-scale features of two types of pooling, namely local saliency and global smoothness, that is, perform max pooling with a 2×2 window and a stride of 2, and the output resolution is reduced to H / 2×W / 2, perform average pooling with a 2×2 window and a stride of 2, and the output resolution is reduced to H / 2×W / 2. Concatenate P max and P avg in the channel dimension to obtain 2C channels. 1×1 convolution (Wp): compress the number of channels from 2C back to C to generate the fused feature map F pool :

[0105]

[0106] F pool = Conv2D(Concat(P max , P avg ); W p )

[0107]

[0108] Among them, the dimension of W p is (C_out, C_in, K, K), C_out is the number of output channels (i.e., the number of convolutional kernels), C_in is the number of input channels (consistent with C of X), and K is the size of the convolutional kernel 1×1.

[0109] (1) Interpolation path: first downsample the input to H / 2×W / 2 through bilinear interpolation, and then use a 3×3 convolution

[0110] (W i ) to further extract features, and the output is F interp :

[0111] F interp = Conv2D(BilinearInterp(X, scale_factor = 0.5); W i )

[0112]

[0113] W iIts dimension is (C_out, C_in, K, K), where C_out is the number of output channels (i.e., the number of convolutional kernels), C_in is the number of input channels (consistent with C of X), K is the size of the convolutional kernel 3×3, and the dimension of Finterp is (B, C, H / 2, W / 2), where B is the batch size, H / 2×W / 2 is the spatial size, and C is the number of channels.

[0114] (2) Dynamic fusion: adaptively balance the contributions of the pooling and interpolation paths

[0115] γ = σ(w γ ), Y2 = γ·F pool + (1 - γ)·F interp

[0116]

[0117] Among them, the weight parameter w γ is a learnable parameter, mapped to [0, 1] through the Sigmoid function (σ). If γ → 1: the model depends on the pooling path (suitable for large-sized defects). If γ → 0: the model depends on the interpolation path (suitable for retaining details).

[0118] Step 2.2.3 Bidirectional attention interaction

[0119] (1) Branch 1 → Branch 2 (guiding attention): use high-resolution details to guide the selection of low-resolution semantic features.

[0120]

[0121] Branch 1 performs feature downsampling to align the high-resolution features to the resolution of Branch 2. Then, an attention feature map is generated to identify the regions in Branch 2 that need to be strengthened (such as the context positions corresponding to the details).

[0122] Among them, is the feature map downsampled by 4-fold average pooling, W g1 , W g2 represents the convolutional weight parameter for generating the attention map, σ represents the Sigmoid function, compressing the attention values to [0, 1], A 1→2 represents the guiding attention map, identifying the regions in the second branch that need to be strengthened, ⊙ represents element-wise multiplication, represents the features of the second branch after attention weighting, Y1 represents the output feature map of the first branch, and Y2 is the output feature map of the second branch.

[0123] (2) Branch 2 → Branch 1 (reverse attention): enhance the detail features through context information.

[0124]

[0125] Among them, represents the feature map after upsampling 4 times by bilinear interpolation, W r1 , W r2 represents the convolutional weight for generating the reverse attention map, A 2→1 represents the reverse attention map, suppressing irrelevant background noise, represents the first-branch feature after context enhancement.

[0126] Step 2.2.4 Multi-scale Feature Fusion

[0127] (1) Resolution Alignment

[0128]

[0129] (2) Feature Concatenation

[0130]

[0131] (3) Final Output

[0132]

[0133] Among them, Y combined represents the result of double-branch feature concatenation, W f represents the weight of the fusion convolution, and Output represents the final output feature.

[0134] Step 2.2.5 Establish Joint Loss Function

[0135] (1) Total Loss Function:

[0136]

[0137] Among them, is the main output loss (directly optimizing the final output accuracy), is the branch consistency loss (constraining the alignment of double-branch features), is the attention regularization loss (preventing the degradation of the attention map), β ∈ [0, 1] is the main loss weight, λ is the regularization coefficient, and the three parts of the loss are jointly optimized to balance the output accuracy and branch consistency.

[0138] (2) Main Output Loss:

[0139]

[0140] Among them, O is the module output, T is the true label, and the purpose of the main output loss is to supervise the consistency between the final output of the module and the true label, and directly optimize the pixel-level accuracy of the final output.

[0141] (3) Branch Consistency Loss:

[0142]

[0143] Feature alignment is performed on two branches. The features of branch 1 are downsampled, and the features of branch 2 are upsampled. The branch consistency loss aims to constrain the consistency of the dual-branch features, promote cross-resolution feature alignment, and force the two branches to learn similar feature representations at different resolutions, as well as prevent branch 1 from overly focusing on noisy details or branch 2 from being overly smoothed.

[0144] (4) Attention regularization loss

[0145]

[0146] Forward attention regularization:

[0147] Reverse attention regularization:

[0148] The goal of the attention regularization loss is to prevent the attention map from degrading (all 0 or all 1) and maintain the activity of the attention mechanism. It can constrain the average activation value of the attention map, avoid being overly sparse or saturated, and ensure that the attention mechanism can dynamically adjust the attention area.

[0149] (5) Gradient backpropagation

[0150] The gradient of the total loss is the weighted sum of each component:

[0151]

[0152] Θ: All trainable parameters of the model (including convolutional weights, dynamic fusion parameter Wγ, attention weights Wg, Wr, etc.). α, λ: Hyperparameters that control the weights of the main loss and the regularization loss, respectively.

[0153] Key gradient components:

[0154] Dynamic fusion parameter w γ , and the weight wγ is updated through the chain rule to adaptively adjust the contributions of the pooling and interpolation paths:

[0155] Attention weight W g :

[0156] where * represents the convolution operation, represents the gradient of the loss function with respect to the attention map A 1→2 , reflecting the impact of the attention map on the loss, is the intermediate feature, is the indicator function that only retains the gradients of the activated neurons, Represents the gradient propagation of the convolution operation, is the input feature map.

[0157] In summary, the joint loss function drives the final output to approximate the true label through the main loss, the branch consistency loss constrains the multi-scale feature alignment, and the regularization term maintains the effectiveness of the attention mechanism. The three work together to achieve precision priority: the main loss dominates the learning of key features; stability guarantee: branch consistency prevents the model from collapsing to a single branch; mechanism activity: regularization ensures the dynamic adjustment ability of attention, and it can be adjusted and adapted to different task requirements in real applications.

[0158] Step 3: Input the feature map processed by the Elmix module into the YOLOv11n model for photovoltaic module defect detection.

[0159] The above method is experimentally verified below. YOLOv11n, YOLOv11n-SE, YOLOv11n-CBAM, and the proposed YOLOv11n-Elmix of the present invention are respectively sampled to process the same photovoltaic module defect images, and the detection performance results are shown in Table 1.

[0160] Table 1 Comparison of detection performance of different detection method models

[0161]

[0162] Traditional YOLOv11n predicts targets through multi-scale detection heads, but high-frequency details are easily lost in the deep network, resulting in poor detection effects for small targets (such as micro-cracks in photovoltaic panels). Due to the enhancement of Elmix, the detail features of 2H×2W are retained through PixelShuffle and directly used for small target localization. Dynamically fuse pooling and interpolation features to enhance context semantics (such as distinguishing hot spots from reflections). Compared with the fixed downsampling of YOLOv11n, the hybrid downsampling strategy of Elmix improves the recall rate of tiny defects while maintaining the speed.

[0163] YOLOv11 is prone to false detections on images with complex backgrounds and relies on post-processing to filter out false positives. Elmix has a bidirectional attention mechanism. From branch 1 to branch 2, it can generate spatial weights using high-resolution details to suppress background interference regions. While from branch 2 to branch 1, the low-resolution semantic feedback reverse attention can filter out noises such as border shadows. In photovoltaic detection, the precision of YOLOv11n-Elmix detection has increased. Compared with other improvement schemes, such as unidirectional attention modules like YOLOv11n+SE / CBAM, YOLOv11n-Elmix achieves cross-resolution closed-loop optimization and is significantly better than these schemes. That is, the precision of the present invention has increased by 2.2%, the recall has increased by 2.5%, mAP@0.5 has increased by 1.6%, and mAP@0.5:0.95 has increased by 1%.

[0164] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and the purpose is to enable those familiar with this technology to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. All equivalent transformations or modifications made according to the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for detecting defects in photovoltaic modules that combines Elmix to enhance the performance of YOLOv11n, characterized in that, It includes the following steps: Step 1: Establish an EL photovoltaic module image dataset and preprocess the image dataset; Step 2: Construct an Elmix module and input the dataset into the Elmix module for processing; the Elmix module includes two branches. The first branch: through a three-level processing process of convolution operation + pixel recombination + cross-branch attention guidance, realize the efficient extraction and dynamic enhancement of local detail features of the dataset images; The second branch: After extracting feature maps using max pooling and average pooling and then splicing them, extract feature maps using the method of interpolation downsampling, and design a weight value to adaptively balance the pooling and interpolation processes; the first branch generates guiding attention for the second branch, and the second branch generates reverse attention for the first branch. The feature maps of the two branches are weighted by attention features, and after resolution alignment, they are spliced; Step 3: Input the feature maps processed by the Elmix module into the YOLOv11n model for photovoltaic module defect detection.

2. The method for detecting photovoltaic module defects by combining Elmix to enhance the performance of YOLOv11n according to claim 1, characterized in that, The Elmix module inputs the original image Where: B is the batch size, C is the number of channels, H×W is the spatial dimension, the first branch is the detail enhancement branch, with a high-resolution path of 2H×2W; the second branch is the context encoding branch, with a multiple downsampling path of H / 2×W / 2.

3. The photovoltaic module defect detection method for enhancing the performance of YOLOv11n by combining Elmix according to claim 2, wherein, The process of the first branch is specifically as follows: (1) Use initial convolution to extract basic features and expand the channels by 4 times: Among them, Conv2D is a convolution operation, X is the input feature map, and W c1 is the weight matrix of the initial convolution kernel, is the intermediate feature map of the first branch; (2) Use PixelShuffle upsampling to convert the channel dimension into spatial resolution: Among them, F1 (2) is the feature map after upsampling F1 (1) by PixelShuffle, and its dimension is (B, C, 2H, 2W); (3) Use detail enhancement convolution to refine local features at high resolution: Among them, W c2 is the weight matrix of the detail enhancement convolution kernel, and Y1 represents the output feature map of the first branch.

4. The photovoltaic module defect detection method for enhancing the performance of YOLOv11n by combining Elmix according to claim 2, characterized in that, The process of the second branch is specifically as follows: (1) Use a dual-pooling path to fuse the multi-scale features of the two poolings, that is, local saliency and global smoothness: F pool = Conv2D(Concat(P max , P avg )); W p ) Among them, MaxPool2D represents max pooling with a 2×2 window and a stride of 2, and P max represents the feature after max pooling, AvgPool2D represents average pooling with a 2×2 window and a stride of 2, and P avg represents the feature after average pooling, Concat represents the concatenation operation; F pool represents the fused feature map; W p represents the weight of the Conv2D convolution; (2) Interpolation path: F interp = Conv2D(BilinearInterp(X, scale_factor = 0.5); W i ) Among them, F interp represents the feature after bilinear interpolation, and BilinearInterp represents the bilinear interpolation operation; W i represents the weight of the 3×3 convolution kernel; scale_factor = 0.5 means that the image size becomes half of the original size; (3) Dynamic fusion, adaptively balance the contributions of the pooling and interpolation paths: γ = σ(w γ ), Y2 = γ·F pool +(1 - γ)·F interp Among them, the weight parameter w γ is a learnable parameter, mapped to [0, 1] through the Sigmoid function (σ), and Y2 is the output feature map of the second branch.

5. The photovoltaic module defect detection method for enhancing the performance of YOLOv11n by combining Elmix according to claim 3 or 4, characterized in that The first branch and the second branch perform bidirectional attention interaction. The first branch generates guiding attention for the second branch, and the second branch generates reverse attention for the first branch. Specifically as follows: (1) First branch → Second branch, the first branch performs feature downsampling, aligns the high-resolution features to the resolution of the second branch, then generates an attention feature map to identify the regions in the second branch that need to be strengthened, and uses high-resolution details to guide the selection of low-resolution semantic features: Among them, is the feature map after downsampling by a factor of 4 average pooling, W g1 , W g2 represents the convolution weight parameter for generating the attention map, σ represents the Sigmoid function, which compresses the attention value to [0, 1], A 1→2 represents the guiding attention map, identifying the regions to be strengthened in the second branch, ⊙ represents element-wise multiplication, represents the features of the second branch after attention weighting, Y1 represents the output feature map of the first branch, and Y2 is the output feature map of the second branch; (2) Second branch → First branch, enhance the detail features through context information: Among them, represents the feature map after upsampling 4 times by bilinear interpolation, W r1 , W r2 represents the convolutional weight for generating the reverse attention map, A 2→1 represents the reverse attention map, which suppresses irrelevant background noise, represents the first-branch feature after context enhancement.

6. The photovoltaic module defect detection method for enhancing the performance of YOLOv11n by combining Elmix according to claim 5, characterized in that, The feature maps of the two branches are weighted by attention features, and after resolution alignment, they are spliced. Specifically as follows: (1) Resolution alignment Among them, represents the feature after the final alignment of the first branch, represents the feature of the first branch after context enhancement; (2) Feature splicing Among them, represents the second-branch feature after attention weighting, Y combined represents the concatenation result of the double-branch features; (3) Final output Among them, W f represents the weight of the fusion convolution, and Output represents the final output feature.

7. The method for detecting photovoltaic module defects by combining Elmix to enhance the performance of YOLOv11n according to any one of claims 2 to 6, characterized in that, During the process of processing the dataset using the Elmix module, a joint loss function is also established. Specifically as follows: (1) Total loss function: Among them, is the main output loss, is the branch consistency loss, is the attention regularization loss, α ∈ [0, 1] is the main loss weight, λ is the regularization coefficient, and the three parts of the loss are jointly optimized to balance the output accuracy and branch consistency; (2) Main output loss: Among them, O is the module output, T is the ground truth label, and C out represents the number of channels of the model output. The main output loss aims to supervise the consistency between the final output of the module and the ground truth label, and directly optimize the pixel-level accuracy of the final output; H' and W' respectively represent the height and width of the output feature map after certain processing. After the operations of the first branch and the second branch of the Elmix module, the size of the output feature map is represented by H′ and W′; (3) Branch consistency loss: Among them, represents the first-branch downsampled feature, represents the second-branch upsampled feature; feature alignment is performed on the two branches, with the first-branch features downsampled and the second-branch features upsampled. The branch consistency loss aims to constrain the consistency of the dual-branch features, promote cross-resolution feature alignment, and force the two branches to learn similar feature representations at different resolutions, as well as prevent the first branch from overly focusing on noisy details or the second branch from being overly smoothed; (4) Attention regularization loss Positive attention regularization: Inverse attention regularization: The goal of the attention regularization loss is to prevent the attention map from degrading and maintain the activity of the attention mechanism. It can constrain the average activation value of the attention map, avoid being overly sparse or saturated, and ensure that the attention mechanism can dynamically adjust the attention area; (5) Gradient backpropagation The gradient of the total loss is the weighted sum of each component: Among them, Θ is the set of model parameters, including all trainable parameters of the model, and α, λ are hyperparameters that respectively control the weights of the main loss and the regularization loss; Key gradient components: Dynamic fusion parameter w γ : Attention weight W g : where * represents the convolution operation, represents the gradient of the loss function with respect to the attention map A 1→2 which reflects The impact of the attention map on the loss, is the intermediate feature, It is an indicator function that only retains the gradients of the activated neurons. It represents the gradient propagation of the convolutional operation. is the input feature map.