Photovoltaic cell multi-defect detection method based on image intelligent identification
By building an intelligent photovoltaic cell defect recognition model, the problems of poor feature extraction robustness and high false detection rate in photovoltaic cell detection are solved, and multi-scale perception and high-precision defect detection are achieved to meet industrial-grade detection needs.
Patent Information
- Application Number
- CN202510888389.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-26
AI Technical Summary
Existing photovoltaic cell defect detection technology has the disadvantages of strong dependence on feature extraction and poor robustness, serious missed detection of small defects, insufficient multi-scale perception, blurred features of some defects, and a high false detection rate due to excessive image interference information. It is difficult to stably identify multiple types of defects under complex lighting and interference backgrounds.
A photovoltaic cell defect intelligent recognition model is constructed, including a backbone network, RE module, PM module, SF module, Thead module and several IFM modules. Through multi-scale feature extraction, attention mechanism and context modeling, multi-task collaborative feature fusion and interference suppression are achieved to improve detection accuracy.
It achieves high-precision, low-false-detection identification of surface defects of multiple types of photovoltaic cells under complex lighting and interference backgrounds, meets industrial-grade detection needs, and improves detection efficiency and accuracy.
Smart Images

Figure CN120707550A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the application field of cross-integration of solar photovoltaic detection technology and artificial intelligence image processing, and specifically to a photovoltaic cell multi-defect detection method based on image intelligent recognition. Background Art
[0002] Photovoltaic cells, the core energy conversion unit in solar systems, have a quality that directly determines the power generation efficiency and service life of PV modules. However, in actual production, due to complex manufacturing processes and variable environmental conditions, PV cells are prone to various surface defects, such as cracks, broken grids, cold solder joints, dark spots, corner debris, and fingerprint stains. If these defects are not detected promptly and accurately, they can easily trigger hot spot effects, leading to cell failure or even damage to the entire module. Traditional PV cell defect detection methods primarily include manual visual inspection, IV curve analysis based on electrical characteristics, and early machine vision technology. Manual inspection relies on operator experience, resulting in low efficiency, strong subjectivity, and a high rate of missed detections. While methods based on electrical parameters can detect some hidden defects, they cannot accurately locate and classify them. Traditional image processing methods rely primarily on manual feature design such as edge detection and threshold segmentation. These methods are extremely sensitive to lighting conditions, image quality, and background noise, making them unsuitable for detecting diverse defect types and complex scenes. With the rapid development of intelligent image recognition technology, deep learning, particularly end-to-end image recognition models represented by convolutional neural networks (CNNs), has become a key research direction for PV cell surface defect detection. Deep learning methods can automatically extract multi-level image features, and have stronger characterization capabilities and higher detection accuracy than traditional methods. However, current photovoltaic cell defect detection technology based on deep learning still faces the following key technical challenges in practical applications. Currently, the main problems faced by photovoltaic cell multi-defect detection are as follows:
[0003] 1. Feature extraction is highly dependent and has poor robustness:
[0004] Traditional image processing-based detection methods rely heavily on manually designed features and lack robustness to complex backgrounds in images, such as grid lines, stains, reflected light, and low-contrast areas, and cannot be generalized to all defect types.
[0005] 2. Serious omissions of small defects and insufficient multi-scale perception:
[0006] Photovoltaic defects vary greatly in size, and small cracks and broken wires are particularly prone to being overlooked. Existing models, lacking an effective multi-scale feature fusion mechanism, struggle to accurately identify both large and small objects.
[0007] 3. Some defect characteristics are vague and samples are scarce:
[0008] Defects such as black spots, short circuits, and fingerprints often appear as low-contrast or weak texture structures, with unclear feature boundaries. Furthermore, due to the limited annotation costs and the limited number of training samples, the model's learning capabilities are limited.
[0009] 4. The image has a lot of interference information and the false detection rate is high:
[0010] Surface images of photovoltaic cells often contain non-defective noise such as reflection patterns, irregular silicon wafer edges, and pseudo defects, which can easily be misjudged by the model as real defects, resulting in a high false detection and missed detection rate.
[0011] Therefore, there is an urgent need for an image recognition method with multi-scale perception capabilities, strong robustness and intelligent feature mining capabilities, which can stably identify various types and forms of photovoltaic cell surface defects under complex lighting and interference backgrounds, improve detection efficiency and accuracy, and meet the actual needs of industrial intelligent detection. Summary of the Invention
[0012] In response to the needs of the prior art, the present invention provides a photovoltaic cell multi-defect detection method based on image intelligent recognition, the purpose of which is to meet the needs of high-precision and high-efficiency industrial-grade photovoltaic module defect detection and diagnosis.
[0013] A photovoltaic cell multi-defect detection method based on image intelligent recognition includes the following steps:
[0014] Step 1: Construct a photovoltaic cell defect intelligent recognition model; the photovoltaic cell defect intelligent recognition model includes a backbone network, RE module, PM module, SF module, Thead module and several IFM modules;
[0015] Step 2: Train and verify the photovoltaic cell defect intelligent recognition model using the training set and validation set to obtain the optimal photovoltaic cell defect intelligent recognition model;
[0016] Step 3: Send the feature map to be detected into the optimal photovoltaic cell defect intelligent recognition model;
[0017] Step 4: Output the photovoltaic cell defect location and corresponding defect type.
[0018] Further: Step 3 is specifically as follows:
[0019] Step 3.1: The feature map to be detected is processed by the backbone network to generate three sets of multi-scale feature maps, which are shallow features, middle features, and deep features respectively;
[0020] Step 3.2: After the three groups of multi-scale feature maps are compressed and transformed by the RE module, they are subjected to group convolution operations to obtain shallow group convolution features, middle group convolution features, and deep group convolution features;
[0021] Step 3.3: The deep features are processed by the PM module to obtain feature P7 and used as the main branch input of the first IFM module. The deep group convolution features are processed by the CGR module and used as the auxiliary branch input of the first IFM module. The output of the first IFM module is feature P6.
[0022] Step 3.4: Feature P6 is upsampled and used as the main branch input of the second IFM module. The middle-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the second IFM module. The output of the second IFM module is used as feature P5.
[0023] Step 3.5: Feature P5 is upsampled and used as the main branch input of the third IFM module. The low-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the third IFM module. The output of the third IFM module is used as feature P4.
[0024] Step 3.6: The shallow, middle, and deep convolutional features are fed into the SF module, which is used to optimize the relationship between features and capture contextual information.
[0025] Step 3.7: Fuse the output of the SF module with feature P4 to obtain feature P3;
[0026] Step 3.8: Send features P3 to P7 to the Thead module, which includes a feature extractor and two task alignment predictors;
[0027] Step 3.8.1: Obtain task interaction features from features P3 to P7 through the feature extractor;
[0028] Step 3.8.2: Both task alignment predictors process the task interaction features and classify and locate them respectively.
[0029] Step 3.9: Fuse the classification information with the positioning information to obtain the classification scores and regression bounding boxes of multiple types of defects in the photovoltaic cell image.
[0030] Further: In step 3.2, the RE module consists of a 1×1 convolution layer and a nonlinear activation function ReLU; the group convolution operation includes a two-dimensional convolution layer, a group normalization layer, and a ReLU activation function operation in sequence.
[0031] Further: In step 3.3, the PM module includes a channel attention module and a spatial attention module.
[0032] Further, the processing process of the IFM module in step 3 is:
[0033] Step S1: The auxiliary branch input is sequentially passed through the CGR module and upsampled to obtain high-level input features ( );
[0034] Step S2: Perform maximum pooling and average pooling operations and spatial attention modules on the high-level input features to generate masks;
[0035] Step S3: The mask branch is multiplied with the main branch input after 1-mask operation to obtain the background attention feature ( ); The other branch is directly multiplied with the main branch input to obtain the foreground attention feature ( );
[0036] Step S4: Send the background attention features and foreground attention features into the CS module to obtain FP interference and FN interference respectively;
[0037] Step S4: Add the main branch input, FN interference and high-level input features and perform GR operation to obtain the preliminary corrected foreground features, where the learnable parameters Multiplied by FN interference to adjust the impact of FN interference in feature correction; GR operation is group normalization and Activated operation combination;
[0038] Step S5: After subtracting the initially corrected foreground features from the FP interference, the final corrected foreground features are obtained by GR operation, which is the output of the IFM module. Multiplied by FP interference to adjust the suppression strength of FP interference.
[0039] Further: In step 3.8.1, task interaction features The calculation formula is as follows:
[0040]
[0041] in, are features P3 to P7, represents k consecutive convolutional layers (k is 6 in this embodiment), represents different convolutional layers and Belong to 1 to , yes Activation function;
[0042] Further: Step 3.8.2 is specifically as follows:
[0043] Step 3.8.2.1: Calculate the classification task-specific feature weight w1 and the localization task-specific feature weight w2 based on the task interaction features.
[0044] Step 3.8.2.2: Multiply the feature weights by the task interaction features to obtain the specific task features for the classification task or localization task;
[0045] Step 3.8.2.3: The specific task features are sequentially processed through 1×1 convolutional layers, function and 2×2 convolutional layer to obtain classification results and positioning results respectively.
[0046] Further: The calculation formulas for feature weight w1 and feature weight w2 are both: ,in, yes activation function, yes activation function, It is the task interaction feature.
[0047] The beneficial effects of the present invention are as follows: the present invention constructs an intelligent recognition model for photovoltaic cell defects through a backbone network, an RE module, a PM module, an SF module, a Thead module and several IFM modules, so that the feature map to be detected is subjected to feature compression and transformation in sequence through the RE module on the basis of extracting multi-scale semantic information through the backbone network, and the channel and spatial attention mechanism are introduced by the PM module to improve the target area response, and multi-layer feature fusion and interference suppression are realized in the IFM module through foreground / background separation and context modeling; at the same time, the SF module further optimizes the cross-layer feature relationship, and the Thead module completes the difference modeling and coordination of classification and positioning information through feature extraction and task alignment predictor, finally realizing the feature extraction and fusion process from shallow to deep and multi-task collaboration, thereby obtaining accurate photovoltaic cell detection defect locations and corresponding defect types. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flow chart of the present invention;
[0049] Figure 2 This is the structural diagram of the intelligent identification model for photovoltaic cell defects;
[0050] Figure 3 It is the structural block diagram of the IFM module;
[0051] Figure 4 This is the structural block diagram of the SF module. DETAILED DESCRIPTION
[0052] The present invention will be described in detail below with reference to the accompanying drawings. The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention. The directional terms such as left, center, right, top, and bottom in the embodiments of the present invention are merely relative concepts or are based on the normal use state of the product, and should not be considered as restrictive.
[0053] A photovoltaic cell multi-defect detection method based on image intelligent recognition includes the following steps:
[0054] Step 1: Construct a photovoltaic cell defect intelligent recognition model; the photovoltaic cell defect intelligent recognition model includes a backbone network, RE module, PM module, SF module, Thead module and several IFM modules;
[0055] Step 2: Train and verify the photovoltaic cell defect intelligent recognition model using the training set and validation set to obtain the optimal photovoltaic cell defect intelligent recognition model;
[0056] Step 2.1: This method first collects representative and diverse high-resolution image samples from actual photovoltaic cell production inspection lines. The collected images cover 12 common defect types, including cracks, fingerprints, black cores, thick lines, star-shaped cracks, corner defects, debris, scratches, horizontal misalignment, vertical misalignment, printing errors, and short circuits, ensuring coverage of typical and complex defect forms. A total of 4,500 image samples were collected and organized. Each image was manually annotated by professionals, including the precise location of the defect in the image (framed by a bounding box) and the corresponding defect category label. After annotation, all images were uniformly normalized and formatted to form a high-quality supervised learning dataset with standardized structure and accurate labels, providing a solid data foundation for subsequent model training.
[0057] Step 2.2: After labeling, the dataset was divided into training, validation, and test sets in a ratio of 7:2:1, with 3,150 images, 900 images, and 450 images, respectively, for model training, parameter adjustment, and performance evaluation. Class-balanced stratified sampling was used during the partitioning process to ensure that all defect types were represented in each subset, preventing data skew from causing unstable model performance.
[0058] Step 2.3: Perform various data augmentation operations on the training data to improve the robustness and generalization ability of the model; the enhancement methods are as follows:
[0059] Multi-scale scaling to enhance the model’s ability to perceive defects of different sizes;
[0060] Random horizontal or vertical flipping to improve the model's spatial transformation invariance;
[0061] Light disturbance, simulating actual ambient light changes;
[0062] Random cropping and filling to enhance the ability to identify local defects;
[0063] Occlusion simulation operations such as Cutout and Mosaic improve the model's ability to cope with image interference;
[0064] Step 2.4: Train and verify the photovoltaic cell defect intelligent recognition model to obtain the optimal photovoltaic cell defect intelligent recognition model;
[0065] Step 3: Send the feature map to be detected into the optimal photovoltaic cell defect intelligent recognition model; specifically including the following steps:
[0066] Step 3.1: The feature map to be detected is processed by the backbone network to generate three sets of multi-scale feature maps, namely shallow features, mid-level features, and deep features. The backbone network uses the standard ResNet-50 structure. The spatial downsampling ratios of shallow features, mid-level features, and deep features are 8 times, 16 times, and 32 times compared to the input image, respectively, ensuring the multi-scale feature expression capability from low-level to high-level layers. Shallow features retain more detailed information, while deep features have stronger semantic expression capabilities.
[0067] Step 3.2: After the three groups of multi-scale feature maps are compressed and transformed by the RE module, they are subjected to group convolution operations to obtain shallow group convolution features, middle group convolution features, and deep group convolution features;
[0068] The RE module consists of a 1×1 convolutional layer (Conv) and a nonlinear activation function ReLU. The 1×1 convolution operation is used to adjust the channel dimension of the feature map, thereby reducing redundant calculations and improving computational efficiency. The ReLU activation function enhances the nonlinear expression capability of the features, helping to capture more complex feature patterns. The final output feature map retains the key semantic information in the original features while having a more compact expression form, which is convenient for subsequent module processing.
[0069] The group convolution operation consists of the following three operations in sequence: 2D convolution layer (Conv2d): The number of input and output channels is 256, the convolution kernel size is 3×3, the stride is 2, and the padding is 1. This convolution layer not only extracts local spatial features but also downsamples the feature map size with a stride of 2, improving the computational efficiency of the model and expanding the receptive field.
[0070] Group Normalization (GroupNorm): Sets the total number of channels to 256 and divides them into 32 groups for normalization, with each group containing 8 channels. Compared with BatchNorm, GroupNorm still has a good normalization effect when small batches or single sample inputs are used, improving model training stability.
[0071] ReLU activation function: Apply ReLU nonlinear activation operation to the normalized feature map to enhance the model's expressiveness and nonlinear modeling capabilities, helping to extract more discriminative features;
[0072] Step 3.3: The deep features are processed by the PM module to obtain feature P7 and serve as the main branch input of the first IFM module. The deep group convolution features are processed by the CGR module and serve as the auxiliary branch input of the first IFM module. The output of the first IFM module is feature P6. The overall structure of the PM module consists of two core sub-modules: the channel attention module (CA_Block) and the spatial attention module (SA_Block). The processing process of the PM module is as follows: the input features first pass through CA_Block and SA_Block respectively; CA_Block: This module adaptively assigns weights to each channel by modeling the global dependency between channels, thereby enhancing the channel response with strong discrimination of the target category; SA_Block: This module focuses on the spatial dimension and generates a spatial attention map, which enables the network to highlight the spatial position of the target area while suppressing background interference.
[0073] Then the output feature maps of CA_Block and SA_Block are added element-wise (element-wise addition) to achieve attention fusion of channel dimension and spatial dimension;
[0074] The final output fused feature map not only retains deep semantic information, but also has good spatial positioning capabilities, effectively improving the focus on the target area, and is especially suitable for detection tasks of small targets or low-contrast targets;
[0075] Step 3.4: Feature P6 is upsampled and used as the main branch input of the second IFM module. The middle-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the second IFM module. The output of the second IFM module is used as feature P5.
[0076] Step 3.5: Feature P5 is upsampled and used as the main branch input of the third IFM module. The low-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the third IFM module. The output of the third IFM module is used as feature P4.
[0077] The processing of the IFM module in steps 3.3, 3.4, and 3.5 is as follows:
[0078] Step S1: The auxiliary branch input is sequentially passed through the CGR module and upsampled to obtain high-level input features ( );
[0079] Step S2: The high-level input features are subjected to maximum pooling and average pooling operations and a spatial attention module to generate a mask. The core of the maximum pooling and average pooling operations and the spatial attention module (SAM module) is to generate a spatial attention map by using the spatial information of the input high-level feature map through maximum pooling and average pooling operations.
[0080] Step S3: The mask branch is multiplied with the main branch input after 1-mask operation to obtain the background attention feature ( ); The other branch is directly multiplied with the main branch input to obtain the foreground attention feature ( ); The 1-mask operation is specifically to subtract the mask from a tensor with the same shape as the mask feature map and all values are 1; the foreground attention feature enables the model to process the target information more focused by enhancing the saliency of the target area, while the background attention feature ensures that the model can ignore irrelevant background noise by reducing background interference;
[0081] Step S4: Send the background attention features and foreground attention features to the CS module (multi-scale context information extraction module) to obtain FP interference and FN interference respectively;
[0082] Step S4: Add the main branch input, FN interference and high-level input features and perform GR operation to obtain the preliminary corrected foreground features, where the learnable parameters Multiplied by FN interference to adjust the impact of FN interference in feature correction; GR operation is group normalization and Activated operation combination;
[0083] Step S5: After subtracting the initially corrected foreground features from the FP interference, the final corrected foreground features are obtained by GR operation, which is the output of the IFM module. Multiplied by FP interference to adjust the suppression strength of FP interference; learnable parameters and Initialized to 1, this operation effectively reduces FN interference and FP interference, thereby improving the recognition accuracy of foreground targets;
[0084] The IFM module is designed to fuse two types of features from different sources and embed an attention modulation mechanism to model the semantic differences between the main and auxiliary branches, thereby improving fusion accuracy. The two inputs of the IFM module are: on the one hand, the high-level semantic features generated by the PM module (the main branch), which have strong object localization capabilities; on the other hand, the original or structured features at the current image level (the auxiliary branch), which preserve local details.
[0085] Step 3.6: The shallow, middle, and deep convolutional features are fed into the SF module, which is used to optimize the relationship between features and capture contextual information.
[0086] Step 3.7: Fuse the output of the SF module with feature P4 to obtain feature P3; achieve the integration and enhancement of multimodal information, thereby further improving the diversity and robustness of feature expression;
[0087] Step 3.8: Send features P3 to P7 to the Thead module, which includes a feature extractor and two task alignment predictors;
[0088] Step 3.8.1: Obtain task interaction features from features P3 to P7 through the feature extractor;
[0089] Task interaction features The calculation formula is as follows:
[0090]
[0091] in, are features P3 to P7, represents k consecutive convolutional layers (k is 6 in this embodiment), represents different convolutional layers and Belong to 1 to , yes Activation function;
[0092] Step 3.8.2: Both task alignment predictors process the task interaction features and classify and locate the information respectively. The specific steps include:
[0093] Step 3.8.2.1: Calculate the classification task-specific feature weight w1 and the localization task-specific feature weight w2 based on the task interaction features. The calculation formulas for feature weight w1 and feature weight w2 are both: ,in, yes activation function, yes activation function, is the task interaction feature;
[0094] Step 3.8.2.2: Multiply the feature weights by the task interaction features to obtain the specific task features for the classification task or localization task;
[0095] Step 3.8.2.3: The specific task features are sequentially processed through 1×1 convolutional layers, function and 2×2 convolution layer to obtain classification results (classification information) and positioning results (positioning information) respectively;
[0096] Step 3.9: Fuse the classification information with the positioning information to obtain the classification scores and regression bounding boxes of multiple types of defects in the photovoltaic cell image;
[0097] Step 4: Output the photovoltaic cell defect location and corresponding defect type.
[0098] In addition, in step 3.8.2.3, to optimize the classification results (classification information) and positioning results (positioning information), you can perform the following steps:
[0099] Step 3.8.2.3.1: Introduce the task alignment weight map and generate two task alignment maps through convolution operation 、 , the classification prediction and fine-tuning positioning prediction are adjusted respectively through these two task alignment maps; the formulas of the two task alignment maps are:
[0100]
[0101]
[0102] in, , , , represents the spatial position of the tensor, Represents channel, alignment graph and alignment graph From the characteristics Automatic learning in
[0103] Step 3.8.2.3.2: Introduce the task alignment metric and calculate the task alignment of each anchor point based on the preliminary prediction results;
[0104] The task alignment metric is defined as: ,in, is the classification score, is the intersection-over-union (IoU) value. To further improve alignment accuracy, we introduced task alignment learning. Task alignment learning dynamically selects high-quality anchors and optimizes sample allocation by designing a new anchor alignment metric. The alignment of each anchor is measured by a combination of classification score and IoU. This metric helps the network dynamically select high-quality anchors, thereby improving training efficiency and prediction accuracy. Task alignment learning selects anchors with high task alignment as positive samples and the remaining anchors as negative samples. This allocation strategy effectively addresses the problem of non-maximum suppression and ensures that training samples better meet the requirements of task alignment.
[0105] Step 3.8.2.3.3: Combine the prediction deviations of classification and positioning to construct a task alignment loss function, guide the model to focus on high-quality predictions, and alleviate the task conflict problem; in order to deal with the task alignment problem between classification tasks and positioning tasks; we introduce the task alignment metric , to dynamically adjust the learning process of classification and localization tasks, avoid feature conflicts between tasks, and thus improve the performance of target detection; the task alignment loss function consists of two main parts, classification loss and localization loss; its steps are:
[0106] Classification loss:
[0107]
[0108] in, represents the number of positive anchor points, represents the number of negative anchor points, It represents the weight parameter, Representation is a regularized task alignment metric that is dynamically calculated based on the alignment of the classification and localization tasks. is the classification score of the positive anchor point, is the classification loss, which represents the error between the classification score and the task alignment metric, is the classification score of the negative anchor, is the classification score of the negative sample;
[0109] Positioning loss:
[0110]
[0111] in, represents the number of positive anchor points, is the predicted object bounding box, is the true bounding box, It is a generalized intersection-over-union loss, used to calculate the error between the true bounding box and the predicted bounding box. During training, focal loss is used to alleviate the imbalance between positive and negative samples, especially by giving greater weight to samples that are difficult to classify. By weighting the regression loss, task alignment loss enables the model to focus on the regression of high-quality anchors and reduce the negative impact of low-quality anchors on training.
[0112] Total loss function: The ultimate goal of the task alignment loss function is to optimize the classification and localization tasks, ensuring that they can work together effectively and avoid conflicts; the final loss function is the weighted sum of the classification loss and the localization loss:
[0113]
[0114] in, and is a parameter that balances classification loss and regression loss, we set .
[0115] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A photovoltaic cell multi-defect detection method based on image intelligent recognition, characterized by: The following steps are involved: Step 1: Construct a photovoltaic cell defect intelligent recognition model; the photovoltaic cell defect intelligent recognition model includes a backbone network, RE module, PM module, SF module, Thead module and several IFM modules; Step 2: Train and verify the photovoltaic cell defect intelligent recognition model using the training set and validation set to obtain the optimal photovoltaic cell defect intelligent recognition model; Step 3: Send the feature map to be detected into the optimal photovoltaic cell defect intelligent recognition model; Step 4: Output the photovoltaic cell defect location and corresponding defect type.
2. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 1, characterized in that: Step 3 is as follows: Step 3.1: The feature map to be detected is processed by the backbone network to generate three sets of multi-scale feature maps, which are shallow features, middle features, and deep features respectively; Step 3.2: After the three groups of multi-scale feature maps are compressed and transformed by the RE module, they are subjected to group convolution operations to obtain shallow group convolution features, middle group convolution features, and deep group convolution features; Step 3.3: The deep features are processed by the PM module to obtain feature P7 and used as the main branch input of the first IFM module. The deep group convolution features are processed by the CGR module and used as the auxiliary branch input of the first IFM module. The output of the first IFM module is feature P6. Step 3.4: Feature P6 is upsampled and used as the main branch input of the second IFM module. The middle-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the second IFM module. The output of the second IFM module is used as feature P5. Step 3.5: Feature P5 is upsampled and used as the main branch input of the third IFM module. The low-level group convolution features are processed by the CGR module and used as the auxiliary branch input of the third IFM module. The output of the third IFM module is used as feature P4. Step 3.6: The shallow, middle, and deep convolutional features are fed into the SF module, which is used to optimize the relationship between features and capture contextual information. Step 3.7: Fuse the output of the SF module with feature P4 to obtain feature P3; Step 3.8: Send features P3 to P7 to the Thead module, which includes a feature extractor and two task alignment predictors; Step 3.8.1: Obtain task interaction features from features P3 to P7 through the feature extractor; Step 3.8.2: Both task alignment predictors process the task interaction features and classify and locate them respectively. Step 3.9: Fuse the classification information with the positioning information to obtain the classification scores and regression bounding boxes of multiple types of defects in the photovoltaic cell image.
3. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 2, characterized in that: In step 3.2, the RE module consists of a 1×1 convolutional layer and a nonlinear activation function ReLU; the group convolution operation includes a two-dimensional convolutional layer, a group normalization layer, and a ReLU activation function operation in sequence.
4. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 2, characterized in that: In step 3.3, the PM module includes a channel attention module and a spatial attention module.
5. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 2, characterized in that: The processing of the IFM module in step 3 is as follows: Step S1: The auxiliary branch input is sequentially passed through the CGR module and upsampled to obtain high-level input features ( ); Step S2: Perform maximum pooling and average pooling operations and spatial attention modules on the high-level input features to generate masks; Step S3: The mask branch is multiplied with the main branch input after 1-mask operation to obtain the background attention feature ( ); The other branch is directly multiplied with the main branch input to obtain the foreground attention feature ( ); Step S4: Send the background attention features and foreground attention features into the CS module to obtain FP interference and FN interference respectively; Step S4: Add the main branch input, FN interference and high-level input features and perform GR operation to obtain the preliminary corrected foreground features, where the learnable parameters Multiplied by FN interference to adjust the impact of FN interference in feature correction; GR operation is group normalization and Activated operation combination; Step S5: After subtracting the initially corrected foreground features from the FP interference, the final corrected foreground features are obtained by GR operation, which is the output of the IFM module. Multiplied by FP interference to adjust the suppression strength of FP interference.
6. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 2, characterized in that: In step 3.8.1, task interaction features The calculation formula is as follows: ; in, are features P3 to P7, represents k consecutive convolutional layers (k is 6 in this embodiment), represents different convolutional layers and Belong to 1 to , yes Activation function.
7. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 2, characterized in that: Step 3.8.2 is as follows: Step 3.8.2.1: Calculate the classification task-specific feature weight w1 and the localization task-specific feature weight w2 based on the task interaction features. Step 3.8.2.2: Multiply the feature weights by the task interaction features to obtain the specific task features for the classification task or localization task; Step 3.8.2.3: The specific task features are sequentially processed through 1×1 convolutional layers, function and 2×2 convolutional layer to obtain classification results and positioning results respectively.
8. The photovoltaic cell multi-defect detection method based on image intelligent recognition according to claim 7, characterized in that: The calculation formulas for feature weight w1 and feature weight w2 are both: ,in, yes activation function, yes activation function, It is the task interaction feature.