Incremental defect detection method for photovoltaic cells based on hierarchical distillation of salient features
Through the incremental learning method of significant feature graded distillation, the problem that the photovoltaic cell defect detection model cannot continue to learn when the defect category increases is solved, and the model maintains the detection ability of old categories while learning new categories, meeting the demand for quality inspection for rapid updates and iterations.
Patent Information
- Application Number
- CN202310629870.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-05-31
AI Technical Summary
The existing CNN-based photovoltaic cell defect detection model cannot be continuously learned when the defect category increases, resulting in the detection performance of the old category deterioration and requires restarting training, which takes a long time and is difficult to meet the requirements of quality inspection tasks for rapid update iteration.
The incremental defect detection method based on significance feature hierarchical distillation is adopted. The teacher model is the same as the student model, and the student model is initialized using the parameters of the teacher model, and the student model is performed with the significance feature hierarchical distillation, so that the student model can be trained incrementally so that it can maintain the detection ability of the old category while learning new categories.
In the dynamic and open photovoltaic cell quality inspection process, the model can be quickly iteratively updated, and continuously learn new categories of defects without forgetting old categories, meeting the quality inspection needs for rapid deployment and iteration of detection models.
Smart Images

Figure CN116630285B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photovoltaic cell defect detection, and specifically is a photovoltaic cell incremental defect detection method based on significant feature graded distillation. Background Art
[0002] Defect detection is an essential step to ensure the quality of photovoltaic cells. Compared with manual visual inspection, computer vision inspection methods based on convolutional neural networks (CNNs) have many advantages such as high accuracy, strong robustness, and fast detection speed. Therefore, they are widely used in photovoltaic cell defect detection tasks.
[0003] The CNN-based detection method needs to determine the types of defects that need to be detected according to the current quality inspection requirements, establish a data set, and train the model to achieve effective detection of the current defect category. However, as quality inspection requirements increase, the types of defects that need to be detected may gradually increase. The model does not have the ability to learn continuously. If only new categories of defect samples and labels are used to incrementally train the current model, its detection performance for old categories will drop significantly. In order to adapt to changes in detection requirements, it is usually necessary to integrate all labeled samples of old and new categories of defects and retrain the model. This training method has high time complexity. Every time the defect category increases, the training needs to be restarted, which is time-consuming and difficult to meet the quality inspection task's requirements for rapid update and iteration of the detection model and rapid deployment.
[0004] Therefore, in order to adapt to the dynamic changes of defect categories during the quality inspection process, this application proposes a method based on hierarchical distillation of significant features. When the number of defect categories that need to be detected increases, the model can be quickly iterated and updated, and continue to learn in the dynamic and open actual photovoltaic cell quality inspection process. Summary of the invention
[0005] In view of the deficiencies in the prior art, the technical problem that the present invention intends to solve is to provide a photovoltaic cell incremental defect detection method based on significant feature graded distillation.
[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0007] A photovoltaic cell incremental defect detection method based on significant feature hierarchical distillation, characterized in that the method comprises the following steps:
[0008] Step 1: Establish a dataset of defects in old categories of photovoltaic cells;
[0009] Step 2: Build a defect detection model, use the old category defect data set of photovoltaic cells to train the defect detection model, and use the trained defect detection model as the original defect detection model;
[0010] Step 3: Establish a new category defect dataset for photovoltaic cells;
[0011] Step 4: Use the original defect detection model as the teacher model. The student model has the same architecture as the teacher model. The output dimension of the student model is the sum of the number of new defect categories and the number of old defect categories. The student model is initialized using the parameters of the teacher model.
[0012] The teacher model and the student model are subjected to hierarchical distillation of significant features, and the process of hierarchical distillation of significant features is:
[0013] Extract Q groups of features of the teacher model and the student model, each group of features is the key layer at the same position of the feature extraction part and the feature fusion part of the teacher model and the student model, obtain the feature map of the key layer output corresponding to the teacher model and the student model, calculate the spatial attention mask and the channel attention mask on the feature map, and obtain the binary mask of the separated foreground and background features of the teacher model detection result;
[0014] The new category defect dataset of photovoltaic cells is input into the teacher model and the initialized student model at the same time, and the student model is incrementally trained based on knowledge distillation; the loss function of incremental training is:
[0015] L=L dis +L det (4)
[0016] Among them, L represents the total loss of the student model, L det represents the detection loss generated by the student model learning new categories, L dis represents the loss generated by the hierarchical distillation of the overall significant features, which is used to maintain the model's ability to detect defects of old categories, L dis It is expressed as:
[0017]
[0018] in, and The feature distillation loss and attention distillation loss for any set of features can be expressed as:
[0019]
[0020]
[0021] Among them, M is a binary mask for separating foreground and background features, which is obtained from the detection results of the teacher model; and and are the spatial attention masks and channel attention masks of the teacher model and the student model for the qth group of features, respectively. and After zero mean processing, it is expressed as and α and β are hyperparameters for balancing the distillation loss of foreground and background features, α + β = 1. The teacher model’s salient features about the old category are adaptively extracted through spatial attention and channel attention, while the foreground and background regions of the feature map are separated for hierarchical distillation to balance the overwhelming distillation loss caused by a large range of backgrounds.
[0022] Step 5: Use the trained student model as the final defect detection model for photovoltaic cell defect detection; when the number of defect categories that need to be detected increases, repeat steps 3 and 4 to retrain the student model.
[0023] Furthermore, in the fourth step, the process of calculating the spatial attention and channel attention masks of the feature map of the key layer extracted by the teacher model is as follows: the features of the teacher model C×H×W dimensions are mapped to tensors of H×W dimensions and C dimensions respectively, and the activation-based teacher model spatial attention mask is obtained according to equations (4) and (5): and channel attention mask
[0024]
[0025]
[0026] in is the feature map of the teacher model, H, W, C represent the height, width and number of channels of the feature set. By similar methods, the spatial attention mask of the student model can be obtained. and channel attention mask
[0027] Furthermore, in the fourth step, the binary mask M is obtained from the detection result of the teacher model. If there are old category defects in the training sample, the teacher model outputs the detection result of the training sample and generates a binary mask M based on the prediction box of the detection result:
[0028]
[0029] Where b represents the prediction box of the teacher model for the old category in the training sample. The pixels in the prediction box of the binary mask M are set to 1, and the other pixels are set to 0. At this time, the scale of M is the same as that of the training sample. When participating in the operation in formula (3), the upsampling operation is performed to change the scale of the binary mask M so that its height H and width W are consistent with the zero-mean processed feature f T and f S same.
[0030] Furthermore, in the fourth step, the student model learns the detection loss L generated by the new category detThe expression is:
[0031] L det =λ 1 L cls +λ 2 L obj +λ 3 L box (1)
[0032] Where, L cls , L obj and L box They are category cross entropy loss, confidence cross entropy loss and positioning cross entropy loss, λ 1 , 2 and λ 3 are all hyperparameters, λ 1 +λ 2 +λ 3 =1.
[0033] Furthermore, the original defect detection model is any target detection model, such as Faster R-CNN, YOLO series models, etc.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] 1. In the actual photovoltaic cell quality inspection process, the types of defects that need to be detected may continue to increase. The existing CNN-based model does not have the ability to continue learning. When the defect categories that need to be detected increase, all data from new and old categories must be integrated to retrain the model to adapt to the changes in categories. In order to avoid retraining, this application proposes a hierarchical distillation method of significant features, so that the model does not forget old categories while learning new categories, and continues to learn in the dynamic and open photovoltaic cell quality inspection tasks, meeting the quality inspection needs for rapid update and iteration of detection models and rapid deployment.
[0036] 2. This application proposes hierarchical distillation of significant features. By calculating spatial attention and channel attention on the feature maps of the key layers of the teacher model, the significant features of the old categories of the teacher model are focused on and distilled adaptively. At the same time, the detection results of the teacher model are used to separate the foreground and background features for hierarchical distillation to avoid redundant distillation losses caused by large-scale background features, thereby obtaining a better stability-plasticity balance. The method of the present invention guides the distillation process and separates the foreground and background hierarchical distillation by performing spatial attention and channel attention on the teacher model, and uses the teacher model and the student model to transfer the knowledge of the teacher model to the student model, thereby solving the problem of class incremental target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0038] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods, but the protection scope of the present application is not limited thereto.
[0039] The present invention is a photovoltaic cell incremental defect detection method based on significant feature graded distillation, comprising the following steps:
[0040] Step 1: Establish a dataset of defects in old categories of photovoltaic cells;
[0041] The defect categories required for current quality inspection are taken as old categories. Based on electroluminescent imaging technology, old category defect images of photovoltaic cells are obtained through industrial cameras. The defect areas in the images are marked and category labels are added to obtain the old category defect dataset of photovoltaic cells.
[0042] Step 2: Build a defect detection model and train it, and use the trained defect detection model as the original defect detection model; select a suitable target detection model as the defect detection model according to the quality inspection requirements, such as the Faster R-CNN model, the YOLO series model, etc.; randomly divide the photovoltaic cell old category defect data set obtained in the first step into a training set and a test set in a ratio of 8:2, use the Mosaic data enhancement method to expand the training set, and use the expanded training set to train the defect detection model; calculate the loss in the training process through the loss function of the following formula;
[0043] L det =λ 1 L cls + λ 2 L obj + λ 3 L box (1)
[0044] Where, L det is the total training loss (i.e., the subsequent detection loss), L cls , L obj and L box They are category cross entropy loss, confidence cross entropy loss and positioning cross entropy loss, λ 1 , 2 and λ 3 are all hyperparameters, λ 1 +λ 2 +λ 3 =1;
[0045] The trained defect detection model is tested using the test set, and the model parameters are adjusted through back propagation until the loss converges; the trained defect detection model is used for defect detection of photovoltaic cells. At this time, the trained defect detection model can detect old category defects and is recorded as the original defect detection model;
[0046] Step 3: Establish a new category defect dataset for photovoltaic cells;
[0047] Assume that the original defect detection model can detect three types of defects: hidden cracks, broken grids, and black spots; if the defect category to be detected changes, and the defect categories to be detected are hidden cracks, broken grids, black spots, and linear defects, then hidden cracks, broken grids, and black spots are used as base defect categories, and linear defects are used as new defect categories; based on electroluminescent imaging technology, the defect image of photovoltaic cells of the new defect category is obtained through an industrial camera, and the position of the defect area in the image is marked, and a category label is added;
[0048] Step 4: Use the original defect detection model as the teacher model. The student model has the same architecture as the teacher model. The output dimension of the student model is the sum of the number of new defect categories and the number of old defect categories. The student model is initialized using the parameters of the teacher model.
[0049] The new category defect dataset of photovoltaic cells is input into the teacher model and the student model for dual network training. During the incremental training process, the teacher model does not update the parameters, and the student model optimizes the model by back propagation according to the loss function, where the loss function can be expressed as:
[0050] L=L det +L dis (2)
[0051] Where L det L is the detection loss, which is used to learn new categories and enable the teacher model to acquire the ability to detect defects in new categories; dis L is introduced to represent the loss caused by the graded distillation of the overall significant features. dis By generating additional regularization terms, the teacher model is constrained to forget old knowledge.
[0052] The hierarchical distillation of the salient features guides the distillation process through the spatial and channel attention of the feature map of the key layer of the teacher model, focuses on the salient features of the old category for distillation, separates the foreground and background, introduces hyperparameters to grade the foreground and background feature distillation, and avoids the imbalance of foreground and background distillation losses.
[0053] Specifically, extract the features of Q groups of teacher models and student models. Each group of features is the key layer at the same position of the feature extraction part and feature fusion part of the teacher model and the student model. For example, using Faster R-CNN as the detection model, take the outputs of C3, C4, and C5 of the feature extraction part of the teacher model and the student model and the outputs of P3, P4, and P5 of the feature fusion part; using YOLOv5 as the detection model, take the outputs of the second C3 module, the third C3 module, and the SPPF module of the feature extraction part of the teacher model and the student model and the three outputs of the bottom-up branch of the feature fusion part. Calculate the feature distillation loss and attention distillation loss between each group of features respectively, and sum them up to get the overall salient feature hierarchical distillation loss L dis .
[0054] For any set of features of the teacher model and the student model and First, the feature distillation loss is calculated, It is expressed as:
[0055]
[0056] Among them, M is a binary mask for separating foreground and background features, and are the spatial attention mask and channel attention mask of the qth group of feature teacher models, respectively. and After zero mean processing, it is expressed as and α and β are hyperparameters for balancing the distillation loss of foreground and background features, and α + β = 1. C, H, and W are the number of channels, height, and width of the set of features, respectively. The superscript T represents the student, and the superscript S represents the teacher.
[0057] Furthermore, the features of the teacher model C×H×W dimensions are mapped into tensors of H×W dimensions and C dimensions respectively, and the spatial attention mask of the teacher model based on the activation of this group of features is obtained according to formulas (4) and (5): and channel attention mask
[0058]
[0059] in is the feature map of the teacher model, H, W, C represent the height, width and number of channels of the feature map. By similar methods, the spatial attention and channel attention masks of the student model can be obtained. and
[0060] Furthermore, the binary mask M is obtained from the detection results of the teacher model. If there are old category defects in the training sample, the teacher model outputs the detection results of the training sample and generates a binary mask M based on the prediction box of the detection result:
[0061]
[0062] Where b represents the prediction box of the teacher model for the old category in the training sample. The pixels in the prediction box of the binary mask M are set to 1, and the other pixels are set to 0. At this time, the scale of M is the same as that of the training sample. When participating in the operation in formula (3), the upsampling operation is performed to change the scale of the binary mask M so that its height H and width W are consistent with the zero-mean processed feature f T and f S same.
[0063] Furthermore, forcing the spatial and channel attention masks of the student model to mimic those of the teacher model, the attention distillation loss can be expressed as:
[0064]
[0065] Finally, the overall significant feature hierarchical distillation loss L dis It is the sum of the feature distillation loss and the attention loss between each set of features, expressed as:
[0066]
[0067] The hierarchical distillation of significant features in the present invention does not have entity modules such as convolution kernels. The calculated distillation loss is used to maintain the ability to recognize old categories and optimize the model together with the loss of learning new categories. As shown in formula (2), the model can maintain the ability to recognize old categories while learning new categories, avoiding the shortcomings of the prior art that only new category data is used to train the original model (a model that can recognize old categories), and the loss is calculated according to formula (1) to optimize the model so that it can obtain the ability to recognize new categories, but at the same time the ability to recognize old categories will be lost. The teacher model and the student model are completely independent, and the feature maps of the key layers of the teacher model and the student model are taken for distillation loss calculation. The method proposed in the present application enables the model to update the model only using a data set containing new category defects, thereby achieving the ability to detect new and old category defects and avoiding retraining.
[0068] The target detection model in the present invention includes a feature extraction part and a feature fusion part. The teacher model only participates in the training process of the student model in the fourth step. The trained student model can detect new and old class defects.
[0069] During the incremental training process, only the network parameters of the student model are optimized until the total loss function converges and the incremental training process ends; the trained student model can detect both old category defects and new category defects; the trained student model is tested using the old category defect dataset and the new category defect dataset of photovoltaic cells.
[0070] Step 5: Use the trained student model as the final defect detection model for photovoltaic cell defect detection; when the number of defect categories to be detected increases, repeat steps 3 and 4 to retrain the student model so that the student model has the ability to detect new types of defects.
[0071] In summary, faced with the problem of the continuous increase in the number of defect categories that need to be detected in the dynamic and open photovoltaic cell quality inspection task, this application proposes a significant feature graded distillation method (see formulas (3) and (7)) to calculate the distillation loss. By introducing additional regularization terms, while the model learns new categories, it constrains old knowledge from being covered, giving the model the ability to continuously learn, which is of great significance for realizing continuous quality monitoring of photovoltaic cells.
[0072] Example 1
[0073] In this embodiment 1, the old category defects are three categories: hidden cracks, broken grids and black spots, and the new category defects are linear defects. For the old category defect dataset of photovoltaic cells, there are 903 hidden crack sample images, 1480 broken grid sample images, and 1009 black spot sample images in the training set, and 2280 hidden crack sample images, 11995 broken grid sample images, and 3824 black spot sample images in the test set. There are 2001 linear defect sample images, of which 773 are used for training and 1228 are used for testing. The original size of the image is 1024×1024 pixels.
[0074] The YOLOv5s model is selected as the defect detection model. The defect detection model is trained and tested using the old category defect dataset of photovoltaic cells. The loss is calculated according to formula (1) to obtain the trained defect detection model, which is recorded as the original defect detection model, namely the teacher model.
[0075] The architecture of the student model is the same as that of the teacher model. The output dimension of the student model is the sum of the number of new defect categories and the number of old defect categories, that is, the output dimension of the student model in this embodiment is 4. The student model is initialized using the teacher model parameters and incrementally trained. The output feature maps of the second C3 module, the third C3 module, and the SPPF module of the feature extraction part of the teacher model and the student model are selected and recorded as and Select the three output feature maps of the bottom-up branch of the teacher model and the student model feature fusion part, denoted as and The loss is calculated by formula (8); the trained student model is used as the final defect detection model, and the four types of defect sample images are input into the final defect detection model to test the detection performance of the final defect detection model for all categories. At the same time, the method of the present invention is compared with the existing common methods. The comparison results are shown in Table 1.
[0076] Table 1
[0077]
[0078] Among them, the model used in method one is obtained by retraining the original defect detection model using the old and new categories of photovoltaic cell defect data to reconstruct the data set, and the model used in method two is obtained by incrementally training the original defect detection model using only the new category defect data set of photovoltaic cell defects. The test results show that the mean average precision (mAP) of the method of the present invention is 76.8%, and the incremental training time is 1.56 hours. Compared with method two, mAP is improved by 41.7%, and compared with method one, mAP has only a 3.8% performance gap, but the training time is shortened by 5.3 hours. On the basis of ensuring the detection performance as much as possible, the method of the present invention has the ability to detect new and old categories of defects, and the model has the ability to continuously learn, which can meet the requirements of rapid iteration and update of the model during quality inspection.
[0079] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A photovoltaic cell incremental defect detection method based on hierarchical distillation of salient features, It is characterized in that The method comprises the following steps: Step 1: Establish a dataset of defects in old categories of photovoltaic cells; Step 2: Build a defect detection model, use the old category defect data set of photovoltaic cells to train the defect detection model, and use the trained defect detection model as the original defect detection model; Step 3: Establish a new category defect dataset for photovoltaic cells; Step 4: Use the original defect detection model as the teacher model. The student model has the same architecture as the teacher model. The output dimension of the student model is the sum of the number of new defect categories and the number of old defect categories. The student model is initialized using the parameters of the teacher model. The teacher model and the student model are subjected to hierarchical distillation of significant features, and the process of hierarchical distillation of significant features is: Extract Q groups of features of the teacher model and the student model, each group of features is the key layer at the same position of the feature extraction part and the feature fusion part of the teacher model and the student model, obtain the feature map of the key layer output corresponding to the teacher model and the student model, calculate the spatial attention mask and the channel attention mask on the feature map, and obtain the binary mask of the separated foreground and background features of the teacher model detection result; For any set of features of the teacher model and the student model and Calculating feature distillation loss and attention distillation loss in, Among them, M is a binary mask for separating foreground and background features, and are the spatial attention mask and channel attention mask of the qth group of feature teacher models, respectively. and After zero mean processing, it is expressed as and α and β are hyperparameters for balancing the distillation loss of foreground and background features, α+β=1; C, H, and W are the number of channels, height, and width of the set of features, respectively; Overall significant feature fractional distillation loss L dis It is the sum of the feature distillation loss and the attention distillation loss between each set of features, expressed as: The new category defect dataset of photovoltaic cells is input into the teacher model and the initialized student model at the same time, and the student model is incrementally trained based on knowledge distillation. The loss function of incremental training is: L=L dis +L det (4) Among them, L represents the total loss of the student model, L det represents the detection loss generated by the student model learning new categories, L dis Represents the loss of overall significant features resulting from fractional distillation; Train the student model by minimizing the above loss; Step 5: Use the trained student model as the final defect detection model for photovoltaic cell defect detection; when the number of defect categories that need to be detected increases, repeat steps 3 and 4 to retrain the student model.
2. The photovoltaic cell incremental defect detection method based on significant feature hierarchical distillation according to claim 1, It is characterized in that The process of calculating the spatial attention and channel attention masks for the feature maps of the key layers extracted by the teacher model is as follows: for any set of extracted features and The features of the teacher model C×H×W dimensions They are mapped into tensors of H×W and C dimensions respectively, and the spatial attention mask of the activated teacher model based on this group of features is obtained according to equations (4) and (5): and channel attention mask Among them, H, W, and C represent the height, width, and number of channels of the feature map.
3. The photovoltaic cell incremental defect detection method based on significant feature hierarchical distillation according to claim 1, It is characterized in that The binary mask M is obtained from the detection result of the teacher model. If there are old category defects in the training sample, the teacher model outputs the detection result of the training sample and generates a binary mask M according to the prediction box of the detection result: Where b represents the prediction box of the teacher model for the old category in the training sample; the pixels of the binary mask M in the prediction box are set to 1, and the other pixels are set to 0; at this time, the scale of M is the same as that of the training sample. When participating in the operation in formula (3), the upsampling operation is performed to change the scale of the binary mask M so that its height H and width W are the same as the feature f after zero mean processing. T and f S same.
4. The photovoltaic cell incremental defect detection method based on significant feature hierarchical distillation according to claim 1, It is characterized in that The student model learns the detection loss L generated by the new category det The expression is: L det =λ 1 L cls +λ 2 L obj +λ 3 L box (1) Where, L cls , L obj and L box They are category loss, confidence loss and positioning loss, λ 1 , 2 and λ 3 are all hyperparameters, λ 1 +λ 2 +λ 3 =1.
5. The photovoltaic cell incremental defect detection method based on significant feature hierarchical distillation according to claim 1, It is characterized in that The original defect detection model is any target detection model, including one of the FasterR-CNN and YOLO series models.
Citation Information
Patent Citations
Solar cell surface defect detection method based on feature-guided channel distillation
CN114511532A
Multi-data-set X-ray security check image target detection method based on dual-mode network
CN116071704A