A PCB board defect detection method based on model compression

By designing a lightweight feature extraction module and using knowledge distillation method in the PCB defect detection model, the problem of excessive model parameters and calculation amount in the prior art is solved, and the compression of model parameters and calculation amount is achieved, while maintaining detection accuracy.

CN114897845BActive Publication Date: 2025-05-06NANJING UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210550075.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-06
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The existing PCB defect detection model based on deep learning has problems with excessive model parameters and calculation amount during the deployment process, resulting in high performance requirements for the deployed devices.

Method used

By designing a lightweight feature extraction module to replace the feature extraction module in the YOLOv5 network, and combining the knowledge distillation method, the teacher network is introduced to supervise the feature map of the student network model, thereby achieving compression of model parameters and calculation amount.

Benefits of technology

While maintaining the average detection accuracy (mAP) of the original PCB defect detection model, the compression of model parameters and calculation amount is achieved, greatly reducing the calculation amount and reducing the performance requirements for deployed devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897845B_ABST
    Figure CN114897845B_ABST
Patent Text Reader

Abstract

The present invention discloses a PCB board defect detection method based on model compression, comprising: preprocessing the original defect image, forming a training sample by regional segmentation; designing a lightweight feature extraction module, replacing the feature extraction module in the YOLOv5 network with the lightweight feature extraction module to perform network compression, and constructing a compressed network; before knowledge distillation is performed, the YOLOv5 network is trained by a PCB sample, and its parameters will not be updated during the knowledge distillation process; based on the YOLOv5 network, in each iterative training process of the compression network, the local and global knowledge distillation modules are used to calculate the distillation loss, and the parameters of the knowledge distillation network are updated by a back propagation algorithm; defect detection is performed by the compression network, and the output detection frame is screened to obtain the final defect prediction frame. The present invention maintains the average detection accuracy of the original network model while compressing the parameter amount and calculation amount of the original model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to model compression in industrial production defect detection based on deep learning, and in particular to a PCB board defect detection method based on model compression. Background Art

[0002] With the development of deep learning and computer vision technology, there are currently a variety of PCB defect algorithms based on deep learning methods. The patent with application number CN202111244986.3 and name of PCB surface defect detection method based on improved YOLOv5 algorithm discloses a PCB defect detection method based on improved YOLOv5 algorithm, which improves the performance of the original network model through adaptive feature fusion and target weighted fusion methods. The patent with application number CN202010417205.5 and name of PCB board defect detection algorithm based on convolutional neural network discloses a PCB defect detection method based on convolutional neural network, which improves the average detection accuracy (mAP) of the network by adding an attention module; however, these methods often have the disadvantages of introducing additional model parameters, increasing detection time and increasing the amount of calculation, so they have higher performance requirements for the related equipment deployed by the algorithm. In order to solve the deployment problem of industrial PCB defect detection models, it is necessary to achieve the compression of network model parameters and calculation amount (FLOPs) under the premise of ensuring the original average detection accuracy (mAP). Summary of the invention

[0003] The purpose of the present invention is to provide a PCB board defect detection method based on model compression, which can achieve compression of model parameters and calculation amount (FLOPs) while maintaining the average detection accuracy (mAP) of the original PCB defect detection model, thereby reducing the amount of calculation.

[0004] The technical solution to achieve the purpose of the present invention is: a PCB board defect detection method based on model compression, comprising the steps of:

[0005] Step S1: pre-processing the defective image of the original PCB and forming a training sample by region segmentation;

[0006] Step S2: Design a lightweight feature extraction module, replace it with the feature extraction module in the YOLOv5 network to perform network compression, and build a compressed network;

[0007] Step S3: Before knowledge distillation is performed, the YOLOv5 network is trained using the PCB samples generated in step S1. During the knowledge distillation process, the parameters of the network will not be updated.

[0008] Step S4: Based on the YOLOv5 network trained in S3, during each iterative training process of the compression network constructed in step S2, the local and global knowledge distillation modules are used to calculate the distillation loss, and the parameters of the compression network are updated through the back propagation algorithm;

[0009] Step S5: PCB board defect detection is performed using the compression network trained in step S4, and the output detection frame is screened to obtain the final defect prediction frame.

[0010] Furthermore, the step 1 specifically includes: randomly cropping the original PCB defect data set to form a 600×600 image and updating the relative position of the corresponding annotations; dividing the enhanced file into a training validation set and a test set in a ratio of 8:2, and the sample ratio of the training set to the validation set in the training validation set is 3:2.

[0011] Furthermore, the input end of the YOLOv5 network trained in step S3 processes the training data as follows: extracting features of the input image through the C3 feature extraction module, performing multi-scale fusion through FPN and PAN, and generating a priori boxes of different numbers and sizes through the detection head.

[0012] Furthermore, the process of multi-scale fusion processing of FPN includes: using upsampling to perform multi-scale fusion on features of different scales from top to bottom; the process of multi-scale fusion processing of PAN includes: using downsampling to perform multi-scale fusion on features of different scales from bottom to top.

[0013] Furthermore, the step 2 specifically includes: replacing the convolution layer of the C3 module BottleNeck in the YOLOv5 network with the Ghost convolution model, and replacing the C3 module in the YOLOv5 network with the compressed C3 module for feature extraction.

[0014] Furthermore, the step S4 specifically includes:

[0015] S4-1: Generate corresponding foreground and background masks through the labels of training samples;

[0016] S4-2: Generate a weight mask of the distillation loss based on the generated foreground and background masks;

[0017] S4-3: Generate spatial attention mask through the spatial attention matrix of the feature map after feature extraction of PCB training sample features;

[0018] S4-4: Generate a channel attention mask through the channel attention matrix of the feature map after feature extraction of the PCB training sample features;

[0019] S4-5: Extracting global semantic information of feature graph through global feature extraction module;

[0020] S4-6: Calculate the local distillation loss of the knowledge distillation network and the YOLOv5 network through the generated foreground and background masks, weight masks, spatial attention masks, and channel attention masks;

[0021] S4-7: Calculate the global distillation loss of the knowledge distillation network and the YOLOv5 network based on the global semantic information of the feature map extracted in step S4-5;

[0022] S4-8: Calculate the total distillation loss through the local distillation loss and the global distillation loss.

[0023] Further, the local distillation loss is:

[0024]

[0025] Among them, L F is the local distillation loss, α, β, γ represent the weight coefficients of the foreground and background masks, F represents the feature map, M represents the foreground mask, and A s represents the spatial attention mask, A c represents the channel attention mask, S represents the weight mask, f represents the adaptive convolution layer, i, j represent the pixel coordinates in the feature map, H, W represent the height and width of the feature map respectively, C represents the channel dimension of the feature map, and k represents the channel dimension corresponding to the pixel.

[0026] Furthermore, the distillation loss function for calculating the global feature map is:

[0027]

[0028] Among them, λ represents the weight coefficient, F represents the global feature, and L represents the loss function.

[0029] Further, the total distillation loss is:

[0030] L d =L G +L F

[0031] Among them, L G is the total distillation loss.

[0032] Furthermore, the output detection frame is screened by removing redundant frames through confidence judgment and non-maximum suppression (NMS) method.

[0033] Compared with the prior art, the present invention has the following significant effects: by using a lightweight convolution module and a local-global-based knowledge distillation method (FGD), and by introducing a teacher network to supervise the feature graphs of the student network model, the present invention can achieve compression of model parameters (Parameters) and computational complexity (FLOPs) while maintaining the average detection accuracy (mAP) of the original PCB defect detection model, thereby greatly reducing the amount of computation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a structural diagram of the teacher network for PCB defect detection based on YOLOv5n of the present invention.

[0035] Figure 2 This is a network structure diagram of the C3 and C3Ghost modules of the present invention.

[0036] Figure 3 This is a diagram of the compressed PCB defect detection network structure of the present invention.

[0037] Figure 4 A flow chart for generating the foreground and background masks of the present invention.

[0038] Figure 5 A flow chart for weight mask generation of the present invention.

[0039] Figure 6 A flow chart for generating the attention mask of the present invention.

[0040] Figure 7 This is a network structure diagram of the global feature extraction module of the present invention.

[0041] Figure 8 This is a diagram of the PCB defect detection training framework based on knowledge distillation of the present invention.

[0042] Fig. 9 This is a detection effect diagram of the PCB defect detection algorithm after compression distillation of the present invention. DETAILED DESCRIPTION

[0043] The following will provide a complete and clear description of the technical solutions in the implementation of the present invention in conjunction with the accompanying drawings in the examples of the present invention.

[0044] like Figure 8 As shown, the present invention provides a PCB defect detection method based on model compression, which sequentially includes an image preprocessing process, a teacher network training process, a student network model compression process, a student network model training process, and a defect detection and identification process, and the specific process is as follows:

[0045] 1. Image preprocessing process

[0046] Image preprocessing process: The original PCB defect samples are randomly cropped and data augmented.

[0047] The specific processing process is:

[0048] A1: Randomly crop the original PCB defect sample image to form a sample image of 600×600 size.

[0049] A2: Perform data enhancement on the cropped image by methods such as brightness transformation and mirror flipping.

[0050] 2. Teacher network training process

[0051] Teacher network training process: reference Figure 1 This example provides a PCB defect detection algorithm based on YOLOv5. Since YOLOv5 is a well-known network, its structure is not described here. The specific processing process is as follows:

[0052] B1: Use the feature extraction module (Backbone) to extract features from the network model. In this operation, it should be noted that the enhanced data set in step 1 is used for training, such as Figure 1 As shown in the Backbone section.

[0053] B2: Use the feature fusion module (Neck) to implement the multi-scale fusion of features in step B1, such as Figure 1 As shown in the Neck section, the specific methods include:

[0054] (1) Use convolution operation to downsample features, and perform top-down multi-scale feature fusion through feature pyramid network (FPN). The specific process includes: first, downsampling is performed through convolution operation to obtain a smaller feature map, and then the feature map is upsampled and summed element by element with the previous layer feature.

[0055] (2) Deconvolution is used for upsampling, and bottom-up multi-scale feature fusion is achieved through a bidirectional fusion network (PAN). The specific process includes: downsampling the shallow features fused by the FPN network to make them the same scale as the feature map of the previous layer, and fusing the shallow features with the deep features from bottom to top by element-by-element summation.

[0056] B3: The feature map fused at the Neck end is sent to the detection head, the feature map after the detector convolution is grid-divided for prediction of the prediction box, the corresponding anchor box is generated through the grid, the confidence of each anchor box, the prediction box coordinates and the output vector for classification are predicted, and the final prediction box and prediction category are obtained through non-maximum suppression (NMS).

[0057] B4: Use the PCB defect samples in the training set in step 1 as training data and perform end-to-end training on the YOLOv5-based PCB defect detection network built in steps B1 to B3.

[0058] 3. Student network model compression process

[0059] Student network model compression process: refer to Figure 2 , Figure 3 ,This example provides a compression algorithm for PCB defect detection network based on YOLOv5. The specific processing process is:

[0060] C1: Reference Figure 2 , Ghost convolution is used to replace the traditional convolution of the BottleNeck part of the C3 feature extraction module in the original algorithm to achieve compression of the C3 feature extraction network module.

[0061] C2: Reference Figure 3 , use the compressed C3 network model to replace the C3 module in the teacher network model, realize the network compression of the teacher model, and use the compressed network model as the student model.

[0062] 4. Student network model training process

[0063] Student network model training process: refer to Figures 4 to 8 This example provides a knowledge distillation training method for a compressed PCB defect detection student model. The specific processing process is as follows:

[0064] D1: Reference Figure 4 , during the training process, the real labels are used to generate foreground and background masks to achieve the foreground and background area division of the PCB feature map. The foreground mask M i,j The calculation method is as follows, where (i, j) represents the pixel coordinates in the feature map, and r represents the foreground area. The corresponding background mask (1-M i,j ):

[0065]

[0066] D2: Reference Figure 5, the weighted mask is calculated by the foreground and background masks in step D1, and the weighted mask S i,j The calculation process is as follows, where H r With W r Represents the height and width of the real box, i.e. the foreground area r, N bg Indicates the number of pixels of the background mask.

[0067]

[0068]

[0069] D3: Reference Figure 6 , the spatial attention mask and channel attention mask are calculated through the feature map spatial attention and channel attention, the spatial attention G of the feature map s With channel attention G c The calculation process is as follows, where H, W represent the height and width of the feature map, C represents the channel dimension of the feature map, and F represents the feature map:

[0070]

[0071]

[0072] Spatial attention mask A of feature map s (F) and channel attention mask A c (F) The calculation process is as follows:

[0073]

[0074]

[0075] D4: Calculate the local distillation loss L through the foreground background mask, weight mask, spatial attention mask and channel mask calculated in steps D1, D2 and D3 F , where α, β, and γ represent weight coefficients, L represents the loss function, T and S represent the teacher network model and the student network model respectively, and F represents the feature map.

[0076]

[0077] D5: Reference Figure 7 , perform global feature extraction of knowledge distillation node feature graphs of teacher network model and student network model. The calculation process of global feature R(F) is as follows, where W k , W v1 , W v2 represents the convolutional layer, LN represents the regularization layer, and N p Represents the number of pixels in the feature map:

[0078]

[0079] D6: Calculate the distillation loss of the global features extracted in step D5, the global loss L G The calculation process is as follows, where F T With F S Represent the characteristic graphs of the teacher network and the student network respectively, and λ represents the weight coefficient:

[0080] L G =λ·∑(R(F T )-R(F S )) 2

[0081] D7: Calculate the total distillation loss L by using the local distillation loss and global distillation loss calculated in steps D4 and D6 d , the calculation process is as follows:

[0082] L d =L G +L F

[0083] D8: Reference Figure 8 , the PCB defect detection algorithm based on YOLOv5n is used as the teacher network, and the YOLOv5n model compressed by the C3Ghost module is used as the student network. The feature graph branches of the feature fusion part of the student network and the teacher network are derived and recorded as Neck1, Neck2, and Neck3 respectively. During the training process, the weights of the teacher network trained in step 2 are loaded, and the knowledge distillation is performed by performing distillation loss calculations from steps D1 to D8 on the feature graph inference results of Neck1, Neck2, and Neck3 of the teacher network and the inference results during the student network training process.

[0084] 5. Defect detection and identification process

[0085] Student network model training process: The student network model trained in step 4 is used to detect and identify PCB board defects. The specific processing process is as follows:

[0086] E1: During model inference, the student network model is separated independently.

[0087] E2: During PCB defect detection and identification, the student network loads the training weights of step 4 and stops updating the gradient of the parameters. The output of the network model is obtained by inputting PCB image samples into the network model after distillation and compression.

[0088] E3: Use confidence judgment and NMS to remove redundant boxes in the output results to achieve PCB defect location and identification. Use the PCB defect detection student model in step 4 to perform defect detection. The detection effect is referenced Fig. 9 .

[0089] Table 1 Comparison of average accuracy, computational complexity and model parameters of network models

[0090]

[0091]

[0092] The comparison between the method of the present invention and the YOLOv5n and YOLOv5n methods is shown in Table 1. The comparative data show that the method of the present invention can achieve the compression of model parameters (Parameters) and calculation amount (FLOPs) while maintaining the average detection accuracy (mAP) of the original PCB defect detection model.

Claims

1. A PCB board defect detection method based on model compression, characterized in that: Includes steps: Step S1: pre-processing the defective image of the original PCB and forming a training sample by region segmentation; Step S2: Design a lightweight feature extraction module, replace it with the feature extraction module in the YOLOv5 network to perform network compression, and build a compressed network; Step S3: Before knowledge distillation is performed, the YOLOv5 network is trained using the PCB samples generated in step S1. During the knowledge distillation process, the parameters of the network will not be updated. Step S4: Based on the YOLOv5 network trained in S3, during each iterative training process of the compression network constructed in step S2, the local and global knowledge distillation modules are used to calculate the distillation loss, and the parameters of the compression network are updated through the back propagation algorithm; Step S5: PCB board defect detection is performed using the compression network trained in step S4, and the output detection frame is screened to obtain the final defect prediction frame; The step S2 specifically includes: replacing the convolution layer of the C3 module BottleNeck in the YOLOv5 network with the Ghost convolution model, and replacing the C3 module in the YOLOv5 network with the compressed C3 module to perform feature extraction; The step S4 specifically includes: S4-1: Generate corresponding foreground and background masks through the labels of training samples; S4-2: Generate a weight mask of the distillation loss based on the generated foreground and background masks; S4-3: Generate spatial attention mask through the spatial attention matrix of the feature map after feature extraction of PCB training sample features; S4-4: Generate a channel attention mask through the channel attention matrix of the feature map after feature extraction of the PCB training sample features; S4-5: Extracting global semantic information of feature graph through global feature extraction module; S4-6: Calculate the local distillation loss of the compression network and the YOLOv5 network through the generated foreground and background masks, weight masks, spatial attention masks, and channel attention masks; S4-7: Calculate the global distillation loss of the compression network and the YOLOv5 network based on the global semantic information of the feature map extracted in step S4-5; S4-8: Calculate the total distillation loss by combining the local distillation loss and the global distillation loss.

2. The PCB board defect detection method based on model compression according to claim 1 is characterized in that: The step S1 specifically includes: randomly cropping the original PCB defect data set to form a 600×600 image and updating the relative position of the corresponding annotations; dividing the enhanced file into a training validation set and a test set in a ratio of 8:2, and the sample ratio of the training set to the validation set in the training validation set is 3:

2.

3. The PCB board defect detection method based on model compression according to claim 1 is characterized in that: The input end of the YOLOv5 network trained in step S3 processes the training data as follows: extracting features from the input image through the C3 feature extraction module, performing multi-scale fusion through FPN and PAN, and generating a priori boxes of different numbers and sizes through the detection head.

4. The PCB board defect detection method based on model compression according to claim 3 is characterized in that: The processing process of the FPN for multi-scale fusion includes: using upsampling to perform multi-scale fusion on features of different scales from top to bottom; the processing process of the PAN for multi-scale fusion includes: using downsampling to perform multi-scale fusion on features of different scales from bottom to top.

5. The PCB board defect detection method based on model compression according to claim 1 is characterized in that: The local distillation loss is: Among them, L F is the local distillation loss, α, β, γ represent the weight coefficients of the foreground and background masks, F represents the feature map, M represents the foreground mask, and A s represents the spatial attention mask, A c represents the channel attention mask, S represents the weight mask, f represents the adaptive convolution layer, i, j represent the pixel coordinates in the feature map, H, W represent the height and width of the feature map respectively, C represents the channel dimension of the feature map, k represents the channel dimension corresponding to the pixel, L represents the loss function, represents the spatial attention mask of the teacher network model, represents the spatial attention mask of the student network model, represents the channel attention mask of the teacher network model, Represents the channel attention mask of the student network model.

6. The PCB board defect detection method based on model compression according to claim 1 is characterized in that: The global distillation loss is: Among them, λ represents the weight coefficient, F represents the global feature, and L represents the loss function. represents the global feature map of the teacher network model, Represents the global feature map of the student network model.

7. The PCB board defect detection method based on model compression according to claim 5 is characterized in that: The total distillation loss is: L d =L G +L F Among them, L G is the global distillation loss.

8. The PCB board defect detection method based on model compression according to claim 1 is characterized in that: The output detection frame is screened as follows: redundant frames are removed by confidence judgment and non-maximum suppression (NMS) method.

Citation Information

Patent Citations

  • PCB defect detection algorithm based on convolutional neural network

    CN111709910A

  • PCB surface defect detection method based on improved YOLOv5 algorithm

    CN114372949A

  • Target detection method and target detection terminal based on knowledge distillation

    CN113743514A

  • Image target detection method and system based on lightweight neural network model

    CN114332666A