PCB defect classification method based on combination of YOLOv8 network and Transform
By combining YOLOv8 network with Transformer in PCB defect classification, and introducing an adaptive fine-grained channel attention mechanism and convolutional gating linear unit, the problems of large model size, low accuracy and slow detection speed in the prior art are solved, and efficient and accurate defect classification is achieved.
Patent Information
- Application Number
- CN202510085655.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing PCB defect classification methods have problems such as large model size, low recognition accuracy, and slow detection speed. It is difficult to optimize the number of parameters and calculation efficiency while improving the accuracy.
The PCB defect classification method based on the combination of YOLOv8 network and Transformer is adopted. Through the reasonable allocation of input channels, combined with the adaptive fine-grained channel attention mechanism and convolutional gating linear units, the network structure is optimized to improve computing efficiency and feature extraction capabilities.
It significantly improves the accuracy of the model, improves the perception of complex defects, and reduces the computational complexity and parameter quantity, achieving a good balance of performance and efficiency.
Smart Images

Figure CN120014340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of printed circuit board surface defect recognition, and in particular to a PCB defect classification method based on a combination of a YOLOv8 network and a Transformer. Background Art
[0002] As the basic platform of electronic devices, PCB (printed circuit board) plays a vital role in modern electronic products, and its quality directly determines the reliability and performance of electronic products. However, during the manufacturing process of PCB, various defects such as short circuit, open circuit and missing holes may occur, which will significantly reduce the electrical performance and even cause the failure of the whole machine. Therefore, in order to ensure product quality, timely detection and repair of PCB defects has become the key. Traditionally, automatic optical inspection (AOI) equipment has been widely used in the field of PCB defect detection. Through high-speed and high-precision visual processing technology, AOI can efficiently detect mounting errors and welding defects, thereby improving detection efficiency and reducing labor costs. However, AOI equipment is expensive and has high requirements for the detection environment, which has become its limitation.
[0003] In recent years, with the rapid development of deep learning technology, deep learning-based defect classification methods have become an important tool for solving PCB defect identification problems. Deep learning methods use convolutional neural networks (CNNs) to efficiently and accurately analyze images and can identify various PCB defects, including small target defects. Initial target detection models such as VGG and ResNet solved the gradient vanishing problem in deep networks by introducing deeper network structures and residual connections, further improving classification accuracy. Subsequently, the Inception series of models enhanced the model's ability to express diverse features through multi-scale convolutional structures, adapting to the diverse needs of PCB defects.
[0004] With the deepening of research, deep learning models for PCB defect classification have gradually integrated a variety of advanced technologies, such as MobileNet and EfficientNet, which have improved the ability of real-time classification by reducing the number of parameters and computational complexity. The introduction of the attention mechanism SE module or CBAM module effectively focuses on the key feature area and enhances the ability to classify tiny defects. In addition, transfer learning using pre-trained models (such as the ImageNet model) has achieved good classification results in small sample scenarios. However, the above methods all have more or less disadvantages such as large model size, low recognition accuracy, and slow detection speed.
[0005] Therefore, in the study of PCB defect classification methods, how to improve the accuracy while optimizing the number of parameters and computational efficiency has become the core difficulty of current technological development. This background provides an important research direction for further exploring new and efficient PCB defect classification algorithms. To this end, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is proposed. Summary of the invention
[0006] The technical problem to be solved by the present invention is: how to enable the defect classification model to obtain higher recognition performance with fewer parameters and lower computational complexity, while being able to effectively detect various types of defects, and provide a PCB defect classification method based on the combination of YOLOv8 network and Transformer.
[0007] The present invention solves the above technical problems through the following technical solutions, and the present invention comprises the following steps:
[0008] S1: Sample pretreatment
[0009] Preprocess the PCB defect sample data samples;
[0010] S2: Network construction
[0011] Combine the YOLOv8 classification network with two Transformer modules and assign input channels to obtain a preliminary PCB defect classification network;
[0012] S3: Network structure optimization
[0013] Based on the preliminary PCB defect classification network, an adaptive fine-grained channel attention mechanism module is introduced, and a convolutional gated linear unit is introduced into the Transformer module to obtain the final PCB defect classification network.
[0014] S4: Model training
[0015] The final PCB defect classification network is trained using the training set to obtain a PCB defect classification model that meets the performance indicators;
[0016] S5: Test result output
[0017] The PCB defect classification model is used to detect PCB defects on an independent test set to obtain the PCB defect classification detection results.
[0018] Furthermore, in step S1, the specific processing process is as follows:
[0019] S11: Randomly crop the defects in all images of the original PCB dataset to meet the data requirements of the classification task;
[0020] S12: Divide the images into a training set, a validation set, and a test set according to a set ratio, and adjust the image size in the training set to a set size.
[0021] Furthermore, in the step S11, the defect categories include missing holes, mouse bites, open circuits, short circuits, burrs, and copper debris.
[0022] Furthermore, in the step S2, in the backbone network CSPDarknet-53 of the YOLOv8 classification network, a first Transformer module is introduced between the fourth CBS module and the fifth CBS module, and a second Transformer module is introduced between the fifth CBS module and the SPPF layer.
[0023] Furthermore, after the fourth CBS module, the feature map is divided into two parts, entering the convolution branch and the Transformer branch respectively. The feature channel ratio of the convolution branch and the Transformer branch is 3:1. One part of the feature map enters the Transformer branch and is processed by the first Transformer module to capture global features using the multi-head self-attention mechanism. The other part of the feature map enters the convolution branch and is processed by the bottleneck layer to extract local features.
[0024] After the fifth CBS module, the feature map is divided into two parts, entering the convolution branch and the Transformer branch respectively. The feature channel ratio of the convolution branch and the Transformer branch is 1:1. One part of the feature map enters the Transformer branch and is processed by the second Transformer module. The multi-head self-attention mechanism is used to capture global features. The other part of the feature map enters the convolution branch and is processed by the bottleneck layer for local feature extraction.
[0025] Furthermore, in the step S3, in the backbone network CSPDarknet-53 of the YOLOv8 classification network, an adaptive fine-grained channel attention mechanism module is added after the second C2f module, and the processing process is as follows:
[0026] S301: Input feature map F∈R C×H×W , the channel descriptor U∈R is generated by global average pooling operation C :
[0027]
[0028] Among them, C, H and W represent the number of channels, length and width of the feature map respectively, and F n (i,j) represents the nThe value of the channel feature map at position (i, j), GAP(x) represents the global average pooling function;
[0029] S302: Introduce the band matrix B for local channel interaction and extract local features U through weighted operations of adjacent channels lc :
[0030]
[0031] Among them, the number of adjacent channels is k, b i are the elements in the band matrix B;
[0032] S303: The diagonal matrix D is used to capture the dependencies between all channels, thereby obtaining the global feature U gc :
[0033]
[0034] in, c Indicates the number of channels, d i are the elements in the diagonal matrix D;
[0035] S304: Introduce the correlation matrix M to characterize the relationship between channels:
[0036]
[0037] S305: Calculate the fused global channel weight and local channel weight and
[0038]
[0039] S306: Calculate the final weight W:
[0040]
[0041] Among them, θ is a learnable parameter, σ represents the Sigmoid activation function;
[0042] S307: The final weight W is multiplied by the input feature map F to generate the final output feature map F * .
[0043] Furthermore, in step S3, the specific processing process of the convolutional gated linear unit is as follows:
[0044] S311: By introducing a depth-wise separable convolution with a kernel size of 3×3 into the gated branch of the traditional GLU;
[0045] S312: Use a 1×1 convolution kernel to perform linear combinations between channels on the output of the depthwise convolution, i.e., point-by-point convolution operations.
[0046] Compared with the prior art, the present invention has the following advantages: the PCB defect classification method based on the combination of the YOLOv8 network and the Transformer establishes a PCB defect classification model based on the YOLOv8 classification network and the Transformer, combining the advantages of both; the feature extraction capability is enhanced and the computational efficiency is improved by rationally allocating input channels, and the ability of the model to fuse global and local information is enhanced by adding an adaptive fine-grained channel attention mechanism; the convolutional gated linear unit is used to reduce the computational complexity while enhancing the ability of the Transformer to express nonlinear features, thereby significantly improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of the process of PCB defect classification method in an embodiment of the present invention;
[0048] Figure 2 1 is an example of various defect samples of printed circuit boards in the NEU-CLS database in an embodiment of the present invention, wherein (a) is a burr, (b) is a missing hole, (c) is a short circuit, (d) is a copper scrap, (e) is an open circuit, and (f) is a rat bite;
[0049] Figure 3 is a schematic diagram of the structure of an adaptive fine-grained channel attention mechanism module in an embodiment of the present invention;
[0050] Figure 4 is a schematic diagram of the structure of a convolutional gated linear unit in an embodiment of the present invention;
[0051] Figure 5 It is a schematic diagram of the overall structure of the model in an embodiment of the present invention;
[0052] Figure 6 It is a comparison chart of parameter quantities and accuracy of different models in PCB defect classification performance in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.
[0054] This embodiment provides a technical solution: a PCB defect classification method based on a combination of a YOLOv8 network and a Transformer, comprising the following steps:
[0055] S1: Sample pretreatment
[0056] In step S1, the following two sub-steps are included:
[0057] S11: Obtain a public dataset of surface defects on printed circuit boards (PCBs), and perform necessary preprocessing on the images in the dataset according to the task requirements, including randomly cropping all defects in each image, and dividing them into training and test sets in a ratio of 7:1:2; the distribution of the dataset is shown in Table 1. Examples of typical 6 types of defect samples, namely spur, missing hole, short, spurious copper, open circuit, and mouse bite, are shown in Table 1. Figure 2 As shown in (a)-(f);
[0058] S12: The images in the training set and the validation set are uniformly resized to 224×224. The purpose of unifying the size of all images in the training set and the validation set is to adapt to the input size requirements of the model. Different input sizes will make it impossible for the computing device to build a unified batch for training.
[0059] Table 1 Dataset distribution
[0060] Defect Category Training set Validation set Test Set Hole 347 50 100 Rat bite 344 49 99 Circuit Breaker 337 48 97 Short Circuit 343 49 99 glitch 341 49 98 Miscellaneous copper 352 50 101 total 2064 295 594
[0061] S2: Combine the YOLOv8 classification network with two Transformer modules to enhance the feature extraction capability and improve the computational efficiency by properly allocating input channels;
[0062] In step S2, the following sub-steps are included:
[0063] Since the Transformer's self-attention mechanism has a high computational complexity, directly applying it to all channels will significantly increase the amount of computation. Therefore, the present invention reasonably distributes the number of input channels, optimizes the computational efficiency, and achieves the purpose of extracting global features. The feature channel is divided into two branches, one part enters the convolution branch for local feature extraction. The other part is processed by the Transformer branch (Transformer module) and uses the multi-head self-attention mechanism (MHSA) to capture global features. Compared with traditional convolutional neural networks, Transformer is better at modeling long-range dependencies, which is particularly important in PCB defect detection scenarios because some minor defects may require global context information to determine their type or location. Due to the high computational complexity of the Transformer structure, if the feature map with a high number of channels is directly processed, the training efficiency may be reduced. To this end, the present invention first allocates the feature channels to the convolution branch and the Transformer branch in a ratio of 3:1 through parameter optimization and experimental verification. This division strategy effectively balances the computational cost and performance of the two modules. The convolution branch obtains more channels, which helps to enhance the efficiency of local feature extraction. The Transformer branch retains enough channels to process global context information and improve the model's ability to perceive complex defects. After the feature extraction of the convolution and Transformer branches is completed, before the feature map enters the SPPF layer, the number of channels is redivided in a 1:1 ratio for feature extraction again.
[0064] S3: Further optimization of network model
[0065] In step S3, the following sub-steps are included:
[0066] S31: Add an adaptive fine-grained channel attention mechanism module (FCA) after the second C2f module of YOLOv8 to enhance the model's ability to fuse global and local information. The definition is as follows:
[0067]
[0068]
[0069] Among them, the feature map F∈R C×H×W , C, H and W represent the number of channels, length and width of the feature map respectively, and the channel descriptor U∈R is generated by global average pooling operation C ; F n (i,j) represents the nThe value of the channel feature map at position (i, j), GAP(x) represents the global average pooling function, which can be used to transform the shape of the feature map F from C×H×W to C×1×1; then, the band matrix B is introduced for local channel interaction, and local features are extracted through weighted operations of adjacent channels. The form of the band matrix is B=[b1,b2,b3,...,b k ],U lc represents local information, and k represents the number of adjacent channels. In order to obtain global channel information and enhance its representation ability, the diagonal matrix D is used to capture the dependencies between all channels. The form of the diagonal matrix D is D = [d1, d2, d3, …, d c ],U gc represents global information, c represents the number of channels, and M represents the correlation matrix, which is used to represent the complex relationship between channels. and Respectively represent the fused global channel weight and local channel weight, θ is a learnable parameter, and σ represents the Sigmoid activation function. This method promotes the interaction between the two, and at the same time, it also effectively avoids redundant cross-correlation operations between global and local information. Finally, the calculated weight W is multiplied by the input feature map F to generate the final output feature map. F is the input feature map, and is the final output feature map.
[0070] S32: Adding convolutional gated linear units (CGLU) to the Transformer module can reduce computational complexity while enhancing the Transformer's ability to express nonlinear features. By introducing 3×3 depthwise separable convolutions in the gated branch of the traditional GLU, the convolution operation is divided into two parts: depthwise convolution and pointwise convolution, which reduces the computational complexity and number of parameters of the convolution. For the input feature map H×W×C in And the output feature map H×W×C out , the computational complexity of standard convolution is H×W×C in ×C out ×k×k. Where: H and W represent the height and width of the feature map; C in , C out Indicates the number of input and output channels; k is the size of the convolution kernel (such as 3×3). Depthwise separable convolution uses a single-channel convolution kernel to perform convolution operations on each channel of the input feature map independently (depthwise convolution), and does not fuse features between different channels. Then, a 1×1 convolution kernel is used to perform linear combinations between channels on the output of the depthwise convolution (pointwise convolution). The computational complexity is H×W×C in ×k×k+H×W×C in ×C out .
[0071] S33: The performance of the model after adding the Transformer module, the adaptive fine-grained channel attention mechanism module and the convolutional gated linear unit is verified again on an independent test set. The results are shown in Table 2. It can be seen from Table 2 that the above improvements are helpful for improving the performance of the model.
[0072] Table 2 Results of performance improvement of various aspects of the model
[0073]
[0074] S4: Verify the performance of the proposed model on an independent test set. The detailed plan is as follows:
[0075] S41: In order to verify the effectiveness of the improved model of the present invention, the results of the improved model are compared with those of other classification models on the test set. The comparison results are shown in Table 3:
[0076] Table 3 Classification model experimental comparison
[0077] Model Accuracy Recall TOP-1 Accuracy Parameter quantity (M) / floating point number (G) CSPDarknet53 97.6% 94.5% 96.7% 5.08 / 12.4 EfficientViT 95.9% 93.1% 95.6% 2.41 / 4.6 MobileNetV4 93.2% 91.1% 92.9% 4.14 / 17.8 ResNet50 96.5% 95.1% 96.1% 26.1 / 68.8 ConvNeXtv2 96.6% 95.2% 95.6% 4.12 / 4.2 Model of the present invention 98.5% 97.9% 99% 3.83 / 10.5
[0078] From the experimental results in the table, we can see that the model of the present invention outperforms other comparison models in terms of precision, recall and TOP-1 accuracy, reaching 98.5%, 97.9% and 99% respectively, showing a strong performance advantage. At the same time, the number of parameters and floating-point operations of the model of the present invention are 3.83M and 10.5×10 9 , achieving a good balance between performance and efficiency. In comparison, the CSPDarknet53 model is inferior to the model of the present invention in terms of precision and TOP-1 accuracy, but the number of parameters is 5.08M. Although the number of parameters and floating-point operations of the EfficientViT model is small, its precision, recall and TOP-1 accuracy are also relatively low. Although ResNet50 is close to the model of the present invention in terms of precision and recall, its number of parameters and floating-point operations are higher, which will lead to a large computational overhead. In addition, although the ConvNeXtv2 model has great advantages in the number of parameters and floating-point operations, its precision and recall are inferior to the model of the present invention. Overall, the model of the present invention achieves a better balance between performance indicators and computational efficiency, proving the effectiveness of the improvement of the model of the present invention.
[0079] In summary, the PCB defect classification method based on the combination of YOLOv8 network and Transformer in the above embodiment establishes a PCB defect classification model based on YOLOv8 classification network and Transformer, combining the advantages of both; by reasonably allocating input channels to enhance the ability of feature extraction and improve computational efficiency, and by adding an adaptive fine-grained channel attention mechanism to enhance the model's ability to fuse global and local information; by using convolutional gated linear units to reduce computational complexity while enhancing the ability of Transformer to express nonlinear features, the accuracy of the model is significantly improved. The effectiveness of the improvement has been demonstrated through various comparative experiments, and the accuracy has been successfully increased from 96.7% to 99%, with a model size of only 3.83M, providing a feasible solution for PCB defect classification. The model proposed in the present invention has a high accuracy rate, and is expected to be deployed on edge devices in the future to be put into production and application.
[0080] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A PCB defect classification method based on the combination of YOLOv8 network and Transformer, characterized in that: The following steps are involved: S1: Sample pretreatment Preprocess the PCB defect sample data samples; S2: Network construction Combine the YOLOv8 classification network with two Transformer modules and assign input channels to obtain a preliminary PCB defect classification network; S3: Network structure optimization Based on the preliminary PCB defect classification network, an adaptive fine-grained channel attention mechanism module is introduced, and a convolutional gated linear unit is introduced into the Transformer module to obtain the final PCB defect classification network. S4: Model training The final PCB defect classification network is trained using the training set to obtain a PCB defect classification model that meets the performance indicators; S5: Test result output The PCB defect classification model is used to detect PCB defects on an independent test set to obtain the PCB defect classification detection results.
2. According to claim 1, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is characterized in that: In step S1, the specific processing process is as follows: S11: Randomly crop the defects in all images of the original PCB dataset to meet the data requirements of the classification task; S12: Divide the images into a training set, a validation set, and a test set according to a set ratio, and adjust the image size in the training set to a set size.
3. According to claim 2, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is characterized in that: In the step S11 , the defect categories include missing holes, mouse bites, open circuits, short circuits, burrs, and copper debris.
4. According to claim 1, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is characterized in that: In the step S2, in the backbone network CSPDarknet-53 of the YOLOv8 classification network, a first Transformer module is introduced between the fourth CBS module and the fifth CBS module, and a second Transformer module is introduced between the fifth CBS module and the SPPF layer.
5. According to claim 4, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is characterized in that: After the fourth CBS module, the feature map is divided into two parts, entering the convolution branch and the Transformer branch respectively. The feature channel ratio of the convolution branch and the Transformer branch is 3:
1. One part of the feature map enters the Transformer branch and is processed by the first Transformer module to capture global features using the multi-head self-attention mechanism. The other part of the feature map enters the convolution branch and is processed by the bottleneck layer for local feature extraction. After the fifth CBS module, the feature map is divided into two parts, entering the convolution branch and the Transformer branch respectively. The feature channel ratio of the convolution branch and the Transformer branch is 1:
1. One part of the feature map enters the Transformer branch and is processed by the second Transformer module. The multi-head self-attention mechanism is used to capture global features. The other part of the feature map enters the convolution branch and is processed by the bottleneck layer for local feature extraction.
6. A PCB defect classification method based on the combination of YOLOv8 network and Transformer according to claim 1 or 5, characterized in that: In step S3, in the backbone network CSPDarknet-53 of the YOLOv8 classification network, an adaptive fine-grained channel attention mechanism module is added after the second C2f module, and the processing process is as follows: S301: Input feature map F∈R C×H×W , the channel descriptor U∈R is generated by global average pooling operation C : Among them, C, H and W represent the number of channels, length and width of the feature map respectively, and F n (i,j) represents the n The value of the channel feature map at position (i, j), GAP(x) represents the global average pooling function; S302: Introduce the band matrix B for local channel interaction and extract local features U through weighted operations of adjacent channels lc : Among them, the number of adjacent channels is k, b i are the elements in the band matrix B; S303: The diagonal matrix D is used to capture the dependencies between all channels, thereby obtaining the global feature U gc : in, c Indicates the number of channels, d i are the elements in the diagonal matrix D; S304: Introduce the correlation matrix M to characterize the relationship between channels: S305: Calculate the fused global channel weight and local channel weight and S306: Calculate the final weight W: Among them, θ is a learnable parameter, σ represents the Sigmoid activation function; S307: The final weight W is multiplied by the input feature map F to generate the final output feature map F * .
7. According to claim 6, a PCB defect classification method based on the combination of YOLOv8 network and Transformer is characterized in that: In step S3, the specific processing process of the convolutional gated linear unit is as follows: S311: By introducing a depth-wise separable convolution with a kernel size of 3×3 into the gated branch of the traditional GLU; S312: Use a 1×1 convolution kernel to perform linear combinations between channels on the output of the depthwise convolution, i.e., point-by-point convolution operations.
Citation Information
Cited By
Lightweight Transform defect detection method and device for high-density PCB (Printed Circuit Board)
CN120726034A
Lightweight transformer defect detection method and device for high-density PCBs
CN120726034B
Defect identification method based on domain feature fusion
CN120912525A