High-precision real-time printed circuit board defect target detection method
By using the EfficientFormerV2 model and C2f-ELA module in the detection of defect targets of printed circuit boards, combined with convolution and Transformer's self-attention module, the problem of poor recognition of small and medium-sized targets in the existing technology is solved, and high-precision real-time detection is achieved to meet the detection needs of the industry.
Patent Information
- Application Number
- CN202510087703.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
The existing printed circuit board defect detection technology has poor small and medium-sized target recognition effect, and the recognition accuracy of printed circuit board defect detection is not high.
A high-precision real-time object detection method for printed circuit board defects is proposed. The EfficientFormerV2 model is used as the reference feature network, combined with convolution and Transformer's self-attention module, and the target feature information is highlighted in the feature fusion stage through the C2f-ELA module.
It realizes high-precision detection of printed circuit board defect targets in environments such as images, videos, cameras and other environments in real-time detection of printed circuit board defect targets in real-time, and has fewer model parameters, and is simple and efficient in structure.
Smart Images

Figure CN120070340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, in particular to object detection technology, and specifically relates to a high-precision real-time printed circuit board defect object detection method. Background Art
[0002] Object detection is an important task in the field of computer vision, and its purpose is to identify and locate multiple target objects to be detected in an image. Existing object detection methods can be divided into two categories, namely traditional methods and deep learning-based methods. With the rise and rapid development of deep learning technology, the current deep learning-based object detection technology has become the mainstream method and has been successfully applied to various industries, such as smart cities, intelligent transportation, intelligent manufacturing and other fields.
[0003] Printed Circuit Boards (PCBs) are key interconnects for assembling electronic components and play a major role in many electronic products and devices. At the same time, the quality of PCBs controls the quality of the entire electronic products and devices. Therefore, in order to ensure the quality of electronic products and devices, the quality inspection of PCBs is an important and necessary step. If defects and flaws of PCBs can be detected before leaving the factory, certain economic losses and disputes can be effectively avoided.
[0004] Traditional PCB defect detection methods have many limitations, such as long detection time, high detection cost, and small detection scale. Currently, they can no longer meet the needs of contemporary automated production. At present, with the continuous research and development of artificial intelligence and deep learning technologies, they have become the best choice to replace traditional algorithms. Deep learning-based object detection methods automatically learn object features, significantly improving detection performance. Currently, deep learning-based object detection methods can be generally divided into methods based on Convolutional Neural Network (CNN) and methods based on Transformer technology. CNN-based object detection methods can be further divided into two-stage object detection methods and single-stage object detection methods. Two-stage object detection methods first generate candidate regions, then use CNN to extract image region features, and finally predict region classification and object position regression. Representative methods include the Faster RCNN series of algorithms [3]; single-stage object detection methods directly use CNN to extract object features from the entire image, and finally perform object classification and object bounding box regression on the output feature map. Representative methods include the YOLO series of algorithms. In recent years, the Transformer model has achieved great success in the field of natural language processing, and its self-attention mechanism has powerful modeling capabilities. DETR (Detection Transformer) introduces Transformer into the field of object detection, creating a new detection paradigm. This algorithm first uses CNN to extract image features, then encodes the features through an encoder, and uses a decoder to generate the category and position of the object. Finally, the Hungarian matching algorithm is used to align and detect object features. DETR does not require additional candidate region generation and post-detection processing steps, simplifying the object detection process. Although DETR has achieved good results, its training process is slow and the computational complexity is high. To solve these problems, subsequent models such as DINO have significantly accelerated the detection speed and improved the detection performance.
[0005] Nowadays, deep learning-based object detection technology has been applied to PCB defect detection. Wu Ruilin et al. proposed the "Defect Detection Method for Class-Incremental Printed Circuit Boards Based on YOLOX". This method applies the knowledge distillation method to the output features and intermediate features of the model to promote the transfer of knowledge of old defect categories, enabling small models to also have the efficient feature detection ability of large models. Wang Longye et al. proposed a method based on the YOLOv5s algorithm, introducing an attention mechanism to enhance the channel features of the feature map, and at the same time using a weighted bidirectional feature pyramid network in the feature fusion stage to improve the model's detection ability for small target defects on printed circuit boards. Finally, there are also some works that implement electronic circuit board defect detection based on the YOLO series of algorithms.
[0006] In order to solve the technical defects existing in the above-mentioned prior art, the applicant proposes a high-precision real-time method for detecting defective targets on printed circuit boards. Summary of the Invention
[0007] The object of the present invention is to solve the technical problem of poor recognition effect of small targets in the existing printed circuit board defect detection technology, and at the same time to improve the recognition accuracy of the defective target detection on printed circuit boards, and to propose a high-precision real-time method for detecting defective targets on printed circuit boards.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0009] A high-precision real-time method for detecting defective targets on printed circuit boards, comprising the following steps:
[0010] Step 1: Use the network to collect defective pictures of printed circuit boards and integrate them with the open-source data set, and screen and sort the images.
[0011] Step 2: Further process the collected images, screen suitable images, and use methods such as Gaussian filtering to process the screened images to obtain a batch of high-quality defective images of printed circuit boards.
[0012] Step 3: Use the Labelimg tool to annotate the data set, obtain a txt file, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0013] Step 4: Construct a high-precision real-time defective target detection model for printed circuit boards.
[0014] Step 5: Put the constructed data set into the high-precision real-time defective target detection model for printed circuit boards, train to obtain a trained model, and use accuracy, recall rate, and mean average precision of all categories as evaluation indicators.
[0015] Step 6: Use the trained high-precision real-time defective target detection model for printed circuit boards to perform image prediction on the detection image for accurate recognition.
[0016] In Step 1, it includes the following steps:
[0017] Step 1-1: Search for images of defective printed circuit boards through the Internet by keywords and download them.
[0018] Step 1-2: Find the open-source defective image data set of printed circuit boards on the open-source data website, and download the corresponding label file at the same time.
[0019] In step 2, specifically: first, the collected printed circuit board defect images are manually screened to eliminate images with low resolution and poor quality, and then preprocessed. The image preprocessing is carried out on the Pycharm platform, and a batch of high-quality printed circuit board defect images are obtained through Gaussian filtering method.
[0020] In step 3, it includes the following steps:
[0021] Step 3-1: Use the Labelimg picture annotation tool to annotate the collected printed circuit board defect images. The defects are divided into six categories: missing hole, mouse bite, open circuit, short circuit, spur, and spurious copper. Among them, missing hole is "missing_hole", mouse bite is "mouse_bite", open circuit is "open_circuit", short circuit is "short", spur is "spur", and spurious copper is "spurious_copper"; obtain the annotated txt file;
[0022] Step 3-2: After completing the annotation work, divide the dataset according to the ratio of 8:1:1 to obtain the training set, validation set, and test set.
[0023] In step 4, a high-precision printed circuit board defect target detection model constructed includes a feature extraction part, a feature fusion part, and a model prediction part;
[0024] The input end of the feature extraction part is used to input images. The output end of the feature extraction part is respectively connected to the input of the feature fusion part and the SPPF module. The output of the feature fusion part is connected to the detection head module in the model prediction part;
[0025] The feature extraction part uses EfficientFormerV2;
[0026] Specifically, the feature extraction part is as follows:
[0027] The input image is input to the Stem module. The output of the Stem module is connected to the input of the first local feature extraction module. The output of the first local feature extraction module is connected to the input of the first downsampling fusion module. The output of the first downsampling fusion module is connected to the input of the second local feature extraction module. The output of the second local feature extraction module is denoted as feature F 1 , the output of the second local feature extraction module is connected to the input of the second downsampling fusion module. The output of the second downsampling fusion module is connected to the input of the first local global feature extraction fusion module. The output of the first local global feature extraction fusion module is denoted as feature F 2, the output of the first local-global feature extraction and fusion module is connected to the input of the third downsampling and fusion module, the output of the third downsampling and fusion module is connected to the input of the second local-global feature extraction and fusion module, and the output of the second local-global feature extraction and fusion module is denoted as feature F 3 ;
[0028] The specific content of the feature fusion part is: feature F 3 is input to the SPPF module, the output of the SPPF module is connected to the input of the first upsampling module, the output of the first upsampling module and feature F 2 are connected to the input of the first Concat module, the output of the first Concat module is connected to the input of the first C2f-ELA module, and the output of the first C2f-ELA module is denoted as feature F ce2 , the output of the first C2f-ELA module is connected to the input of the second upsampling module, the output of the second upsampling module and feature F 1 are connected to the input of the second Concat module, the output of the second Concat module is connected to the input of the second C2f-ELA module, and the output of the second C2f-ELA module is denoted as feature F ce1 , the input of the second C2f-ELA module is connected to the input of the first CBS module, the input of the first CBS module and feature F ce2 are connected to the input of the third Concat module, the output of the third Concat module is connected to the input of the third C2f-ELA module, and the output of the third C2f-ELA module is denoted as feature F ce21 , the output of the third C2f-ELA module is connected to the input of the second CBS module, the output of the second CBS module and feature F 3 are connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the fourth C2f-ELA module, and the output of the fourth C2f-ELA module is denoted as feature F ce3 , finally, feature F ce1 , feature F ce21 and feature F ce3 are fed into the model prediction module;
[0029] The specific content of the model prediction part is: feature F obtained from the feature fusion part ce1 is input to the Detect1 module, feature F ce21 is input to the Detect2 module, feature F ce3 is input to the Detect3 module.
[0030] The specific content of C2f-ELA is: the input feature F is sent into the input of the first ELA module, and then split into two features denoted as feature F e1 and feature F e2 , and the output feature F of the first ELA modulee1 is sent to the first Bottleneck module, and the output of the first Bottleneck module is denoted as feature F b , and then the feature F is split b to obtain feature F b1 and feature F b2 , feature F b1 is sent into the input of the second ELA module. The output of the second ELA module is connected to the input of the second Bottleneck module, and the output of the second Bottleneck module is denoted as feature F eb , and finally the feature F b2 , feature F e2 and feature F eb are sent to the fifth Concat module. The output of the fifth Concat module is sent to the third CBS module, and the output feature of the third CBS module is denoted as feature F o .
[0031] Specifically, for SPPF: the input feature F i is input to the fourth CBS module, and the output feature of the fourth CBS module is denoted as feature F i1 . The output of the fourth CBS module is connected to the input of the first max-pooling layer module, and the output feature of the first max-pooling layer module is denoted as feature F m1 . The output of the first max-pooling layer module is connected to the input of the second max-pooling layer module, and the output feature of the second max-pooling layer module is denoted as feature F m2 . The input of the second max-pooling layer module is connected to the output of the third max-pooling layer module, and the output feature of the third max-pooling layer module is denoted as feature F m3 , feature F i1 , feature F m1 , feature F m2 and feature F m3 are connected to the input of the sixth Concat. The output of the sixth Concat is connected to the output of the fifth CBS module.
[0032] Specifically, for CBS: the input feature is sent into the input of the first convolutional layer. The output of the first convolutional layer is connected to the input of the batch normalization layer, and the output of the batch normalization layer is connected to the activation function SiLU.
[0033] Specifically, for the EfficientFormerV2 network:
[0034] In the Stem module, the input features are connected to the second convolutional layer, the output of the second convolutional layer is connected to the second batch normalization layer, the output of the second batch normalization layer is connected to the first GeLU activation function, the connection of the first GeLU activation function is connected to the third convolutional layer, the output of the third convolutional layer is connected to the third batch normalization layer, and the output of the third batch normalization layer is connected to the second GeLU activation function;
[0035] In the local feature extraction module, the input features are connected to the input of the first feature extraction block, and the output of the first feature extraction block is connected to the input of the second feature extraction block; where the first feature extraction block and the second feature extraction block each contain a 1×1 convolutional layer, a feature extraction block batch normalization layer, a GeLU activation function, a 3×3 depthwise separable convolutional layer, a batch normalization layer, a GeLU activation function, a 1×1 convolution, a batch normalization layer, and a residual connection layer;
[0036] In the downsampling fusion module, the input features are connected to the input of the first local feature extraction module, the output of the first local feature extraction module is connected to the input vector of the first global feature extraction module, and the output of the first global feature extraction module is connected to the input vector of the second global feature extraction module; where the global feature extraction module contains a 3×3 convolutional layer and a batch normalization downsampling layer, three 1×1 convolutional layers and batch normalization layers respectively used to obtain Query, Key, and Value for performing self-attention operations, an upsampling layer for restoring the dimension of the feature map, a 1×1 convolutional layer and a batch normalization layer, and a feature extraction block;
[0037] In the local-global feature extraction fusion module, the input features are connected to the input of the third local feature extraction module, the output of the third local feature extraction module is connected to the input of the fourth local feature extraction module, the output of the fourth local feature extraction module is connected to the input of the first local-global feature fusion block, and the output of the first local-global feature fusion block is connected to the input of the fifth local-global feature fusion block;
[0038] In the local-global feature fusion block, the input feature is connected to the input of the first downsampling module Downsample. The output of the first downsampling module Downsample is connected to the input of the fourth convolutional layer. The output of the fourth convolutional layer is connected to the input of the fourth batch normalization layer. The output of the fourth batch normalization layer is connected to the input of the first Attention module. The output of the first Attention module is connected to the input of the first upsampling module Upsample. The output of the first upsampling module Upsample is connected to the input of the fifth convolutional layer. The output of the fifth convolutional layer is connected to the input of the fifth batch normalization layer. The output of the fifth batch normalization layer is connected to the input of the sixth convolutional layer. The output of the sixth convolutional layer is connected to the input of the sixth batch normalization layer. The output of the sixth batch normalization layer is connected to the input of the third GeLU activation function. The output of the third GeLU activation function is connected to the input of the first depthwise separable convolutional layer. The output of the first depthwise separable convolutional layer is connected to the input of the seventh batch normalization layer. The output of the seventh batch normalization layer is connected to the input of the fourth GeLU activation function. The output of the fourth GeLU activation function is connected to the input of the seventh convolutional layer. The input of the seventh convolutional layer is connected to the input of the eighth batch normalization layer.
[0039] Compared with the prior art, the present invention has the following technical effects:
[0040] 1) The present invention proposes an efficient method for detecting defective targets on printed circuit boards, which can achieve high-precision detection of defective targets on printed circuit boards in environments such as images, videos, and cameras in real-world scenarios;
[0041] 2) The present invention proposes an efficient framework for detecting defective targets on printed circuit boards, using the EfficientFormerV2 model as the benchmark feature network, which combines convolution and the self-attention module of Transformer to achieve complete extraction and characterization of the features of the input source (images, videos, cameras, etc.). A brand-new C2f-ELA module is proposed in the feature fusion stage, which combines the visual attention mechanism to highlight the target feature information;
[0042] 3) The present invention can meet the requirements of real-time detection of defective targets on printed circuit boards in the industrial field. While ensuring high-efficient target extraction ability, it has fewer model parameters. At the same time, for the convenience of deployment in actual scenarios, the target detection framework designed by the present invention has a simple and efficient structure. By adopting an end-to-end method, inputting a single picture can directly output the target result, and it can achieve the detection of defective targets on PCBs in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The following further describes the present invention in conjunction with the drawings and embodiments:
[0044] Figure 1 is the overall flowchart of the present invention;
[0045] Figure 2 This is the overall framework diagram of the high-precision real-time printed circuit board defect target detection model in the present invention;
[0046] Figure 3 is Figure 2 the schematic diagram of the CBS module in;
[0047] Figure 4 is Figure 2 the schematic diagram of the C2f-ELA module in;
[0048] Figure 5 is Figure 4 the schematic diagram of the Bottleneck module of the C2f-ELA module in;
[0049] Figure 6 is Figure 2 the schematic diagram of the SPPF structure in;
[0050] Figure 7 is Figure 2 the schematic diagram of the EfficientFormerv2 model in;
[0051] Figure 8 is Figure 4 the schematic diagram of the ELA module in;
[0052] Figure 9 This is the experimental result index diagram of the YOLOv8s model on the printed circuit board defect dataset in the embodiment;
[0053] Figure 10 This is the experimental result index diagram of the present invention on the printed circuit board defect dataset in the embodiment;
[0054] Figure 11 Example pictures of printed circuit board defects in the embodiment;
[0055] Figure 12 This is the detection result diagram of the present invention on the printed circuit board defect pictures. Detailed implementation manners
[0056] A high-precision real-time printed circuit board defect target detection method includes the following steps:
[0057] Step 1: Use the network to collect printed circuit board defect pictures and integrate them with the open-source dataset, and screen and sort the images;
[0058] Step 2: Further process the collected images, screen suitable images, and use methods such as Gaussian filtering to process the screened images to obtain a batch of high-quality printed circuit board defect images;
[0059] Step 3: Use the Labelimg tool to annotate the dataset, obtaining txt files, and divide them into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0060] Step 4: Construct a high-precision real-time printed circuit board defect target detection model;
[0061] Step 5: Put the constructed dataset into the high-precision real-time printed circuit board defect target detection model, train to obtain the trained model, and use accuracy, recall rate, and the mean average precision of all categories as evaluation metrics;
[0062] Step 6: Use the trained high-precision real-time printed circuit board defect target detection model to perform image prediction on the detection image for accurate recognition.
[0063] In Step 1, the following steps are included:
[0064] Step 1-1: Search the Internet for keywords "printed circuit board defects" to obtain images and download them;
[0065] Step 1-2: Find an open-source printed circuit board defect image dataset on an open-source data website and download the corresponding label file at the same time.
[0066] In Step 2, specifically: First, manually select the collected printed circuit board defect images, select and delete images with low resolution and poor quality, and then perform preprocessing. Perform image preprocessing on the Pycharm platform, and use the Gaussian filtering method to process to obtain a batch of high-quality printed circuit board defect images.
[0067] In Step 3, the following steps are included:
[0068] Step 3-1: Use the Labelimg picture annotation tool to annotate the collected printed circuit board defect images. The defects are divided into six categories: missing hole, mouse bite, open circuit, short circuit, spur, and spurious copper. Among them, missing hole is "missing_hole", mouse bite is "mouse_bite", open circuit is "open_circuit", short circuit is "short", spur is "spur", and spurious copper is "spurious_copper"; obtain the annotated txt files;
[0069] Step 3-2: After completing the annotation work, divide the dataset according to the ratio of 8:1:1 to obtain the training set, the validation set, and the test set.
[0070] In Step 4, a constructed high-precision printed circuit board defect target detection model includes a feature extraction part, a feature fusion part, and a model prediction part;
[0071] The input end of the feature extraction part is used to input images. The output ends of the feature extraction part are respectively connected to the input of the feature fusion part and the SPPF module. The output of the feature fusion part is connected to the detection head module in the model prediction part;
[0072] The feature extraction part adopts / uses EfficientFormerV2;
[0073] Specifically, the feature extraction part is as follows:
[0074] The input image is input to the Stem module. The output of the Stem module is connected to the input of the first local feature extraction module. The output of the first local feature extraction module is connected to the input of the first downsampling fusion module. The output of the first downsampling fusion module is connected to the input of the second local feature extraction module. The output of the second local feature extraction module is denoted as feature F 1 , the output of the second local feature extraction module is connected to the input of the second downsampling fusion module. The output of the second downsampling fusion module is connected to the input of the first local-global feature extraction and fusion module. The output of the first local-global feature extraction and fusion module is denoted as feature F 2 , the output of the first local-global feature extraction and fusion module is connected to the input of the third downsampling fusion module. The output of the third downsampling fusion module is connected to the input of the second local-global feature extraction and fusion module. The output of the second local-global feature extraction and fusion module is denoted as feature F 3 ;
[0075] Specifically, the feature fusion part is as follows: Feature F 3 is input to the SPPF module. The output of the SPPF module is connected to the input of the first upsampling module. The output of the first upsampling module and feature F 2 are connected to the input of the first Concat module. The output of the first Concat module is connected to the input of the first C2f-ELA module. The output of the first C2f-ELA module is denoted as feature F ce2 , the output of the first C2f-ELA module is connected to the input of the second upsampling module. The output of the second upsampling module and feature F 1 are connected to the input of the second Concat module. The output of the second Concat module is connected to the input of the second C2f-ELA module. The output of the second C2f-ELA module is denoted as feature F ce1 , the input of the second C2f-ELA module is connected to the input of the first CBS module. The input of the first CBS module and feature F ce2 are connected to the input of the third Concat module. The output of the third Concat module is connected to the input of the third C2f-ELA module. The output of the third C2f-ELA module is denoted as feature F ce21, the output of the third C2f-ELA module is connected to the input of the second CBS module, and the output of the second CBS module is connected to the feature F 3 and the input of the fourth Concat module. The output of the fourth Concat module is connected to the input of the fourth C2f-ELA module, and the output of the fourth C2f-ELA module is denoted as feature F ce3 , and finally the feature F ce1 , the feature F ce21 and the feature F ce3 are fed into the model prediction module;
[0076] Specifically for the model prediction part: The feature F obtained from the feature fusion part ce1 is input to the Detect1 module, the feature F ce21 is input to the Detect2 module, and the feature F ce3 is input to the Detect3 module.
[0077] Specifically for C2f-ELA: The input feature F is sent to the input of the first ELA module, and then split into two features denoted as feature F e1 and feature F e2 . The output feature F of the first ELA module e1 is sent to the first Bottleneck module, and the output of the first Bottleneck module is denoted as feature F b , and then the feature F b is split to obtain feature F b1 and feature F b2 . The feature F b1 is sent to the input of the second ELA module. The output of the second ELA module is connected to the input of the second Bottleneck module, and the output of the second Bottleneck module is denoted as feature F eb . Finally, the feature F b2 , the feature F e2 and the feature F eb are sent to the fifth Concat module. The output of the fifth Concat module is sent to the third CBS module, and the output feature of the third CBS module is denoted as feature F o .
[0078] Specifically for SPPF: The input feature F i is input to the fourth CBS module, and the output feature of the fourth CBS module is denoted as feature F i1 . The output of the fourth CBS module is connected to the input of the first max pooling layer module, and the output feature of the first max pooling layer module is denoted as feature F m1 . The output of the first max pooling layer module is connected to the input of the second max pooling layer module, and the output feature of the second max pooling layer module is denoted as feature F m2, the input of the second max-pooling layer module is connected to the output of the third max-pooling layer module, and the output features of the third max-pooling layer module are denoted as feature F m3 , feature F i1 、feature F m1 、feature F m2 and feature F m3 are connected to the input of the sixth Concat, and the output of the sixth Concat is connected to the output of the fifth CBS module.
[0079] Specifically, for CBS: the input features are sent to the input of the first convolutional layer, the output of the first convolutional layer is connected to the input of the batch normalization layer, and the output of the batch normalization layer is connected to the SiLU activation function.
[0080] Specifically, the EfficientFormerV2 network is as follows:
[0081] In the Stem module, the input features are connected to the second convolutional layer, the output of the second convolutional layer is connected to the second batch normalization layer, the output of the second batch normalization layer is connected to the first GeLU activation function, the first GeLU activation function is connected to the third convolutional layer, the output of the third convolutional layer is connected to the third batch normalization layer, and the output of the third batch normalization layer is connected to the second GeLU activation function;
[0082] In the local feature extraction module, the input features are connected to the input of the first feature extraction block, and the output of the first feature extraction block is connected to the input of the second feature extraction block; where the first feature extraction block and the second feature extraction block each contain a 1×1 convolutional layer, a feature extraction block batch normalization layer, a GeLU activation function, a 3×3 depthwise separable convolutional layer, a batch normalization layer, a GeLU activation function, a 1×1 convolution, a batch normalization layer, and a residual connection layer;
[0083] In the downsampling fusion module, the input features are connected to the input of the first local feature extraction module, the output of the first local feature extraction module is connected to the input vector of the first global feature extraction module, and the output of the first global feature extraction module is connected to the input vector of the second global feature extraction module; where the global feature extraction module contains a 3×3 convolutional layer and a downsampling layer with batch normalization, three 1×1 convolutional layers and batch normalization layers respectively for obtaining Query, Key, and Value for performing self-attention operations, an upsampling layer for restoring the dimension of the feature map, a 1×1 convolutional layer and a batch normalization layer, and a feature extraction block;
[0084] In the local-global feature extraction and fusion module, the input feature is connected to the input of the third local feature extraction module. The output of the third local feature extraction module is connected to the input of the fourth local feature extraction module. The output of the fourth local feature extraction module is connected to the input of the first local-global feature fusion block. The output of the first local-global feature fusion block is connected to the input of the fifth local-global feature fusion block;
[0085] In the local-global feature fusion block, the input feature is connected to the input of the first downsampling module Downsample. The output of the first downsampling module Downsample is connected to the input of the fourth convolutional layer. The output of the fourth convolutional layer is connected to the input of the fourth batch normalization layer. The output of the fourth batch normalization layer is connected to the input of the first Attention module. The output of the first Attention module is connected to the input of the first upsampling module Upsample. The output of the first upsampling module Upsample is connected to the input of the fifth convolutional layer. The output of the fifth convolutional layer is connected to the input of the fifth batch normalization layer. The output of the fifth batch normalization layer is connected to the input of the sixth convolutional layer. The output of the sixth convolutional layer is connected to the input of the sixth batch normalization layer. The output of the sixth batch normalization layer is connected to the input of the third GeLU activation function. The output of the third GeLU activation function is connected to the input of the first depthwise separable convolutional layer. The output of the first depthwise separable convolutional layer is connected to the input of the seventh batch normalization layer. The output of the seventh batch normalization layer is connected to the input of the fourth GeLU activation function. The output of the fourth GeLU activation function is connected to the input of the seventh convolutional layer. The input of the seventh convolutional layer is connected to the input of the eighth batch normalization layer.
[0086] In a further explanation, the high-precision real-time printed circuit board defect target detection model adopted by the present invention consists of Figure 2 It can be seen that the input end of the feature extraction part is used to input images. The output end of the feature extraction part is respectively connected to the feature fusion part and the input of the SPPF module. The output of the feature fusion part is connected to the detection head module in the model prediction part;
[0087] Figure 2 The feature extraction part in Figure 7 The structure shown is as follows: The input image first passes through a Stem module, and then respectively passes through a local feature extraction module ( Figure 7 Local in Figure 7 Subsample in Figure 7 Local-Global in Figure 7 Subsample2 in
[0088] The Stem module contains two convolution modules, each of which contains a 3×3 convolution layer, a batch normalization layer, and a GeLU activation function; a local feature extraction module contains two feature extraction blocks, each of which contains a 1×1 convolution layer and a batch normalization layer, a GeLU activation function, a 3×3 depth-separable convolution layer and a batch normalization layer, a GeLU activation function, a 1×1 convolution layer and a batch normalization layer, and a residual connection layer; a downsampling module contains a 3×3 convolution layer and a batch normalization layer; a local-global feature extraction fusion module contains two local feature extraction modules ( Figure 7 Local) and two global feature extraction modules, where the global feature extraction module contains a 3×3 convolution layer and a batch normalization downsampling layer, three 1×1 convolution layers and batch normalization layers are used to obtain Query, Key and Value for performing self-attention operations, an upsampling layer is used to restore the dimension of the feature map, a 1×1 convolution layer and batch normalization layer, and a feature extraction block; a downsampling fusion module contains a local feature extraction module, three 1×1 convolution layers and batch normalization layers are used to obtain Query, Key and Value for performing self-attention operations, a 3×3 convolution layer and batch normalization layer are used to obtain local features ( Figure 7 locality in), two 1×1 convolutional layers are used to transform the spatial dimension ( Figure 7 Talking Head in ) and a feature extraction block.
[0089] Figure 2 The structure of the feature fusion part in is as follows: It includes three feature maps of different sizes output by the EfficientFormerV2 model of the input image, denoted as F 1 、F 2 and F 3 (from front to back), F 3 First, a SPPF module is used to further extract features to obtain F s , then F s Enter the upsampling module Upsample to F u1 , then F u1 and F 2 Concatenate by channel dimension to get F c2 , then F c2 After a C2f-ELA module extracts features, we get F ce2 , subsequent F ce2 After the second upsampling module Upsample, we get F u2 , then F u2 With F 1 Concatenate by channel dimension to get Fc1 , F c1 Continue to extract features through a C2f-ELA module to obtain F ce1 , F ce1 Then, extract features through a CBS module to obtain F c1 , F c1 and F ce2 Concatenate them along the channel dimension to obtain F cc2 , F cc2 Extract features through a C2f-ELA module to obtain F ce2 , F ce2 Extract features through a CBS module to obtain F c2 , F c2 and F 3 Concatenate them along the channel dimension to obtain F c3 , F c3 Extract features through a C2f-ELA module to obtain F ce3 , and finally F ce1 、F ce2 and F ce3 Are sent into the model prediction module.
[0090] Figure 2 The model prediction part in includes the first detection head Detect, the second detection head Detect, and the third detection head Detect; among them, the first detection head, the second detection head, and the third detection head each contain two independent 3×3 convolutional layers and a 1×1 convolutional layer, which are used to predict the target coordinates and class information.
[0091] Figure 3 is the diagram of the CBS module. A CBS module contains a 3×3 convolutional layer, a batch normalization layer, and a SiLU activation function.
[0092] Figure 4 is the diagram of the C2f-ELA module. The input feature map F is first obtained after ELA attention to get F e , and then two copies are made to get F e1 and F e2 , and then F e1 is sent into the Bottleneck module to get F b , and then F b is copied twice to get F b1 and F b2 , and subsequently F b1 continues to be sent into ELA and Bottleneck to get F eb , and finally F e2 、F b2 and F eb are connected along the channel and then sent to a CBS convolutional block to smooth the features to obtain the output feature map Fo .
[0093] Figure 5 It is a Bottleneck module diagram, including two CBS modules and a residual connection block.
[0094] Figure 6 It is a SPPF module diagram, and the input feature map F i First passes through a 1×1 convolutional layer, a batch normalization layer and a SiLU activation function to obtain F i1 , and then F i1 is fed into a max pooling layer with a pooling factor of 5 to obtain F m1 , F m1 is fed into a max pooling layer with a pooling factor of 5 again to obtain F m2 , F m2 is fed into a max pooling layer with a pooling factor of 5 again to obtain F m3 , and finally F i , F m1 , F m2 and F m3 are connected by channels to obtain the output F o .
[0095] Embodiment:
[0096] The present invention as a whole includes two major stages, namely the model training stage and the model inference stage. In these two stages, the overall structure of the model remains unchanged, and what changes are the weight parameters of the entire model, which are optimized as the training progresses, and finally a relatively high detection accuracy is achieved. The model training stage and the model inference stage of the present invention are described in detail below.
[0097] Model training stage:
[0098] Step 1: Prepare a printed circuit board defect target detection dataset;
[0099] The printed circuit board defect dataset is a public dataset released by Peking University. This dataset annotates 6 types of defects (missing holes, rat bites, open circuits, short circuits, strays, and miscellaneous copper), and can be used for object detection, classification, and registration tasks. Among them, there are 693 images suitable for the detection task. In this paper, 569 images are randomly selected as the training set, and 124 images are used as the validation set. In order to enhance the diversity of the dataset, the present invention adopts data augmentation strategies, including rotation, flipping, cropping, and changing the brightness, contrast, and hue of the images. Through these methods, a sufficient number and diversity of printed circuit board defect samples can be accumulated to provide a data basis for model training. Figure 11 Shows a single instance picture of a printed circuit board.
[0100] Step 2: Construct a high-precision real-time printed circuit board defect target detection model;
[0101] To enhance the feature extraction ability of the YOLOv8s
[10] model for images, in the feature extraction stage of the present invention, EfficientFormerV2
[11] is used to replace the convolutional block structure of the original YOLOv8s model. In the feature fusion stage, the C2f-ELA module is proposed by combining ELA (Efficient Local Attention)
[12] to better highlight the target features. The overall framework of a PCB defect target detection model based on the improved YOLOv8s model designed by the present invention is as Figure 2 shown. Figure 2 The CBS module in Figure 3 is as Figure 2 shown. The Upsample in Figure 2 is an upsampling operator, whose function is to double the width and height dimensions of the input feature map, and the implementation method is upsampling interpolation; Concat is a splicing operation, which splices two feature maps with the same number of channels according to the dimension; Detect is the output feature map detection module, which is used to predict the category and coordinate information of the target. Figure 4 shown.
[0102] EfficientFormerV2 has made some modifications to both the structure and search algorithm of the EfficientFormer model. Specifically, in terms of the model structure, in Figure 7 (b) part, EfficientFormerV2 uses depthwise separable convolution to replace the original average pooling layer, and combining local information can effectively improve the integrity of target feature extraction. Compared with traditional convolution operations, depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution, showing the following significant advantages. First, the computational efficiency is significantly improved. Depthwise separable convolution greatly reduces the number of parameters and computational amount required for convolution operations. Second, the local feature extraction ability is enhanced. Depthwise convolution focuses on the spatial domain operation of the input features, so it can capture more fine-grained local patterns, while pointwise convolution effectively fuses the information between channels. This feature makes the model perform better when dealing with small targets in complex backgrounds. In addition, the lightweight design of the model provides convenience for its deployment in practical applications. Due to the significant reduction in the number of parameters, depthwise separable convolution can alleviate the overfitting problem, thereby improving the generalization performance of the model on small-scale datasets. Finally, the flexibility and scalability further enhance the adaptability of the network. Depthwise separable convolution can be seamlessly integrated into different network architectures, and replacing the pooling layer with it can improve both the local feature extraction ability and the global semantic expression ability. InFigure 7 (d) Part of EfficientFormerV2 proposes a Local-Global module that fuses local information and global features, fully leveraging the local information of the target extracted by convolution and the global features of the image focused by self-attention to better represent the target features. Additionally, in Figure 7 (e) Part of EfficientFormerV2 proposes an improved multi-head self-attention module. Local information is incorporated into the self-attention calculation matrix through 3×3 convolution. Additionally, a fully connected layer (labeled Talking Head in the figure) is added between different Heads in self-attention to enhance the information interaction between different Heads. This design significantly improves the model's performance while keeping the number of parameters and computational latency approximately unchanged. First, the fusion of local information enhances the model's ability to perceive fine-grained features. By introducing 3×3 depth convolution into the self-attention calculation matrix, the network captures local context relationships, compensating for the limitations of traditional multi-head self-attention mechanisms in processing spatial information. Second, the communication mechanism between different Heads in self-attention improves the global feature representation ability. Introducing a fully connected layer in the dimension of attention heads not only increases the information flow between different Heads but also breaks the bottleneck of independent information processing by traditional multi-head self-attention. In this way, the model can share and synthesize feature information between different attention heads, thus generating more expressive global representations. In the search algorithm, EfficientFormerV2 proposes a brand-new fine-grained joint search strategy to optimize module parameters and the number of channels to obtain the optimal module parameters and number of channels. This method not only further explores the potential of the search space but also balances depth and width in network design, showing the following significant advantages. First, it optimizes the configuration of network depth and width. By jointly searching the number of modules and channels in each stage, the efficiency of the model in resource-constrained scenarios is improved. Second, global optimality is achieved through fine-grained joint search. Different from traditional methods that only optimize a single dimension, the joint search strategy of EfficientFormerV2 can explore the design space from a global perspective, avoiding the problem of falling into local optimality due to optimizing a single metric.
[0103] The present invention introduces EfficientFormerV2 into the feature extraction part of YOLOv8s, achieving performance optimization for the task of detecting defects on printed circuit boards. By replacing the original convolutional blocks, the present invention effectively enhances the model's ability to extract features from input images, especially showing excellent performance in small object detection tasks. EfficientFormerV2 combines the global feature capture ability of ViT and the local feature extraction efficiency of CNN, and can more comprehensively capture the global dependencies of various defects in printed circuit board images during the feature extraction stage, while characterizing the feature information of small objects (such as micro cracks and defect points) with higher resolution and finer granularity. For the detection scenario of printed circuit boards with complex structures and diverse defect morphologies, the global-local joint representation ability of EfficientFormerV2 can significantly improve the recognition accuracy of the model for defect targets. The above advantages enable the improved model to be applied to the real-time quality control link in electronic manufacturing, helping to quickly locate and repair defects on circuit boards, and improving production efficiency and product quality.
[0104] The structure of the Efficient Local Attention (ELA) is as Figure 8 shown. Given an input feature map F with width, height, and number of channels being H, W, and C respectively, average pooling operations are performed along the horizontal and vertical directions on each channel to obtain F h and F w , and the specific operations are shown in the following formula.
[0105]
[0106] where H and W are the width and height of the input feature map respectively.
[0107] Subsequently, F h and F w are respectively passed through a one-dimensional convolution, group normalization operation, and Sigmoid function to obtain the attention weights at two positions, namely F hw and F ww , as shown in the following formula. The final output result F o is the product of the original input feature map F and the two attention weights F hw , F ww at different positions, as shown in the following formula.
[0108] F hw =σ(G n D h (F h ))
[0109] F ww =σ(G n D w (F w ))
[0110] F o = F × F hw × F ww
[0111] where σ is the Sigmoid function, D h and D w are convolutional layers with 1, and their convolutional kernels are set to 5 and 7 respectively, G n is the group normalization operation.
[0112] The design of the ELA module draws on the pooling strategy of corner attention and introduces a series of optimizations to improve efficiency and performance. Its structure is mainly divided into the following steps. First, obtain the directional features, perform pooling in the horizontal and vertical directions respectively in the spatial dimension, and extract the directional feature vectors containing long-range dependencies. This method effectively retains the rich information of the target position and avoids the interference of irrelevant regions on label prediction. Then, use one-dimensional convolution for local interaction, and use lightweight one-dimensional convolution to perform local interaction on the feature vectors. By adjusting the size of the convolutional kernel, the local interaction range can be flexibly represented, thereby capturing appropriate spatial dependencies. Then, use normalization and activation, and use group normalization to replace batch normalization to improve the stability and generalization ability of small-batch training. Subsequently, combine the non-linear activation function to generate the directional attention prediction. Finally, it is the fusion of attention at different positions. After obtaining the position attention in the horizontal and vertical directions respectively, combine them through a product operation to generate the final global position attention. The ELA attention enhances the spatial information modeling ability. By independently processing the feature vectors in two different directions, the ELA can accurately capture the positions of the regions of interest while retaining long-range dependencies. The ELA design is lightweight enough to use one-dimensional convolution to replace the traditional two-dimensional convolution, greatly reducing the computational overhead and optimizing the computational speed, making it more suitable for deployment on resource-constrained devices. Through the streamlined structure design and lightweight operations, the ELA balances performance and efficiency, providing a better attention mechanism for deep convolutional neural networks.
[0113] Based on the ELA, the present invention proposes a brand-new C2f-ELA module, and the overall structure is as Figure 4 shown. The input feature map F first passes through the ELA attention to obtain F e , and then two copies are obtained to get F e1 and F e2 , and then F e1 is sent into the Bottleneck to obtain F b , where the Bottleneck consists of two 3×3 convolutional layers and a residual connection layer. Then, F b is copied twice to get F b1 and F b2 , and subsequently Fb1 Continue to send it into ELA and Bottleneck to obtain F eb , and finally send F e2 , F b2 and F eb Connect them by channels and then send them to a CBS convolutional block to smooth the features to obtain the output feature map F o .
[0114] The C2f-ELA module proposed by the present invention demonstrates various advantages in the task of detecting defective targets on printed circuit boards, which are mainly reflected in the following aspects. First, it has accurate target positioning ability. The C2f-ELA module strengthens the modeling ability of the target position through the ELA mechanism. The ELA adopts a directional feature separation strategy to independently process the feature vectors in the horizontal and vertical directions, and then generates a global position attention map through attention fusion. This design enables the C2f-ELA module to accurately locate the defective area, and can effectively improve the detection accuracy even when the target size is small or the background is complex. Second, it has enhanced feature representation ability. The interaction and fusion of features in the multi-level path enable the module to simultaneously focus on local details and global structures, improving the recognition ability of complex defective targets. The design of the C2f-ELA module is highly optimized, significantly reducing the computational overhead, and this advantage is particularly suitable for deployment in resource-constrained devices or scenarios with high real-time requirements. In addition, the defective targets in printed circuit boards have diverse forms. The C2f-ELA module combines the directional attention of ELA and the feature extraction ability of Bottleneck, and can capture the significant features of different defects, adapting to various detection requirements, whether it is large-scale targets or tiny structural abnormalities.
[0115] Step 3: Model training framework and device;
[0116] In the model training stage of the present invention, the PyTorch framework is used to train the model, and a server equipped with an NVIDIA GeForce RTX 3090 graphics card is used to train the detection model. The central processing unit of the server is an Inter i9-12900HK, and the memory is 32G.
[0117] Step 4: Experimental index results;
[0118] Figure 9Shows the experimental result index diagram of the YOLOv8s model on the printed circuit board defect dataset. It can be seen from the figure that YOLOv8s trains stably on the electronic circuit board defect dataset, and the final object loss can converge to a relatively small value. During the training phase, the object bounding box loss, DFL loss (Distribution Focal Loss), and object classification loss continuously decrease. After 100 rounds of training, they decrease to 1.08, 0.919, and 0.822 respectively. During the validation phase, the object bounding box loss, DFL loss, and object classification loss of the YOLOv8s model decrease to 1.65, 1.07, and 0.953 respectively after 100 rounds of iteration. In terms of performance metrics, with the optimization of the model, at the 100th round, its Precision, Recall, mAP@0.5, and mAP@0.5:0.95 reach 92.3%, 90.9%, 94.4%, and 50.8% respectively.
[0119] Figure 10 Shows the experimental result index diagram of a high-precision real-time printed circuit board defect target detection model proposed by the present invention on the printed circuit board defect dataset. It can be seen from the figure that after 100 rounds of iterative optimization during the training phase of the model, the object bounding box loss, DFL loss, and object classification loss show a stable and rapid downward trend, and finally converge to 0.9, 0.719, and 0.682 respectively. During the validation phase, these three losses decrease to 1.419, 0.873, and 0.787 respectively. In terms of specific metrics, the method in this paper reaches 97.3%, 93.9%, 97.1%, and 55.4% in Precision, Recall, mAP@0.5, and mAP@0.5:0.95 respectively. Compared with the YOLOv8s model, it is improved by 5.0%, 3.0%, 2.7%, and 4.6% respectively.
[0120] Model inference phase:
[0121] Step 1: Image, video or camera input
[0122] The present invention accepts image, video or camera input during the inference phase. When video and camera are used as input, the present invention extracts the video or camera into frame-by-frame pictures according to the frame rate, and then sends them frame by frame to the next stage.
[0123] Step 2: Construct a high-precision real-time printed circuit board defect target detection model
[0124] In the first step of the present invention, frame-by-frame input images are obtained, and then the pictures are sent into a high-precision real-time printed circuit board defect target detection model for detection to obtain the output results, that is, the target category and bounding box coordinate information predicted for each pixel point on the output feature map.
[0125] Step 3: Detection Post-Processing and Result Output
[0126] In the second step of the present invention, the prediction results obtained from the input image through constructing a high-precision real-time printed circuit board defect target detection model are obtained. Then, in this step, these detection results will first be mapped back from the output image to the original image, that is, a coordinate mapping transformation is performed, and the mapping factor is obtained by dividing the size of the output feature map by the size of the original input image. After that, the present invention first filters out invalid detection frames according to the detection confidence score (default set to 0.25), and then uses the non-maximum suppression post-processing algorithm to filter out overlapping target frames, and finally obtains the final prediction results obtained from the input image through a certain method. Figure 11 is an example image of a printed circuit board, Figure 12 is the detection result image of the method proposed by the present invention on the printed circuit board image.
Claims
1. A high-precision real-time printed circuit board defect target detection method, characterized in that: The following steps are involved: Step 1: Collect PCB defect images from the Internet and integrate them with open source datasets, and filter and organize the images; Step 2: Further process the collected images, select suitable images, and use Gaussian filtering and other methods to process the selected images to obtain a batch of high-quality printed circuit board defect images; Step 3: Use the Labelimg tool to label the dataset, obtain a txt file, and divide it into training set, validation set, and test set according to a certain ratio; Step 4: Build a high-precision real-time printed circuit board defect target detection model; Step 5: Put the constructed data set into the high-precision real-time printed circuit board defect target detection model, train and obtain the trained model, and use precision, recall rate and the average precision of all categories as evaluation indicators; Step 6: Use the trained high-precision real-time printed circuit board defect target detection model to perform image prediction and accurate identification on the inspection image.
2. The method according to claim 1, characterized in that In step 1, the following steps are included: Step 1-1: Search the keyword "printed circuit board defects" through the Internet engine to obtain images and download them; Step 1-2: Find the open source printed circuit board defect image dataset on the open source data website and download the corresponding label file.
3. The method according to claim 1, characterized in that In step 2, specifically: the collected printed circuit board defect images are first manually screened to screen out images with low resolution and quality, and then preprocessed on the Pycharm platform to obtain a batch of high-quality printed circuit board defect images through Gaussian filtering method.
4. The method according to claim 1, characterized in that In step 3, the following steps are included: Step 3-1: Use Labelimg image annotation tool to annotate the collected defect images of printed circuit boards. The defects are divided into six categories: missing hole, mouse bite, open circuit, short circuit, spurious, and spurious copper. The missing hole is "missing_hole", the mouse bite is "mouse_bite", the open circuit is "open_circuit", the short circuit is "short", the spurious is "spur", and the spurious copper is "spurious_copper". Get the annotated txt file; Step 3-2: After completing the labeling work, divide the data set according to a certain ratio to obtain the training set, validation set and test set.
5. The method according to claim 1, characterized in that In step 4, a high-precision printed circuit board defect target detection model is constructed, which includes a feature extraction part, a feature fusion part, and a model prediction part; The input end of the feature extraction part is used to input an image, the output end of the feature extraction part is connected to the input of the feature fusion part and the SPPF module respectively, and the output of the feature fusion part is connected to the detection head module in the model prediction part; The feature extraction part adopts / uses EfficientFormerV2; The feature extraction part is as follows: The input image is input to the Stem module, the output of the Stem module is connected to the input of the first local feature extraction module, the output of the first local feature extraction module is connected to the input of the first downsampling fusion module, the output of the first downsampling fusion module is connected to the input of the second local feature extraction module, the output of the second local feature extraction module is recorded as feature F1, the output of the second local feature extraction module is connected to the input of the second downsampling fusion module, the output of the second downsampling fusion module is connected to the input of the first local-global feature extraction and fusion module, the output of the first local-global feature extraction and fusion module is recorded as feature F2, the output of the first local-global feature extraction and fusion module is connected to the input of the third downsampling fusion module, the output of the third downsampling fusion module is connected to the input of the second local-global feature extraction and fusion module, and the output of the second local-global feature extraction and fusion module is recorded as feature F3; The feature fusion part is as follows: feature F3 is input to the SPPF module, the output of the SPPF module is connected to the input of the first upsampling module, the output of the first upsampling module and feature F2 are connected to the input of the first Concat module, the output of the first Concat module is connected to the input of the first C2f-ELA module, and the output of the first C2f-ELA module is recorded as feature F ce2 , the output of the first C2f-ELA module is connected to the input of the second upsampling module, the output of the second upsampling module and the feature F1 are connected to the input of the second Concat module, the output of the second Concat module is connected to the input of the second C2f-ELA module, and the output of the second C2f-ELA module is recorded as feature F ce1 The second C2f-ELA module input is connected to the first CBS module input, and the first CBS module input and feature F ce2 The output of the third Concat module is connected to the input of the third C2f-ELA module, and the output of the third C2f-ELA module is recorded as feature F. ce21 , the output of the third C2f-ELA module is connected to the input of the second CBS module, the output of the second CBS module and the feature F3 are connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the fourth C2f-ELA module, and the output of the fourth C2f-ELA module is recorded as feature F ce3 , and finally feature F ce1 , Feature F ce21 and feature F ce3 Send it to the model prediction module; The model prediction part is specifically as follows: the feature F obtained by the feature fusion part ce1 Input to Detect1 module, feature F ce21 Input to Detect2 module, feature F ce3 Input to the Detect3 module.
6. The method according to claim 5, characterized in that The specific steps of C2f-ELA are as follows: the input feature F is sent to the input of the first ELA module, and then split into two features recorded as feature F e1 and feature F e2 , the output feature F of the first ELA module e1 And sent to the first Bottleneck module, the output of the first Bottleneck module is recorded as feature F b , then split feature F b Get feature F b1 and feature F b2 , feature F b1 The output of the second ELA module is connected to the input of the second Bottleneck module, and the output of the second Bottleneck module is recorded as feature F eb Finally, the feature F b2 , Feature F e2 and feature F eb It is sent to the fifth Concat module, and the output of the fifth Concat module is sent to the third CBS module. The output feature of the third CBS module is recorded as feature F o .
7. The method according to claim 5, characterized in that The specifics of SPPF are: Input feature F i Input to the fourth CBS module, the output feature of the fourth CBS module is recorded as feature F i1 The output of the fourth CBS module is connected to the input of the first maximum pooling layer module, and the output feature of the first maximum pooling layer module is recorded as feature F m1 , the output of the first maximum pooling layer module is connected to the input of the second maximum pooling layer module, and the output feature of the second maximum pooling layer module is recorded as feature F m2 , the input of the second maximum pooling layer module is connected to the output of the third maximum pooling layer module, and the output feature of the third maximum pooling layer module is recorded as feature F m3 , feature F i1 , Feature F m1 , Feature F m2 and feature F m3 Connected to the input of the sixth Concat, and the output of the sixth Concat is connected to the output of the fifth CBS module.
8. The method according to claim 5, characterized in that The specific details of CBS are as follows: the input features are sent to the input of the first convolutional layer, the output of the first convolutional layer is connected to the input of the batch normalization layer, and the output of the batch normalization layer is connected to the activation function SiLU.
9. The method according to claim 5, characterized in that The EfficientFormerV2 network is specifically: In the Stem module, the input features are connected to the second convolutional layer, the output of the second convolutional layer is connected to the second batch normalization layer, the output of the second batch normalization layer is connected to the first GeLU activation function, the first GeLU activation function is connected to the third convolutional layer, the output of the third convolutional layer is connected to the third batch normalization layer, and the output of the third batch normalization layer is connected to the second GeLU activation function; In the local feature extraction module, the input feature is connected to the input of the first feature extraction block, and the output of the first feature extraction block is connected to the input of the second feature extraction block; wherein the first feature extraction block and the second feature extraction block include a 1×1 convolution layer, a feature extraction block batch normalization layer, a GeLU activation function, a 3×3 depth-separable convolution layer, a batch normalization layer, a GeLU activation function, a 1×1 convolution, a batch normalization layer and a residual connection layer; In the downsampling fusion module, the input feature is connected to the input of the first local feature extraction module, the output of the first local feature extraction module is connected to the input vector of the first global feature extraction module, and the output of the first global feature extraction module is connected to the input vector of the second global feature extraction module; wherein the global feature extraction module includes a 3×3 convolution layer and a batch normalization downsampling layer, three 1×1 convolution layers and batch normalization layers for obtaining Query, Key and Value respectively for performing self-attention operations, an upsampling layer for restoring the dimension of the feature map, a 1×1 convolution layer and batch normalization layer, and a feature extraction block; In the local-global feature extraction and fusion module, the input feature is connected to the input of the third local feature extraction module, the output of the third local feature extraction module is connected to the input of the fourth local feature extraction module, the output of the fourth local feature extraction module is connected to the input of the first local-global feature fusion block, and the output of the first local-global feature fusion block is connected to the input of the fifth local-global feature fusion block; In the local-global feature fusion block, the input feature is connected to the first downsampling module Downsample input, the first downsampling module Downsample output is connected to the fourth convolutional layer input, the fourth convolutional layer output is connected to the fourth batch normalization layer input, the fourth batch normalization layer output is connected to the first Attention module input, the first Attention module output is connected to the first upsampling module Upsample input, the first upsampling module Upsample output is connected to the fifth convolutional layer input, the fifth convolutional layer output is connected to the fifth batch normalization layer input, the fifth batch normalization layer output is connected to the sixth convolutional layer input, the sixth convolutional layer output is connected to the sixth batch normalization layer input, the sixth batch normalization layer output is connected to the third GeLU activation function input, the third GeLU activation function output is connected to the first depthwise separable convolutional layer input, the first depthwise separable convolutional layer output is connected to the seventh batch normalization layer input, the seventh batch normalization layer output is connected to the fourth GeLU activation function input, the fourth GeLU activation function output is connected to the seventh convolutional layer input, and the seventh convolutional layer input is connected to the eighth batch normalization layer input.
Citation Information
Cited By
Printing defect detection method and equipment based on computer vision and medium
CN122391170A