A method for detecting surface defects of printed circuit boards based on deep learning

By improving the RT-DETR network algorithm, the DySample and RetBlockC3 modules were introduced, and the DA-ACFN module was integrated, which solved the problem of detecting small defects on the surface of the printed circuit board, and efficient and accurate defect recognition was achieved, improving detection accuracy and robustness.

CN120031872BActive Publication Date: 2025-07-18HENAN INST OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510496935.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately detect various minor defects on the surface of printed circuit boards, especially due to the complex manufacturing process and the small defect size, which affects product quality and safety.

Method used

The improved RT-DETR network algorithm is adopted, and dynamic upsampling is introduced by introducing the DySample module, replacing it with the RetBlockC3 module, and integrating the DA-ACFN module to optimize feature extraction and fusion to improve the efficiency and accuracy of the model in multi-scale target processing.

Benefits of technology

It improves the target detection accuracy and robustness of printed circuit board defect detection, can more effectively identify multiple defects, reduce calculation complexity, and enhance the feature extraction ability of key areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031872B_ABST
    Figure CN120031872B_ABST
Patent Text Reader

Abstract

A method for detecting surface defects of printed circuit boards based on deep learning, including the RT-DETR network algorithm. The DySample module is introduced in the CCFM sampling stage of the RT-DETR network algorithm, and dynamic upsampling is performed in the CCFM sampling stage; the RetBlockC3 module with the RetBlock structure is used to replace the original RepC3 module in the RT-DETR network algorithm; the DA-ACFN module is integrated into the RT-DETR network to obtain the improved RT-DETR network; the improved RT-DETR network algorithm is trained using the printed circuit board surface defect dataset to obtain the RT-DETR object detection model; the RT-DETR object detection model is used for detecting surface defects of printed circuit boards; thereby improving the object detection accuracy and robustness of the RT-DETR network in printed circuit board defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical image processing, and particularly to a method for detecting surface defects of printed circuit boards based on deep learning. Background Art

[0002] The detection of surface defects of printed circuit boards (PCBs) plays a crucial role in quality control. With the rapid development of computer vision (CV) technology, traditional image processing methods (such as thresholding, edge detection, region-based methods) as well as modern machine learning and deep learning technologies are gradually replacing the inefficient manual detection methods. Among them, optical image processing technology has become the mainstream method for PCB surface defect detection. However, with the rapid development of the electronic information technology industry and the increasing market demand for PCB products, the manufacturing of PCBs is moving towards high reliability, high density, and miniaturization. These trends not only expand the functions of PCBs, increase the circuit density, but also reduce the size, which brings many challenges to the machine vision-based PCB surface defect detection, mainly reflected in the following two aspects:

[0003] 1. Multiple defects caused by complex manufacturing processes: The manufacturing process of PCBs is complex and prone to defects such as open circuits, short circuits, and burrs. These tiny defects may seriously affect the product quality and even cause serious safety hazards, such as an increased risk of electric leakage and even fire, endangering the lives of users. Therefore, it is required that the detection model has the ability to efficiently and accurately identify multiple defects.

[0004] 2. High difficulty in detecting tiny defects: Due to the influence of different manufacturing processes, the defects are usually small in area and tiny in size. For example, the size of PCB surface defects is usually less than 4500 pixels, and the size of burrs is usually less than 300 pixels, which only accounts for 0.005% - 0.07% of a high-resolution PCB image (about 6.5 million pixels). Such a tiny proportion makes it extremely challenging to quickly and accurately detect these defects. Summary of the Invention

[0005] Aiming at the problems in the prior art, the present invention provides a method for detecting surface defects of printed circuit boards based on deep learning, aiming to achieve fast, accurate, and high-speed detection of surface defects of printed circuit boards.

[0006] A method for detecting surface defects of printed circuit boards based on deep learning includes the RT-DETR network algorithm. The deep learning method includes the following steps:

[0007] Step 1: Improve the RT-DETR network algorithm, specifically:

[0008] Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm, and perform dynamic upsampling in the CCFM sampling stage;

[0009] Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that adopts the RetBlock structure;

[0010] Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain the improved RT-DETR network;

[0011] Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR object detection model;

[0012] Step 3: Apply the RT-DETR object detection model to the detection of printed circuit board surface defects.

[0013] Furthermore: The DA-ACFN module includes the ACFN module and the DASI module. The ACFN is optimized by the DASI module to form the improved DA-ACFN module.

[0014] Furthermore: The ACFN module includes a feature splicing module. The first input end of the feature splicing module is connected to the downsampling module and is used to input low-level features. The second input of the feature splicing module is connected to the convolutional module and is used to input middle-level features. The third input of the feature splicing module is connected to the convolutional module and is used to input middle-level features. Several outputs of the feature splicing module are all passed through the depthwise separable convolutional module and then sequentially through the skip connection module and the 1×1 convolutional module for information integration to obtain the integrated information. The integrated information is skip-connected with one output of the feature splicing module to obtain the fused feature map;

[0015] The process of the DASI module optimizing the ACFN module is as follows:

[0016] Step 1.3.1: Receive the fused feature map, and assume the dimension of the fused feature map is C × H × W;

[0017] Step 1.3.2: Divide the fused feature map into four groups of features, namely low-level features, high-level features, background features, and the feature selection part;

[0018] Step 1.3.3: Adopt the Group mechanism to divide the channels, and apply the respective channel attention strategies to the features of different groups to obtain the transformed features of the corresponding dimensions;

[0019] Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function;

[0020] Step 1.3.5: The weighted features are recombined, and after per-channel multiplication (×) and weighted addition (+), channel concatenation is performed to obtain the finally optimized feature representation, completing the optimization of the ACFN module.

[0021] Furthermore, in the improved RT-DETR network, after the input image undergoes preliminary feature extraction through ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features. One output of the last BasicBlock layer is directly fed into the first feature fusion module in the backbone feature extraction part, and the other output is sequentially fed into the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives the feature layers output by other BasicBlocks layers except the last BasicBlock layer. One output of the first DA-ACFN module is fed into the first feature fusion module after passing through the third convolutional layer, and the other output is fed into the Dysample layer in the feature fusion part. The CCFM layer in the feature fusion part is replaced by the second DA-ACFN module.

[0022] Furthermore: Combining Figure 2 As shown, the dynamic upsampling in Step 1.1 specifically includes the following steps:

[0023] Step 1.1.1: Generate an offset of size 2S²×H×W from the input feature map of size C×H×W through a linear projection layer.

[0024] Step 1.1.2: Convert the offset to an appropriate spatial resolution, and finally calculate the final coordinates of the sampling points by adding the offset to the original grid point coordinates.

[0025] Step 1.1.3: Resample the feature map using the generated sampling points to obtain the upsampled feature map.

[0026] Furthermore, the upsampling process of the DySample module is explained by the following formula:

[0027] (1)

[0028] In this equation, 𝑥 represents the input feature map, 𝑥′ represents the output feature map, and 𝑔 is the original sampling grid; bilinear(x) performs fixed multiple scaling on the input feature map by means of bilinear interpolation, interpolating through the weighted average of adjacent pixels; pixel_shuffle() is used as a pixel rearrangement operation, and by rearranging the pixels in the channel dimension to the spatial dimension, the enlarged feature map is evenly distributed in space; reshaping_sample() is used to reshape the sampled feature map.

[0029] Furthermore: DySample includes a static range factor sampling strategy and a dynamic range factor sampling strategy; among them, the static range factor sampling strategy performs bilinear interpolation through a 0.25-fold scaling and then applies pixel shuffle for pixel rearrangement; the dynamic range factor sampling strategy performs bilinear interpolation with a 0.5-fold scaling and combines the dynamically calculated offset for compensation;

[0030] Furthermore: In step 1.2: The RetBlockC3 module includes depthwise separable convolution, layer normalization, Manhattan self-attention mechanism, and a feed-forward neural network. The input features are sequentially sent to the Manhattan self-attention mechanism after depthwise separable convolution and layer normalization. The Manhattan self-attention mechanism calculates the Manhattan distance between different positions in the feature map and assigns different weights to the features at different positions according to the spatial attenuation principle. The output of the Manhattan self-attention mechanism is normalized again through layer normalization, and the normalized features enter the feed-forward neural network for feature transformation and enhancement.

[0031] Furthermore: The output of the RetBlockC3 module is:

[0032] (2)

[0033] Among them, DWConv is the depthwise separable convolution operation of the context local enhancement module, and respectively calculate the attention scores from the vertical and horizontal directions;

[0034] (3)

[0035] (4)

[0036] Among them, is the Manhattan self-attention calculation in the horizontal direction, that is, the attention score calculated along the horizontal direction; is the Manhattan self-attention calculation in the vertical direction, that is, the attention score calculated along the vertical direction; It is a query, key, and value matrix in the horizontal direction; is a normalization function that maps numerical values to the range (0, 1) and is used to calculate attention weights. is a scaling factor. and are the horizontal and vertical decay matrices, respectively; and are the Manhattan distances between two tokens n and m horizontally and vertically, respectively;

[0037] (5)

[0038] (6)

[0039] Among them, γ is a decay coefficient, Xn and Xm are the horizontal coordinates of the nth and mth tokens in the image, respectively, and Yn and Ym are the vertical coordinates of the nth and mth tokens in the image, respectively.

[0040] Advantages of the present invention: By improving the RT-DETR network through the DySample, RetBlockC3, and DA-ACFN modules, the computational complexity of the RT-DETR network is reduced, and it can more effectively focus on important regions in the image to enhance the feature extraction ability of the model, thereby improving the efficiency and accuracy of the network in multi-scale target processing, and is used to improve the object detection accuracy and robustness in printed circuit board defect detection. Brief Description of the Drawings

[0041] Figure 1 is the flow block diagram of the invention;

[0042] Figure 2 is the structural schematic diagram of the DySample module in the present invention;

[0043] Figure 3 is the structural schematic diagram of the RetBlock module applied in the RT-DETR network algorithm;

[0044] Figure 4 is the structural schematic diagram of the DA-ACFN module;

[0045] Figure 5 is the structural schematic diagram of the DASI module;

[0046] Figure 6 is the structural schematic diagram of the RT-DETR;

[0047] Figure 7 is the structural schematic diagram of the visual detection automation device. Detailed Embodiments

[0048] The present invention will be described in detail below with reference to the accompanying drawings. The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention. The terms of orientation such as left, middle, right, up, and down in the embodiments of the present invention are only relative concepts to each other or are referenced based on the normal use state of the product and should not be considered restrictive.

[0049] A method for detecting surface defects of a printed circuit board based on deep learning, as Figure 1 and Figure 6 shown, includes the RT-DETR network algorithm. The deep learning method includes the following steps:

[0050] Step 1: Improve the RT-DETR network algorithm, specifically:

[0051] Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm and perform dynamic upsampling in the CCFM sampling stage; this can reduce the computational complexity while retaining rich semantic information and improving the detection accuracy of the model;

[0052] Combined with Figure 2 shown, the dynamic upsampling in Step 1.1 specifically includes the following steps:

[0053] Step 1.1.1: Generate an offset of size 2S²×H×W from the input feature map of size C×H×W through a linear projection layer;

[0054] Step 1.1.2: Convert the offset to an adapted spatial resolution, and finally calculate the final coordinates of the sampling points by adding the offset to the original grid point coordinates;

[0055] Step 1.1.3: Resample the feature map using the generated sampling points to obtain the upsampled feature map; this can reduce the computational complexity while retaining rich semantic information and improving the detection accuracy of the model;

[0056] The upsampling process of the DySample module is explained by the following formula:

[0057] (1)

[0058] In this equation, 𝑥 represents the input feature map, 𝑥′ represents the output feature map, and 𝑔 is the original sampling grid; bilinear(x) performs fixed-magnification scaling (usually 2× or 4×) on the input feature map using bilinear interpolation. By interpolating the weighted average of adjacent pixels, it improves the smoothness and continuity of the feature map, thereby reducing artifacts that may occur during the upsampling process; pixel_shuffle() is a pixel rearrangement operation that rearranges the pixels in the channel dimension to the spatial dimension, making the enlarged feature map more evenly distributed in space, thus improving the sampling accuracy. For example, in the case of 2× magnification, a 2×2 sub-block in one channel will be remapped to the spatial dimension to generate a new high-resolution feature map; reshaping_sample() is used to reshape the sampled feature map so that its shape fits the subsequent calculation structure of the network, ensuring that the compensated features can be correctly passed to the subsequent network layers to maintain the consistency and stability of the feature information;

[0059] In addition, DySample includes a Static Scope Factor sampling strategy and a Dynamic Scope Factor sampling strategy; among them, the Static Scope Factor sampling strategy performs bilinear interpolation with a 0.25-fold scaling and then applies pixel shuffle for pixel rearrangement, enabling the features to be evenly expanded in the spatial dimension; the Dynamic Scope Factor sampling strategy performs bilinear interpolation with a 0.5-fold scaling and combines it with a dynamically calculated offset for compensation, making it more adaptable during the sampling process, so that the distribution of feature points can be adaptively adjusted according to the input features; in Figure 2 it, s represents the sampling scale factor, which controls the distribution range of the sampling points; sH and sW respectively represent the dynamic scaling coefficients in the height (H) and width (W) directions, which are used to adjust the offsets of the sampling points in different directions; 0.5 and 0.25 respectively correspond to the scaling ratios of the Dynamic Scope Factor and the Static Scope Factor, determining the magnification degree of the feature map; g represents the sampling point grid, which determines the specific sampling positions during the upsampling process to ensure the accuracy and consistency of feature extraction;

[0060] Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that adopts the RetBlock structure; as Figure 3As shown in the figure, the RetBlockC3 module includes depthwise separable convolution (DWConv), layer normalization (LN), Manhattan self-attention mechanism (MaSA), and feed-forward neural network (FFN). The input features are sequentially fed into the depthwise separable convolution (DWConv) and layer normalization (LN), and then sent to the Manhattan self-attention mechanism (MaSA). The Manhattan self-attention mechanism (MaSA) calculates the Manhattan distance between different positions in the feature map and assigns different weights to the features at different positions according to the spatial attenuation principle. The output of the Manhattan self-attention mechanism (MaSA) is normalized again through layer normalization (LN) to ensure the stability and consistency of the features during transmission. The normalized features enter the feed-forward neural network (FFN) for feature transformation and enhancement to improve the expression ability of the target features, and the optimized features are fused into the backbone network to provide a more discriminative feature representation for subsequent object detection. Among them, the depthwise separable convolution (DWConv) can enhance the feature extraction ability while reducing the computational complexity, enabling the network to focus more precisely on the key regions and reducing the interference of redundant information. Layer normalization (LN) can stabilize the feature distribution and improve the convergence speed of the training process. The Manhattan self-attention mechanism (MaSA) enables the model to focus more on the interaction of local information and effectively weakens the interference of the background region on the detection task. It can be seen that RetBlockC3 combines DWConv, LN, the Manhattan self-attention mechanism (MaSA), and FFN, not only improving the feature expression ability while reducing the computational complexity, but also enhancing the network's attention to local regions, making the model have higher accuracy and robustness in multi-scale feature fusion and object detection tasks;

[0061] The output of the RetBlockC3 module is:

[0062] (2)

[0063] Among them, DWConv is the depthwise separable convolution operation of the context local enhancement module, and calculate the attention scores in the vertical and horizontal directions respectively;

[0064] (3)

[0065] (4)

[0066] Among them, is the Manhattan self-attention calculation in the horizontal direction, that is, the attention score calculated along the horizontal direction (H direction); is the Manhattan self-attention calculation in the vertical direction, that is, the attention score calculated along the vertical direction (W direction); are the Query, Key, and Value matrices in the horizontal direction (H); is the normalization function that maps values to the range (0, 1) and is used to calculate the attention weights, is the scaling factor, 、 are the decay matrices in the horizontal (H direction) and vertical (W direction) respectively; and are the Manhattan distances between two tokens n and m in the horizontal (H direction) and vertical (W direction) respectively;

[0067] (5)

[0068] (6)

[0069] where γ is a decay coefficient, Xn and Xm are the horizontal coordinates of the nth and mth tokens in the image respectively, and Yn and Ym are the vertical coordinates of the nth and mth tokens in the image respectively;

[0070] Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain the improved RT-DETR network; where the DA-ACFN module includes the ACFN module (Adaptive Contextual Fusion Network) and the DASI module (Dimension-Aware Selective Integration Module), and the ACFN module is optimized through the DASI module (Dimension-Aware Selective Integration Module) to form the improved DA-ACFN module;

[0071] The ACFN module (Adaptive Contextual Fusion Network) aims to fully exploit the information correlation of features at different scales to improve the overall performance of the detection model, such as Figure 4As shown in the figure, the ACFN module includes a feature concatenation module (Concat). The first input end of the feature concatenation module is connected to the downsampling module (ADown) and is used to input low-level features (P3). Low-level features have rich detailed information but are vulnerable to noise interference. The second input of the feature concatenation module is connected to the convolutional module (Conv) and is used to input middle-level features (P4). Middle-level features are used to achieve a balance between semantic expression and detail retention. The third input of the feature concatenation module is connected to the convolutional module (Conv) and is used to input middle-level features (P5). High-level features have strong semantic information but lack fine object details. These features are extracted by convolutional networks at different levels to form a complete multi-scale feature representation. Secondly, in the multi-scale feature alignment stage, since P3, P4, and P5 have different spatial sizes, downsampling (ADown) and convolution (Conv) are required for alignment to ensure that features at different scales can be effectively fused in the same scale space. After several outputs of the feature concatenation module pass through the depthwise separable convolution module (DWConv), further feature refinement can be performed to reduce the computational complexity and enhance the local information interaction ability. Then, through the skip connection module (ADD) and the 1×1 convolutional module (Conv) in sequence for information integration to obtain integrated information, and a skip connection is made between the integrated information and one output of the feature concatenation module to obtain a fused feature map;

[0072] As Figure 5 shown, the process of the DASI module optimizing the ACFN module (Adaptive Context Fusion Network) is as follows:

[0073] Step 1.3.1: Receive the fused feature map, and assume the dimension of the fused feature map is C × H × W;

[0074] Step 1.3.2: Divide the fused feature map into four groups of features (Group = 4), namely low-level features (Cl), high-level features (Ch), background features (CBS), and feature selection part (CsS). This division method enables the DASI module to accurately identify information at different scales and prevent high-level features from suppressing low-level features during the fusion process;

[0075] Step 1.3.3: Use the Group mechanism to divide the channels and apply respective channel attention strategies to the features of different groups. That is, the DASI module uses different channel attention methods for the features of different groups under the Group mechanism to obtain transformed features of corresponding dimensions;

[0076] Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function;

[0077] Step 1.3.5: The weighted features are recombined. That is, after the DASI module completes the weighted calculation, the features within different channels (i.e., Groups) are rearranged and fused to adapt to the format requirements of subsequent network calculations; and after channel-wise multiplication (×) and weighted addition (+), channel concatenation is performed; to obtain the final optimized feature representation; the optimization of the ACFN module is completed; after being processed by the DASI module, not only can multi-scale feature fusion be achieved, but also the most valuable feature information can be intelligently selected, thus significantly improving the accuracy and robustness of the model in the object detection task; the DASI module realizes further optimization on the basis of the ACFN module. The ACFN module is mainly responsible for the initial fusion of features, but it does not have the ability to intelligently select different scale features by itself, but directly splices and processes information of different scales; the DASI module introduces a feature selection mechanism on this basis, dynamically adjusts the weights of different scale features through a dimension-aware strategy, and ensures that high-level semantic information and low-level detail information can be reasonably combined, thereby further improving the detection performance; combined with Figure 4 (corresponding to Figure 5 the detailed process), the upper part (P3, P4, P5 feature alignment and fusion) belongs to the processing part of ACFN, while the lower part (the part after DWConv and Add) belongs to the working area of DASI; DASI optimizes the features output by ACFN through channel division and selective integration, making it more suitable for the object detection task of different scales; in specific applications, this optimization mechanism is particularly prominent in the small object detection task; since low-level features are crucial in small object detection, DASI effectively improves the detection ability of small objects by enhancing the weights of low-level features; while in the large object detection task, high-level features usually play a dominant role, and DASI significantly improves the detection accuracy of large objects by strengthening the influence of high-level features;

[0078] In the improved RT-DETR network, after the input image undergoes preliminary feature extraction through ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features. One output of the last BasicBlock layer is directly fed into the first feature fusion module (Concat) in the backbone feature extraction part, and the other output is successively fed into the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives the feature layers output by other BasicBlocks layers except the last BasicBlock layer to achieve deeper feature fusion. The traditional Backbone mainly relies on CNN for feature extraction. The introduction of the DA-ACFN module not only enhances the fusion ability of features at different scales but also introduces richer multi-scale information at an early stage, enabling subsequent network layers to make full use of local and global information; the ACFN module is responsible for multi-scale feature alignment to ensure the consistency of features at different scales and selectively enhances key features through the DASI module, thus reducing the interference of redundant information on the detection task; in addition, the DA-ACFN module enables the low-level features (with rich detail information) and high-level features (with stronger semantic information) to be more effectively fused, thereby improving the detection ability for small targets; one output of the first DA-ACFN module is fed into the first feature fusion module (Concat) after passing through the third convolutional layer, and the other output is fed into the Dysample layer in the feature fusion part (Neck); the CCFM layer in the feature fusion part (Neck) is replaced with a second DA-ACFN module to optimize the feature expression ability and enhance the multi-scale feature fusion effect; compared with the original RT-DETR, the improved network enhances the original modules at these two key positions to improve the performance of the model in complex detection scenarios; to enhance the feature expression ability, the DA-ACFN module performs multi-scale feature alignment through the ACFN module, enabling features from different Backbone layers to be combined more smoothly and avoiding information loss caused by scale differences; at the same time, the introduction of the DASI module realizes the selective integration of features. Compared with the simple feature splicing of the CCFM in the original RT-DETR network, the DASI module can adaptively select the most discriminative features, effectively reducing redundant information and improving the effectiveness of features; in addition, the DA-ACFN module further improves the detection accuracy of small targets; the CCFM of the original RT-DETR network may cause small target features to be masked by high-level features during the multi-scale fusion process due to only simple feature fusion. The DA-ACFN adaptively enhances low-level features through the DASI module, thereby improving the detection ability for small targets and ensuring that the model performs more stably and accurately in complex detection tasks.

[0079] Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR object detection model;

[0080] Step 3: Apply the RT-DETR object detection model to the detection of printed circuit board surface defects.

[0081] In this paper, the optimized RT-DETR algorithm is compared with mainstream object detection algorithms such as Faster R-CNN, SSD, YOLOv3, YOLOv5, and YOLOv8. Four metrics, namely Parameters, fps, mAP@0.5, and mAP@0.5:0.95, are used as evaluation criteria. The detection performance of different algorithms for the filter surface defects on the dataset is shown in Table 1.

[0082] Table 1 Comparison of the results of the algorithm of the present invention and other mainstream algorithms

[0083]

[0084] Step 4: Train the optimized YOLOv5 network algorithm through the filter surface defect dataset, and obtain the YOLOv5 object detection network for filter surface defect detection; among them, the filter surface defect dataset includes collecting different defect categories of the filter through a vision device in different environments to make the filter dataset, and then using LabelImg to label the filter dataset, while generating a.txt file, and randomly dividing the filter dataset and the corresponding.txt file into a training set, a validation set, and a test set according to the ratio of 7:1:2.

[0085] The deep learning method for detecting printed circuit board surface defects of the present invention can be applied to vision detection automation equipment, as Figure 7 shown. The equipment includes the following modules: a printed circuit board transportation module 1, an inner surface detection module 2, a turning module 3, an outer surface detection module 4, and a screening module 5. The printed circuit board loading module 1 transports the circuit board to the surface defect detection module 2 through a conveyor belt, and then the circuit board is turned over by the turning module 3 and enters the outer surface detection module 4. After the circuit board passes the outer surface detection, the screening module 5 is used to screen the circuit board for qualified and defective products. The industrial control screen 6 displays the operation status of the equipment in real time, and counts the number of various defective products and the change in the qualification rate of the circuit board. The printer 7 is used to print the detection data. The equipment is equipped with an Ethernet industrial area array camera with the model MV-E200-10GC.

[0086] Based on the above technical solution, taking the production line of a printed circuit board production workshop as the background, the following example is given. In this experiment, 1000 products were randomly produced, and mainstream object detection algorithms such as Faster R-CNN, SSD, YOLOv3, YOLOv5, and YOLOv8 were compared with the improved algorithm of the present invention. The experiment was carried out on the Windows 10 operating system platform, and the environment was configured based on PyTorch 2.1.1 and Python 3.8. The hardware devices were an Intel Core i5-12600KF CPU (main frequency 3.51 GHz), 32 GB of memory, and an NVIDIA GeForce GTX 4070 Ti SUPER graphics card (16 GB of video memory). The training category of the present invention is 6, the number of iterations is set to 150 epochs, the batch size is set to 8, and precision (P), recall (R), average precision (AP), and mean average precision (mAP) are used as performance evaluation indicators.

[0087] Among them, Precision represents the ratio of correct results among all predictions of positive samples; Recall represents the ratio of correctly predicted positive samples among all positive samples; the PR curve takes Recall as the abscissa and Precision as the ordinate. The larger the area below the curve, the better the performance of the model on the dataset. The area enclosed by the curve is the average precision AP, which represents the average of the precision rates for each class; mAP represents the average of the APs for all classes. The larger the mAP value, the better the model performance. The specific calculation formulas for each index are as follows:

[0088] (7)

[0089] (8)

[0090] (9)

[0091] (10)

[0092] In the formula, TP represents true positives; FP represents false positives; FN represents false negatives; n represents the total number of detection categories, i represents the category index, AP is calculated by interpolation method, and when n = 1, 'AP' is mAP.

[0093] In this paper, the optimized RT-DETR algorithm is compared with 10 mainstream object detection algorithms such as Faster R-CNN, RT-DETR, and YOLOv3, and four indicators of Parameters, fps, mAP@0.5, and mAP@0.5:0.95 are used for evaluation. The experimental results are shown in Table 1.

[0094] To ensure a fair comparison, all algorithms maintain the same parameter settings during training, including batch size and number of iterations. The improved RT-DETR algorithm outperforms Faster R-CNN, SSD, YOLOv3, YOLOv5, YOLOv8, Centernet, RetinaNet, CDI-YOLO, Light-PDD, and RT-DETR by 12.4%, 46.7%, 5.1%, 5.0%, 8.0%, 38.8%, 15.9%, 0.9%, 5.4%, and 0.4% respectively in terms of mAP@0.5:0.95. These results indicate that the improved RT-DETR algorithm has been enhanced in both detection accuracy and the detection efficiency of printed circuit board surface defects.

[0095] To further verify the effectiveness of each application strategy in this algorithm model, ablation experiments were conducted in this section. The original RT-DETR was denoted as model (1), RT-DETR + DySample as model (2), RT-DETR + DySample + RetBlock as model (3), and RT-DETR + DySample + RetBlock + DA-ACFN as model (4). ("+" indicates the introduction of this module, and the results are shown in Table 2.)

[0096] Table 2 Comparison of various indicators in ablation experiments

[0097]

[0098] Meanwhile, we adopted a more rigorous statistical analysis to support the conclusion of performance improvement. We conducted a more detailed statistical verification of the model performance. We used the K-fold cross-validation method, which almost equally divided the dataset into 5 parts and conducted 5 experiments. For the important indicator mAP in the field of object detection, we obtained the following experimental results: 0.90, 0.95, 0.94, 0.97, and 0.92. Based on these experimental results, we further calculated the mean (0.936), standard deviation (0.026), and 95% confidence interval ([0.911, 0.961]). It can be seen from these statistical results that the performance of our model is relatively stable across different training / test splits, and the fluctuation range of mAP is small, indicating that the performance of the model is highly reliable. At the same time, these results further support the effectiveness of the improvement method we proposed.

[0099] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting surface defects of printed circuit boards based on deep learning, including the RT-DETR network algorithm, characterized in that: The deep learning method includes the following steps: Step 1: Improve the RT-DETR network algorithm, specifically: Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm and perform dynamic upsampling in the CCFM sampling stage; Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that adopts the RetBlock structure; Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain the improved RT-DETR network; The DA-ACFN module includes the ACFN module and the DASI module. The ACFN module includes a feature splicing module. The first input end of the feature splicing module is connected to the downsampling module and is used to input low-level features. The second input of the feature splicing module is connected to the convolutional module and is used to input middle-level features. The third input of the feature splicing module is connected to the convolutional module and is used to input middle-level features. Several outputs of the feature splicing module are all passed through the depthwise separable convolutional module and then sequentially through the skip connection module and the 1×1 convolutional module for information integration to obtain the integrated information. The integrated information is skip-connected with one output of the feature splicing module to obtain the fused feature map. The ACFN module is optimized by the DASI module to form the improved DA-ACFN module. Among them, the process of the DASI module optimizing the ACFN module is: Step 1.3.1: Receive the fused feature map, and assume the dimension of the fused feature map is C × H × W; Step 1.3.2: Divide the fused feature map into four groups of features, namely low-level features, high-level features, background features, and the feature selection part; Step 1.3.3: Use the Group mechanism to divide the channels and apply respective channel attention strategies to the features of different groups to obtain transformed features of corresponding dimensions; Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function; Step 1.3.5: The weighted features are recombined and channel concatenated after performing per-channel multiplication (×) and weighted addition (+) to obtain the finally optimized feature representation; complete the optimization of the ACFN module; In the improved RT-DETR network, after the input image undergoes preliminary feature extraction through ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features. One output of the last BasicBlock layer is directly fed into the first feature fusion module in the backbone feature extraction part, and the other output is sequentially fed into the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives the feature layers output by other BasicBlocks layers except the last BasicBlock layer; one output of the first DA-ACFN module is fed into the first feature fusion module after passing through the third convolutional layer, and the other output is fed into the Dysample layer in the feature fusion part; the CCFM layer in the feature fusion part is replaced with the second DA-ACFN module; Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR object detection model; Step 3: Use the RT-DETR object detection model for printed circuit board surface defect detection.

2. The method for detecting surface defects of a printed circuit board based on deep learning according to claim 1, wherein: As shown in Figure 2, the dynamic upsampling in Step 1.1 specifically includes the following steps: Step 1.1.1: Generate an offset with a size of 2S²×H×W from the input feature map with a size of C×H×W through a linear projection layer; Step 1.1.2: Convert the offset to an appropriate spatial resolution, and finally calculate the final coordinates of the sampling points by adding the offset to the original grid point coordinates; Step 1.1.3: Use the generated sampling points to resample the feature map to obtain the upsampled feature map.

3. The method for detecting surface defects of a printed circuit board based on deep learning according to claim 2, wherein: The upsampling process of the DySample module is explained by the following formula: (1) In this equation, 𝑥 represents the input feature map, 𝑥′ represents the output feature map, and 𝑔 is the original sampling grid; bilinear(x) performs fixed multiple scaling on the input feature map by bilinear interpolation, interpolating through the weighted average of adjacent pixels; pixel_shuffle() is used as a pixel rearrangement operation, rearranging the pixels in the channel dimension to the spatial dimension so that the enlarged feature map is evenly distributed in space; reshaping_sample() is used to reshape the sampled feature map.

4. The method for detecting surface defects of a printed circuit board based on deep learning according to claim 3, wherein: DySample includes a static range factor sampling strategy and a dynamic range factor sampling strategy; among them, the static range factor sampling strategy performs bilinear interpolation with a 0.25-fold scaling and then applies pixel shuffle for pixel rearrangement; the dynamic range factor sampling strategy performs bilinear interpolation with a 0.5-fold scaling and combines it with the dynamically calculated offset for compensation.

5. The method for detecting surface defects of a printed circuit board based on deep learning according to claim 1, characterized in that: In Step 1.2: The RetBlockC3 module includes depthwise separable convolution, layer normalization, Manhattan self-attention mechanism, and a feed-forward neural network. The input features are sequentially sent to the Manhattan self-attention mechanism after depthwise separable convolution and layer normalization. The Manhattan self-attention mechanism calculates the Manhattan distance between different positions in the feature map and assigns different weights to the features at different positions according to the spatial decay principle. The output of the Manhattan self-attention mechanism is normalized again through layer normalization, and the normalized features enter the feed-forward neural network for feature transformation and enhancement.

6. The method for detecting surface defects of a printed circuit board based on deep learning according to claim 5, wherein: Output of the RetBlockC3 module is as follows: (2) Among them, DWConv is the depthwise separable convolution operation of the context local enhancement module, and are used to calculate the attention scores in the vertical and horizontal directions respectively; (3) (4) Among them, is the horizontal Manhattan self-attention calculation, that is, the attention score calculated along the horizontal direction; is the vertical Manhattan self-attention calculation, that is, the attention score calculated along the vertical direction; is the query, key, and value matrices in the horizontal direction; is the normalization function that maps the value to between (0, 1) and is used to calculate the attention weight, is the scaling factor, and are the horizontal and vertical attenuation matrices respectively; and are the Manhattan distances between two tokens n and m in the horizontal and vertical directions respectively; (5) (6) Among them, γ is a decay coefficient, Xn and Xm are the horizontal coordinates of the nth and mth tokens in the image respectively, and Yn and Ym are the vertical coordinates of the nth and mth tokens in the image respectively.

Citation Information

Patent Citations

  • PCB surface defect detection method and system based on RT-DETR detector

    CN118822989A

  • Automatic driving target detection method based on improved RT-DETR

    CN119314144A