Printed circuit board surface defect detection method based on deep learning
By improving the RT-DETR network algorithm, DySample, RetBlockC3 and DA-ACFN modules were introduced, and the problem of micro defect detection on the surface of the printed circuit board is solved, fast and accurate defect detection is achieved, and detection accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510496935.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art is difficult to detect small defects on the surface of printed circuit boards quickly and accurately, especially the problems of high detection of multiple defects arising in complex manufacturing processes and the problem of high difficulty in detecting small defects.
The improved RT-DETR network algorithm is adopted, and the computing complexity and feature extraction capabilities of the network are optimized by introducing the DySample module, RetBlockC3 module and DA-ACFN module, and the computing complexity and feature extraction capabilities of the network are improved, and the efficiency and accuracy of the model in multi-scale target processing are improved.
It realizes fast and accurate detection of printed circuit board surface defects, improves target detection accuracy and robustness, and can more effectively identify multiple defects and small defects.
Smart Images

Figure CN120031872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optical image processing, and in particular to a printed circuit board surface defect detection method based on deep learning. Background Art
[0002] Surface defect detection of printed circuit boards (PCBs) plays a vital role in quality control. With the rapid development of computer vision (CV) technology, traditional image processing methods (such as thresholding, edge detection, and region-based methods) as well as modern machine learning and deep learning technologies are gradually replacing inefficient manual detection methods. Among them, optical image processing technology has become the mainstream method for PCB surface defect detection. However, with the rapid development of the electronic information technology industry and the increasing market demand for PCB products, PCB manufacturing is moving towards high reliability, high density, and miniaturization. These trends have not only expanded the functions of PCBs, increased line density, but also reduced their size, which has brought many challenges to machine vision-based PCB surface defect detection, which are mainly reflected in the following two aspects:
[0003] 1. Complex manufacturing process leads to multiple defects: The manufacturing process of PCB is complex and prone to defects, such as open circuits, short circuits, and burrs. These tiny defects may seriously affect product quality and even cause serious safety hazards, such as increased risk of leakage and even fire, endangering the life safety of users. Therefore, the detection model is required to have the ability to efficiently and accurately identify multiple defects.
[0004] 2. Difficulty in detecting tiny defects: Due to the influence of different manufacturing processes, defects are usually small in area and tiny in size. For example, the size of PCB surface defects is usually less than 4500 pixels, and the size of burrs is usually less than 300 pixels, which only accounts for 0.005%-0.07% of a high-resolution PCB image (about 6.5 million pixels). Such a tiny proportion makes it a very challenging task to detect these defects quickly and accurately. Summary of the invention
[0005] In view of the problems in the prior art, the present invention provides a printed circuit board surface defect detection method based on deep learning, which aims to achieve fast, accurate and high-speed detection of printed circuit board surface defects.
[0006] A printed circuit board surface defect detection method based on deep learning includes an RT-DETR network algorithm. The deep learning method includes the following steps:
[0007] Step 1: Improve the RT-DETR network algorithm, specifically:
[0008] Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm, and enable the CCFM sampling stage to perform dynamic upsampling;
[0009] Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that adopts the RetBlock structure;
[0010] Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain the improved RT-DETR network;
[0011] Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR target detection model;
[0012] Step 3: Use the RT-DETR object detection model for printed circuit board surface defect detection.
[0013] Further, the DA-ACFN module includes an ACFN module and a DASI module, and the ACFN is optimized by the DASI module to form an improved DA-ACFN module.
[0014] Further, the ACFN module includes a feature splicing module, a first input end of the feature splicing module is connected to the downsampling module and used to input low-level features, a second input end of the feature splicing module is connected to the convolution module and used to input middle-level features, and a third input end of the feature splicing module is connected to the convolution module and used to input middle-level features; several outputs of the feature splicing module are all passed through the depthwise separable convolution module and then sequentially through the jump connection module and the 1×1 convolution module for information integration to obtain integrated information, and the integrated information is jump-connected with one output end of the feature splicing module to obtain a fused feature map;
[0015] The process of DASI module optimizing ACFN module is as follows:
[0016] Step 1.3.1: Receive the fused feature map, and set the dimension of the fused feature map to be C × H × W;
[0017] Step 1.3.2: Divide the fused feature map into four groups of features, namely low-level features, high-level features, background features, and feature selection parts;
[0018] Step 1.3.3: Use the Group mechanism to divide the channels, and apply the respective channel attention strategies to the features of different groups to obtain the transformation features of the corresponding dimensions;
[0019] Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function;
[0020] Step 1.3.5: The weighted features are recombined and channel-wise concatenated through channel-by-channel multiplication (×) and weighted addition (+); to obtain the final optimized feature representation; to complete the optimization of the ACFN module.
[0021] Further, in the improved RT-DETR network, after the input image is subjected to preliminary feature extraction by ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features, one output of the last BasicBlock layer is directly sent to the first feature fusion module in the backbone feature extraction part, and the other output is successively sent to the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives feature layers output by other BasicBlocks layers except the last BasicBlock layer; one output of the first DA-ACFN module is sent to the first feature fusion module after passing through the third convolutional layer, and the other output is sent to the Dysample layer in the feature fusion part; the CCFM layer in the feature fusion part is replaced by the second DA-ACFN module.
[0022] Further: Combine Figure 2 As shown, the dynamic upsampling in step 1.1 specifically includes the following steps:
[0023] Step 1.1.1: Pass the input feature map of size C×H×W through a linear projection layer to generate an offset of size 2S²×H×W;
[0024] Step 1.1.2: Convert the offset to an appropriate spatial resolution, and finally calculate the final coordinates of the sampling point by adding the offset to the original grid point coordinates;
[0025] Step 1.1.3: Resample the feature map using the generated sampling points to obtain the upsampled feature map.
[0026] Further: The upsampling process of the DySample module is explained by the following formula: (1)
[0027] In this equation, represents the input feature map, represents the output feature map, is the original sampling grid; bilinear(x) uses bilinear interpolation to scale the input feature map by a fixed multiple and interpolate the value by the weighted average of adjacent pixels; pixel_shuffle() is a pixel rearrangement operation that rearranges the pixels of the channel dimension to the spatial dimension so that the enlarged feature map is evenly distributed in space; reshaping_sample() is used to reshape the sampled feature map.
[0028] Further, DySample includes a static range factor sampling strategy and a dynamic range factor sampling strategy; wherein the static range factor sampling strategy performs bilinear interpolation by scaling by 0.25 times, and then applies pixel shuffle to rearrange pixels; the dynamic range factor sampling strategy performs bilinear interpolation by scaling by 0.5 times, and compensates with a dynamically calculated offset;
[0029] Further: In step 1.2: The RetBlockC3 module includes depthwise separable convolution, layer normalization, Manhattan self-attention mechanism and feedforward neural network. The input features are sequentially sent to the Manhattan self-attention mechanism after depthwise separable convolution and layer normalization. The Manhattan self-attention mechanism calculates the Manhattan distance between different positions in the feature map, and assigns different weights to features at different positions according to the principle of spatial attenuation. The output of the Manhattan self-attention mechanism is normalized again through layer normalization, and the normalized features enter the feedforward neural network for feature transformation and enhancement.
[0030] Further: Output of RetBlockC3 module for: (2)
[0031] Among them, DWConv is the depth-separable convolution operation of the context local enhancement module. and Calculate the attention scores from the vertical and horizontal directions respectively; (3) (4)
[0032] in, is the Manhattan self-attention calculation in the horizontal direction, that is, the attention score calculated along the horizontal direction; is the vertical Manhattan self-attention calculation, that is, the attention score calculated along the vertical direction; is the query, key, and value matrix in the horizontal direction; is a normalization function that maps values to (0, 1) and is used to calculate the attention weight. is the scaling factor, , are the horizontal and vertical attenuation matrices, respectively; and are the Manhattan distances between two tokens n and m horizontally and vertically respectively; (5) (6)
[0033] Where γ is the attenuation coefficient, Xn and Xm are the horizontal coordinates of the nth and mth token in the image, respectively, and Yn and Ym are the vertical coordinates of the nth and mth token in the image, respectively.
[0034] The beneficial effects of the present invention are as follows: the RT-DETR network is improved through the DySample, RetBlockC3 modules and the DA-ACFN modules, thereby reducing the computational complexity of the RT-DETR network and being able to more effectively focus on important areas in the image to enhance the feature extraction capability of the model, thereby improving the efficiency and accuracy of the network in multi-scale target processing, and being used to improve the target detection accuracy and robustness in printed circuit board defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A flowchart of the invention;
[0036] Figure 2 It is a structural diagram of the DySample module in the present invention;
[0037] Figure 3 It is a schematic diagram of the structure of the RetBlock module used in the RT-DRDA network algorithm;
[0038] Figure 4 It is a schematic diagram of the structure of the DA-ACFN module;
[0039] Figure 5 It is a schematic diagram of the structure of the DASI module;
[0040] Figure 6 It is the structural diagram of RT-DRDA;
[0041] Figure 7 This is a schematic diagram of the structure of the visual inspection automation equipment. DETAILED DESCRIPTION
[0042] The present invention is described in detail below in conjunction with the accompanying drawings. The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention. The directional terms such as left, middle, right, top, and bottom in the embodiments of the present invention are only relative concepts or are based on the normal use state of the product, and should not be considered as restrictive.
[0043] A printed circuit board surface defect detection method based on deep learning, such as Figure 1 and Figure 6 As shown, including the RT-DETR network algorithm, the deep learning method includes the following steps:
[0044] Step 1: Improve the RT-DETR network algorithm, specifically:
[0045] Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm, and perform dynamic upsampling in the CCFM sampling stage; this can reduce the computational complexity while retaining rich semantic information and improving the detection accuracy of the model;
[0046] Combination Figure 2 As shown, the dynamic upsampling in step 1.1 specifically includes the following steps:
[0047] Step 1.1.1: Pass the input feature map of size C×H×W through a linear projection layer to generate an offset of size 2S²×H×W;
[0048] Step 1.1.2: Convert the offset to an appropriate spatial resolution, and finally calculate the final coordinates of the sampling point by adding the offset to the original grid point coordinates;
[0049] Step 1.1.3: Use the generated sampling points to resample the feature map to obtain the upsampled feature map; this can reduce the computational complexity while retaining rich semantic information and improving the detection accuracy of the model;
[0050] The upsampling process of the DySample module is explained by the following formula: (1)
[0051] In this equation, represents the input feature map, represents the output feature map, is the original sampling grid; bilinear(x) uses bilinear interpolation to scale the input feature map by a fixed multiple (usually 2× or 4×), and interpolates the weighted average of adjacent pixels to improve the smoothness and continuity of the feature map, thereby reducing artifacts that may be generated during upsampling; pixel_shuffle() is a pixel rearrangement operation that rearranges the pixels in the channel dimension to the spatial dimension so that the enlarged feature map is more evenly distributed in space, thereby improving sampling accuracy. For example, in the case of 2× enlargement, the 2×2 sub-blocks in a channel will be remapped to the spatial dimension to generate a new high-resolution feature map; reshaping_sample() is used to reshape the sampled feature map so that its shape adapts to the subsequent calculation structure of the network, ensuring that the compensated features can be correctly passed to the subsequent network layers to maintain the consistency and stability of the feature information;
[0052] In addition, DySample includes the Static Scope Factor sampling strategy and the Dynamic Scope Factor sampling strategy. The Static Scope Factor sampling strategy performs bilinear interpolation by scaling by 0.25 times, and then applies pixel shuffle to rearrange pixels, so that the features can be evenly expanded in the spatial dimension. The Dynamic Scope Factor sampling strategy uses bilinear interpolation by scaling by 0.5 times, and combines it with the dynamically calculated offset for compensation, so that it is more adaptable in the sampling process, so that the distribution of feature points can be adaptively adjusted according to the input features. Figure 2 In the above, s represents the sampling scale factor, which controls the distribution range of the sampling points; sH and sW represent the dynamic scaling factors in the height (H) and width (W) directions, respectively, which are used to adjust the offset of the sampling points in different directions; 0.5 and 0.25 correspond to the scaling ratios of the dynamic range factor and the static range factor, respectively, which determine the magnification of the feature map; g represents the sampling point grid, which determines the specific sampling position in the upsampling process to ensure the accuracy and consistency of feature extraction;
[0053] Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that uses the RetBlock structure; Figure 3As shown in the figure, the RetBlockC3 module includes deep separable convolution (DWConv), layer normalization (LN), Manhattan self-attention mechanism (MaSA) and feedforward neural network (FFN). The input features are sequentially sent to the Manhattan self-attention mechanism (MaSA) after deep separable convolution (DWConv) and layer normalization (LN). The Manhattan self-attention mechanism (MaSA) calculates the Manhattan distance between different positions in the feature map and assigns different weights to the features at different positions according to the principle of spatial attenuation. The output of the Manhattan self-attention mechanism (MaSA) is normalized again by layer normalization (LN) to ensure the stability and consistency of the features during transmission. The optimized features enter the feedforward neural network (FFN) for feature transformation and enhancement to improve the expression ability of the target features, and the optimized features are fused into the backbone network to provide more discriminative feature representation for subsequent target detection; among them, the deep separable convolution (DWConv) can enhance the feature extraction ability while reducing the computational complexity, so that the network can focus on key areas more accurately and reduce the interference of redundant information; layer normalization (LN) can stabilize the feature distribution and improve the convergence speed of the training process; Manhattan self-attention mechanism (MaSA) enables the model to focus more on the interaction of local information and effectively weaken the interference of background areas on detection tasks; it can be seen that RetBlockC3 combines DWConv, LN, Manhattan self-attention mechanism (MaSA) and FFN, which not only reduces the computational complexity while improving the feature expression ability, but also enhances the network's attention to local areas, making the model have higher accuracy and robustness in multi-scale feature fusion and target detection tasks; Output of the RetBlockC3 module for: (2)
[0054] Among them, DWConv is the depth-separable convolution operation of the contextual local enhancement module. and Calculate the attention scores from the vertical and horizontal directions respectively; (3) (4)
[0055] in, is the Manhattan self-attention calculation in the horizontal direction, that is, the attention score calculated along the horizontal direction (H direction); is the vertical Manhattan self-attention calculation, i.e., the attention score calculated along the vertical direction (W direction); It is the query, key, and value matrix in the horizontal direction (H); is a normalization function that maps values to (0, 1) and is used to calculate the attention weight. is the scaling factor, , are the horizontal (H direction) and vertical (W direction) attenuation matrices, respectively; and are the Manhattan distances between two tokens n and m in the horizontal (H direction) and vertical (W direction) directions respectively; (5) (6)
[0056] Where γ is the attenuation coefficient, Xn and Xm are the horizontal coordinates of the nth and mth tokens in the image, respectively, and Yn and Ym are the vertical coordinates of the nth and mth tokens in the image, respectively;
[0057] Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain an improved RT-DETR network; wherein the DA-ACFN module includes an ACFN module (Adaptive Contextual Fusion Network) and a DASI module (Dimension Aware Selective Integration Module), and the ACFN module is optimized by the DASI module (Dimension Aware Selective Integration Module) to form an improved DA-ACFN module;
[0058] The ACFN module (Adaptive Context Fusion Network) aims to fully explore the information relevance of features at different scales, thereby improving the overall performance of the detection model, such as Figure 4As shown in the figure, the ACFN module includes a feature concatenation module (Concat). The first input of the feature concatenation module is connected to the downsampling module (ADown) and used to input low-level features (P3). Low-level features have rich detail information but are susceptible to noise interference; the second input of the feature concatenation module is connected to the convolution module (Conv) and used to input middle-level features (P4). Middle-level features are used to strike a balance between semantic expression and detail retention. The third input of the feature concatenation module is connected to the convolution module (Conv) and used to input middle-level features (P5). High-level features have strong semantic information but lack fine target details; these features are extracted by convolutional networks at different levels to form a complete multi-scale feature representation. Secondly, in the multi-scale feature alignment stage, since P3, P4 and P5 have different spatial sizes, they need to be aligned through downsampling (ADown) and convolution (Conv) to ensure that the features of different scales can be effectively fused in the same scale space; the several outputs of the feature splicing module are all passed through the depthwise separable convolution module (DWConv) for further feature extraction to reduce the computational complexity and enhance the local information interaction capability; then the information is integrated through the jump connection module (ADD) and the 1×1 convolution module (Conv) in turn to obtain the integrated information, and the integrated information is jump-connected with one output of the feature splicing module to obtain the fused feature map;
[0059] like Figure 5 As shown in the figure, the process of DASI module optimizing ACFN module (Adaptive Context Fusion Network) is as follows:
[0060] Step 1.3.1: Receive the fused feature map, and set the dimension of the fused feature map to be C × H × W;
[0061] Step 1.3.2: Divide the fused feature map into four groups of features (Group = 4), namely low-level features (Cl), high-level features (Ch), background features (CBS) and feature selection (CsS). This division enables the DASI module to accurately identify information of different scales and prevent high-level features from suppressing low-level features during the fusion process.
[0062] Step 1.3.3: Use the Group mechanism to divide the channels, and apply different channel attention strategies to the features of different groups. That is, the DASI module uses different channel attention methods for the features of different groups under the Group mechanism to obtain the transformation features of the corresponding dimensions.
[0063] Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function;
[0064] Step 1.3.5: The weighted features are recombined, that is, after completing the weighted calculation, the DASI module rearranges and fuses the features in different channels (that is, Group) to adapt them to the format requirements of subsequent network calculations; and performs channel splicing after channel-by-channel multiplication (×) and weighted addition (+); to obtain the final optimized feature representation; to complete the optimization of the ACFN module; after being processed by the DASI module, not only can the fusion of multi-scale features be achieved, but also the most valuable feature information can be intelligently screened out, thereby significantly improving the accuracy and robustness of the model in target detection tasks; the DASI module is further optimized on the basis of the ACFN module. The ACFN module is mainly responsible for the initial fusion of features, but it does not have the ability to intelligently screen features of different scales, but directly splices and processes information of different scales; on this basis, the DASI module introduces a feature selection mechanism, which dynamically adjusts the weights of features of different scales through a dimension perception strategy to ensure that high-level semantic information and low-level detail information can be reasonably combined, thereby further improving detection performance; combined Figure 4 (correspond Figure 5 The detailed process of ACFN is shown in Figure 1. The upper part (P3, P4, P5 feature alignment and fusion) belongs to the processing part of ACFN, while the lower part (the part after DWConv and Add) belongs to the working area of DASI. DASI optimizes the features of ACFN output by channel division and selective integration, making it more suitable for detection tasks of targets of different scales. In specific applications, this optimization mechanism is particularly prominent in small target detection tasks. Since low-level features are crucial in small target detection, DASI effectively improves the detection ability of small targets by enhancing the weight of low-level features. In large target detection tasks, high-level features usually play a leading role. DASI significantly improves the detection accuracy of large targets by strengthening the influence of high-level features.
[0065] In the improved RT-DETR network, after the input image is subjected to preliminary feature extraction by ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features. One output of the last BasicBlock layer is directly sent to the first feature fusion module (Concat) in the backbone feature extraction part, and the other output is sent to the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives the feature layers output by other BasicBlocks layers except the last BasicBlock layer to achieve deeper feature fusion. The traditional Backbone mainly relies on CNN for feature extraction, and the introduction of the DA-ACFN module not only enhances the fusion capability of features of different scales, but also introduces richer multi-scale information in the early stage, so that subsequent network layers can make full use of local and global information. The ACFN module is responsible for multi-scale feature alignment to ensure the consistency of features at different scales, and selectively enhances key features through the DASI module, thereby reducing the interference of redundant information on detection tasks. In addition, the DA-ACFN The module enables a more effective fusion of low-level features (with rich detail information) and high-level features (with stronger semantic information), thereby improving the detection capability of small targets; one output of the first DA-ACFN module is sent to the first feature fusion module (Concat) after the third convolutional layer, and the other output is sent to the Dysample layer in the feature fusion part (Neck); the CCFM layer in the feature fusion part (Neck) is replaced with the second DA-ACFN module to optimize the feature expression capability and enhance the multi-scale feature fusion effect; compared with the original RT-DETR, the improved network enhances the original modules at these two key positions to improve the performance of the model in complex detection scenarios; in order to improve the feature expression capability, the DA-ACFN module performs multi-scale feature alignment through the ACFN module, so that features from different Backbone The features of the layers can be combined more smoothly to avoid information loss caused by scale differences; at the same time, the introduction of the DASI module realizes the selective integration of features. Compared with the CCFM of the original RT-DETR network, which only performs simple feature splicing, the DASI module can adaptively select the most discriminative features, effectively reduce redundant information and improve the effectiveness of features; in addition, the DA-ACFN module further improves the detection accuracy of small targets; the CCFM of the original RT-DETR network only performs simple feature fusion, which may cause the features of small targets to be masked by high-level features in the multi-scale fusion process, while DA-ACFN adaptively enhances low-level features through the DASI module, thereby improving the detection ability of small targets and ensuring that the model performs more stably and accurately in complex detection tasks;
[0066] Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR target detection model;
[0067] Step 3: Use the RT-DETR object detection model for printed circuit board surface defect detection.
[0068] This paper compares the optimized RT-DETR algorithm with mainstream target detection algorithms such as Faster R-CNN, SSD, YOLOv3, YOLOv5, and YOLOv8, and uses four indicators, Parameters, fps, mAP@0.5, and mAP@0.5:0.95, as evaluation criteria. The detection performance of different algorithms for filter surface defects on the dataset is shown in Table 1.
[0069] Table 1 Comparison of the results of the proposed algorithm with other mainstream algorithms
[0070] Step 4: Train the optimized YOLOv5 network algorithm through the filter surface defect dataset, and obtain the YOLOv5 target detection network for filter surface defect detection; wherein, the filter surface defect dataset includes collecting different defect categories of the filter through a visual device under different environments and making a filter dataset, and then labeling the filter dataset through LabelImg, generating a .txt file at the same time, and randomly dividing the filter dataset and the corresponding .txt file into a training set, a validation set, and a test set in a ratio of 7:1:2.
[0071] The printed circuit board surface defect detection deep learning method of the present invention can be applied to visual inspection automation equipment, such as Figure 7 As shown. The equipment includes the following modules: printed circuit board transport module 1, inner surface detection module 2, flip module 3, outer surface detection module 4, screening module 5. The printed circuit board feeding module 1 sends the circuit board to the surface defect detection module 2 through a conveyor belt, and then flips the circuit board through the flip module 3 and enters the outer surface detection module 4. After the circuit board passes the outer surface inspection, the screening module 5 is used to screen the circuit board for qualified and defective products. The industrial control screen 6 displays the operating status of the equipment in real time, and counts the number of various defective products and the changes in the qualified rate of the circuit board. The printer 7 is used to print the detection data. The equipment is equipped with an Ethernet industrial array camera, model MV-E200-10GC.
[0072] According to the above technical solution, the following example is used as the background of the production line of a printed circuit board production workshop. This experiment randomly produces 1,000 products, and compares the mainstream target detection algorithms such as Faster R-CNN, SSD, YOLOv3, YOLOv5, and YOLOv8 with the improved algorithm of the present invention. The experiment was carried out on the Windows 10 operating system platform, and the environment configuration was based on PyTorch 2.1.1 and Python 3.8. The hardware equipment is Intel Core i5-12600KF CPU (main frequency 3.51 GHz), 32 GB memory and NVIDIA GeForce GTX 4070 Ti SUPER graphics card (16 GB video memory). The training categories of the present invention are 6, the number of iterations is set to 150 epochs, the batch size is set to 8, and the precision (P), recall (R), average precision (AP) and mean average precision (mAP) are used as performance evaluation indicators.
[0073] Among them, Precision represents the correct ratio of all the results predicted as positive samples; Recall represents the ratio of all positive samples that are correctly predicted; the PR curve uses Recall as the horizontal axis and Precision as the vertical axis. The larger the area on the lower left of the curve, the better the model effect on the data set. The area enclosed by the curve is the average precision AP, which represents the average value of each category of precision; mAP represents the average value of all categories of AP. The larger the mAP value, the better the model performance. The specific calculation formulas for each indicator are as follows: (7) (8) (9) (10)
[0074] Where TP represents a true positive example, FP represents a false positive example, FN represents a false negative example, n represents the total number of detection categories, i represents the category index, and AP is calculated by interpolation. When n=1, 'AP' is mAP.
[0075] This paper compares the optimized RT-DETR algorithm with 10 mainstream object detection algorithms, including Faster R-CNN, RT-DETR and YOLOv3, and uses four indicators, Parameters, fps, mAP@0.5 and mAP@0.5:0.95, for evaluation. The experimental results are shown in Table 1.
[0076] To ensure a fair comparison, all algorithms maintain the same parameter settings during training, including batch size and number of iterations. The improved RT-DETR algorithm performs 12.4%, 46.7%, 5.1%, 5.0%, 8.0%, 38.8%, 15.9%, 0.9%, 5.4% and 0.4% better than Faster R-CNN, SSD, YOLOv3, YOLOv5, YOLOv8, Centernet, RetinaNet, CDI-YOLO, Light-PDD and RT-DETR on mAP@0.5:0.95, respectively. These results show that the improved RT-DETR algorithm has improved both detection accuracy and efficiency of printed circuit board surface defect detection.
[0077] In order to further verify the effectiveness of each application strategy in this algorithm model, this section conducts an ablation experiment. In the experiment, the original RT-DETR is represented as model (1), RT-DETR + DySample is represented as model (2), RT-DETR + DySample + RetBlock is represented as model (3), and RT-DETR + DySample + RetBlock + DA-ACFN is represented as model (4). (“+” indicates the introduction of this module. The results are shown in Table 2.) Table 2 Comparison of various indicators in ablation experiments
[0078] At the same time, we adopted a more rigorous statistical analysis to support the conclusion of performance improvement. We conducted a more detailed statistical verification of the model performance. We used the K-fold cross-validation method, divided the data set into 5 parts almost equally, and conducted 5 experiments. For the important indicator mAP in the field of object detection, we obtained the following experimental results: 0.90, 0.95, 0.94, 0.97 and 0.92. Based on these experimental results, we further calculated the mean (0.936), standard deviation (0.026) and 95% confidence interval ([0.911, 0.961]). From these statistical results, it can be seen that the performance of our model on different training / testing partitions is relatively stable, and the fluctuation range of mAP is small, indicating that the performance of the model has high reliability. At the same time, these results further support the effectiveness of the improvement method we proposed.
[0079] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A printed circuit board surface defect detection method based on deep learning, including an RT-DETR network algorithm, characterized in that: The deep learning method includes the following steps: Step 1: Improve the RT-DETR network algorithm, specifically: Step 1.1: Introduce the DySample module in the CCFM sampling stage of the RT-DETR network algorithm, and enable the CCFM sampling stage to perform dynamic upsampling; Step 1.2: Replace the original RepC3 module in the RT-DETR network algorithm with the RetBlockC3 module that adopts the RetBlock structure; Step 1.3: Integrate the DA-ACFN module into the RT-DETR network to obtain the improved RT-DETR network; Step 2: Use the printed circuit board surface defect dataset to train the improved RT-DETR network algorithm to obtain the RT-DETR target detection model; Step 3: Use the RT-DETR object detection model for printed circuit board surface defect detection.
2. The printed circuit board surface defect detection method based on deep learning according to claim 1, characterized in that: The DA-ACFN module includes an ACFN module and a DASI module. The ACFN is optimized by the DASI module to form an improved DA-ACFN module.
3. The printed circuit board surface defect detection method based on deep learning according to claim 2, characterized in that: The ACFN module includes a feature splicing module, a first input end of the feature splicing module is connected to the downsampling module and used to input low-level features, a second input end of the feature splicing module is connected to the convolution module and used to input middle-level features, and a third input end of the feature splicing module is connected to the convolution module and used to input middle-level features; The outputs of the feature splicing module are all passed through the depthwise separable convolution module and then sequentially through the skip connection module and the 1×1 convolution module for information integration to obtain integrated information, and the integrated information is skip-connected with one output of the feature splicing module to obtain a fused feature map; The process of DASI module optimizing ACFN module is as follows: Step 1.3.1: Receive the fused feature map, and set the dimension of the fused feature map to be C × H × W; Step 1.3.2: Divide the fused feature map into four groups of features, namely low-level features, high-level features, background features, and feature selection parts; Step 1.3.3: Use the Group mechanism to divide the channels, and apply the respective channel attention strategies to the features of different groups to obtain the transformation features of the corresponding dimensions; Step 1.3.4: The transformed features are weighted after passing through the Sigmoid activation function; Step 1.3.5: The weighted features are recombined and channel-wise concatenated through channel-by-channel multiplication (×) and weighted addition (+); to obtain the final optimized feature representation; to complete the optimization of the ACFN module.
4. The printed circuit board surface defect detection method based on deep learning according to claim 2, characterized in that: In the improved RT-DETR network, after the input image is subjected to preliminary feature extraction by ConvBN and MaxPool2d, multiple BasicBlocks layers are used to gradually extract features. One output of the last BasicBlock layer is directly sent to the first feature fusion module in the backbone feature extraction part, and the other output is sent to the first DA-ACFN module after passing through the first convolutional layer, AIFI layer, and second convolutional layer in the backbone feature extraction part. The first DA-ACFN module also receives feature layers output by other BasicBlocks layers except the last BasicBlock layer; one output of the first DA-ACFN module is sent to the first feature fusion module after passing through the third convolutional layer, and the other output is sent to the Dysample layer in the feature fusion part; the CCFM layer in the feature fusion part is replaced by the second DA-ACFN module.
5. The printed circuit board surface defect detection method based on deep learning according to claim 1, characterized in that: As shown in FIG. 2 , the dynamic upsampling in step 1.1 specifically includes the following steps: Step 1.1.1: Pass the input feature map of size C×H×W through a linear projection layer to generate an offset of size 2S²×H×W; Step 1.1.2: Convert the offset to an appropriate spatial resolution, and finally calculate the final coordinates of the sampling point by adding the offset to the original grid point coordinates; Step 1.1.3: Resample the feature map using the generated sampling points to obtain the upsampled feature map.
6. The printed circuit board surface defect detection method based on deep learning according to claim 5, characterized in that: The upsampling process of the DySample module is explained by the following formula: (1) In this equation, represents the input feature map, represents the output feature map, is the original sampling grid; bilinear(x) uses bilinear interpolation to scale the input feature map by a fixed multiple and interpolate the value by the weighted average of adjacent pixels; pixel_shuffle() is a pixel rearrangement operation that rearranges the pixels of the channel dimension to the spatial dimension so that the enlarged feature map is evenly distributed in space; reshaping_sample() is used to reshape the sampled feature map.
7. The printed circuit board surface defect detection method based on deep learning according to claim 6, characterized in that: DySample includes static range factor sampling strategy and dynamic range factor sampling strategy. The static range factor sampling strategy performs bilinear interpolation by scaling by 0.25 times and then applies pixel shuffle to rearrange pixels. The dynamic range factor sampling strategy performs bilinear interpolation by scaling by 0.5 times and compensates with the dynamically calculated offset.
8. The printed circuit board surface defect detection method based on deep learning according to claim 1, characterized in that: In step 1.2: The RetBlockC3 module includes depthwise separable convolution, layer normalization, Manhattan self-attention mechanism and feedforward neural network. The input features are sent to the Manhattan self-attention mechanism after depthwise separable convolution and layer normalization. The Manhattan self-attention mechanism calculates the Manhattan distance between different positions in the feature map and assigns different weights to features at different positions according to the principle of spatial attenuation. The output of the Manhattan self-attention mechanism is normalized again through layer normalization, and the normalized features enter the feedforward neural network for feature transformation and enhancement.
9. The printed circuit board surface defect detection method based on deep learning according to claim 8, characterized in that: Output of the RetBlockC3 module for: (2) Among them, DWConv is the depth-separable convolution operation of the context local enhancement module. and Calculate the attention scores from the vertical and horizontal directions respectively; (3) (4) in, is the Manhattan self-attention calculation in the horizontal direction, that is, the attention score calculated along the horizontal direction; is the vertical Manhattan self-attention calculation, that is, the attention score calculated along the vertical direction; is the query, key, and value matrix in the horizontal direction; is a normalization function that maps values to (0, 1) and is used to calculate the attention weight. is the scaling factor, , are the horizontal and vertical attenuation matrices, respectively; and are the Manhattan distances between two tokens n and m horizontally and vertically respectively; (5) (6) Where γ is the attenuation coefficient, Xn and Xm are the horizontal coordinates of the nth and mth token in the image, respectively, and Yn and Ym are the vertical coordinates of the nth and mth token in the image, respectively.
Citation Information
Patent Citations
Target specific response attention target tracking method based on twin network
CN111291679A
Weed classification detection method and system based on YOLOv8 improved algorithm
CN118762286A
PCB surface defect detection method and system based on RT-DETR detector
CN118822989A
Automatic driving target detection method based on improved RT-DETR
CN119314144A
Classroom behavior detection method based on dimension diffusion perception adaptive fusion
CN119559691A
Cited By
Surface defect detection classification and segmentation model construction method, system and medium
CN120953187A
Control method of industrial robot
CN121061879A
Control method of industrial robot
CN121061879B
Industrial product surface defect detection method and system based on deep active learning
CN121685503A