A PCB surface defect detection system and method based on improved YOLOv10

By improving the YOLOv10 network model and adopting the Esvg, Seop and Lag modules, the ESL-YOLO detection model is constructed. This solves the problems of missed detection and false detection in existing PCB defect detection methods under small targets and complex background noise, and achieves higher detection accuracy and efficiency.

CN120125585BActive Publication Date: 2025-09-09CHANGCHUN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510611943.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-09
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Existing PCB defect detection methods suffer from high missed detection rate, high false detection rate and decreased bounding box accuracy when dealing with small targets and complex background noise.

Method used

An improved YOLOv10 network model is adopted. By replacing the C2f module in the backbone network with the Esvg module, introducing the Seop module into the P4 and P5 feature map processing of the neck network, and replacing the detection head with the Lag module, an ESL-YOLO detection model is constructed.

Benefits of technology

It improves the model's perception of complex defects, enhances the detection capability of multi-scale targets, reduces false alarm and missed detection rates, and improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125585B_ABST
    Figure CN120125585B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image data processing and relates to a PCB surface defect detection system and method based on an improved YOLOv10. The system is obtained by improving the YOLOv10 network model. The C2f module in the backbone network of the model is replaced with a multi-scale cross-stage fusion module to extract deep features of the image and capture multi-scale information in the input data. A dynamic perception feature information enhancement module is introduced into the neck network and applied to the processing of P4 and P5 feature maps. The detection head of the YOLOv10 network model is replaced with a Lag module. The Lag module performs channel adjustment through a group normalization convolution module. The data output by the shared convolution group is then distributed to the positioning convolution module and the classification convolution module through a shared convolution group. The system can detect small defects even under complex background noise interference and has the advantages of high detection accuracy and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image data processing technology and relates to printed circuit board (PCB) defect detection. Specifically, it relates to a PCB surface defect detection system and method based on an improved YOLOv10. Background Art

[0002] Printed circuit boards (PCBs) are widely used in electronic products due to their high level of integration, lightweight design, low cost, and high stability. However, during the manufacturing process, various factors, including equipment failure, operational errors, and design flaws, can lead to various minor defects on PCBs. If these defects are not detected and addressed promptly, they can seriously affect the performance, quality, and safety of the PCBs.

[0003] PCB defects vary widely, including solder joint defects, missing components, short circuits, and open circuits. The goal of PCB defect detection is to automatically detect these defects through image analysis, thereby achieving efficient and accurate product quality control. Early PCB defect detection methods primarily relied on traditional image processing algorithms, including edge detection, morphological analysis, and template matching. These methods are effective for simple defects but often perform poorly for images with complex backgrounds and noise.

[0004] With the development of deep learning, especially convolutional neural networks (CNNs), the performance of tasks such as image classification and object detection has significantly improved. In PCB defect inspection, deep learning-based object detection algorithms include two-stage and one-stage object detection frameworks. In the two-stage detection algorithm, R-CNN is a region-based convolutional neural network that uses a selective search algorithm to extract potential candidate regions from the input image. It then extracts features from each candidate region and classifies them using a support vector machine (SVM). This transforms the object detection problem into a candidate region classification task. Faster R-CNN improves on R-CNN by introducing a region proposal network (RPN), enabling end-to-end training of the entire object detection system. However, limitations in the candidate region generation process can lead to missed detections or inaccurate localization. In the one-stage algorithm, SSD (Single Shot MultiBox Detector) employs a convolutional neural network (CNN)-based feature extractor and performs object detection on feature maps at multiple scales. It detects objects of varying sizes by applying convolution kernels of different sizes to feature maps at different levels. This multi-scale detection strategy enables SSD to effectively detect objects of varying sizes. YOLO (You Only Look Once), an end-to-end object detection algorithm, achieves high detection accuracy while maintaining speed and is widely used in multiple visual inspection fields. YOLO is also a single-stage object detection algorithm. Unlike SSD, YOLO transforms the object detection problem into a regression problem. It divides the input image into a fixed-size grid and predicts the bounding box and class of the object in each grid cell. While all of the aforementioned methods can detect surface defects on printed circuit boards, their feature extraction capabilities are insufficient when dealing with small objects. This causes minor defects (such as solder cracks and small shorts) to be weakened in deep feature maps, resulting in a high rate of missed detection. Furthermore, in the presence of complex background noise, the model struggles to distinguish minor defects from noise, increasing the false detection rate. Furthermore, the traditional YOLO detection head lacks contextual modeling of defects, resulting in the merging of densely packed defects, reduced bounding box accuracy, and low recall for defects with long-tail distributions. Summary of the Invention

[0005] In view of the above technical problems and defects, the object of the present invention is to provide a PCB surface defect detection system based on an improved YOLOv10. The system adopts a trained ESL-YOLO detection model. The ESL-YOLO detection model is obtained by improving the YOLOv10 network model. The C2f module in the backbone network of the YOLOv10 network model is replaced with the Esvg module. The Seop module is introduced into the neck network of the YOLOv10 network model and applied to the processing of the P4 and P5 feature maps. The detection head of the YOLOv10 network model is replaced with Lag, thereby completing the construction of the ESL-YOLO detection model. The trained ESL-YOLO detection model is then used to identify defects in printed circuit boards. The system can also detect minor defects under complex background noise interference, and has the advantages of high detection accuracy and efficiency.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A PCB surface defect detection system based on improved YOLOv10 is disclosed. The system uses a trained ESL-YOLO detection model, which is obtained by improving the YOLOv10 network model. The last two C2f modules in the backbone network of the YOLOv10 network model are replaced with Esvg modules. In the neck network of the YOLOv10 network model, a Seop module is introduced during the processing of the P4 and P5 feature maps of multi-scale feature fusion. The detection head of the YOLOv10 network model is replaced with a Lag module.

[0008] Among them, the Esvg module is a multi-scale cross-stage fusion module used to extract deep features of the image and capture multi-scale information in the input data. The Esvg module divides the data into two branches. One branch extracts multi-scale features of the image through n ESBlock modules, and the other branch retains the original features. The original features and the multi-scale features of the image are then spliced ​​together and finally output after convolution processing. The ESBlock module uses a multi-branch structure to process the input feature map.

[0009] The Seop module is a dynamic perception feature information enhancement module that processes the data in two paths. One path is output after passing through the average pooling layer, convolution layer, and Softmax function; the other path is output after multi-branch refinement processing of the features through the grouped convolution module; finally, the two outputs are feature weighted and the weighted downsampled feature map is output;

[0010] The Lag module performs channel adjustment through the group normalization convolution module, and then distributes the data output by the shared convolution group to the positioning convolution module and the classification convolution module through the shared convolution group; among them, the positioning convolution module is responsible for regression, and the classification convolution module is responsible for classification.

[0011] As a preferred embodiment of the present invention, the method for extracting feature information by the Esvg module is: by using 1×1 convolution to change the channel number i of the input feature map to the specified output channel number m; then, the input is divided into two branches with a channel number of 0.5×m through a Split operation; wherein, the first branch extracts multi-scale features of the image through n ESBlock modules, and the second branch retains the original features; then the ESBlock module is spliced ​​with the second branch through n connections to obtain a feature map with a channel number of (n+2)×0.5m, and finally the channel number is changed to m through 1×1 convolution to obtain the output.

[0012] As a preferred embodiment of the present invention, the ESBlock module includes a first convolutional layer and a second convolutional layer; wherein, the first convolutional layer adopts 3×3 standard convolution for processing, and the second convolutional layer adopts a multi-scale convolution module.

[0013] As a preferred embodiment of the present invention, the data processing flow of the Seop module is as follows: after adjusting the number of channels through 3×3 global average pooling and 1×1 convolution, (bs, ch, 2h, 2w) is reorganized into (bs, ch, h, w, 4) through a dimension reorganization operation; then normalized using the Softmax function; the other path is downsampled by 2 times through a grouped convolution module to halve the size of the feature map and expand the number of channels to 4 times; then, the output is reshaped into (bs, ch, h, w, 4) through a dimension reorganization operation; where bs is the batch size, ch is the number of channels, h is the height of the feature map, and w is the width of the feature map.

[0014] As a preferred embodiment of the present invention, the shared convolution group includes two 3×3 group normalized convolution modules. The Lag module uses feature maps of three different scales, P3, P4, and P5, in the neck network as input. Channel adjustment is first performed through a 1×1 group normalized convolution module; then, after fusion, it enters a 3×3 group normalized convolution module to further extract spatial information and form a shared feature representation; subsequently, the shared features are passed back to the 3×3 group normalized convolution module to enhance the feature representation, and are assigned to two different task branches, the positioning convolution module and the classification convolution module.

[0015] As a further preference of the present invention, the data processing flow of the multi-scale convolution module is as follows: first, the input feature map is divided into three parts in proportion through the Split operation, of which 50% of the feature maps are directly retained as the main branch, 25% of the feature maps are processed through a 3×3 convolution layer, and the remaining 25% of the feature maps are processed through a 5×5 convolution layer for feature extraction; subsequently, the output feature maps of the three branches are spliced ​​in the channel dimension, and feature fusion is performed through a 1×1 convolution.

[0016] As a further preferred embodiment of the present invention, the grouped convolution module divides the feature map into 4 groups, each group independently performs a 3×3 convolution operation, and controls the number of channels ch=64 and stride=2.

[0017] As a further preferred embodiment of the present invention, the positioning convolution module and the classification convolution module adopt 1×1 convolution, and the output of the positioning convolution module is scale-adaptively calibrated through a Scale function.

[0018] The present invention also provides a PCB surface defect detection method based on improved YOLOv10, which includes the following steps:

[0019] Step 1. Build the ESL-YOLO detection model described above;

[0020] Step 2. Train the ESL-YOLO detection model to obtain a trained ESL-YOLO detection model;

[0021] Step 3. Input the PCB image data into the trained ESL-YOLO detection model for surface defect detection, thereby identifying the location and category of surface defects in the PCB image.

[0022] As a preferred embodiment of the present invention, the process of training the ESL-YOLO detection model is as follows:

[0023] Step a. Obtain a public dataset of PCB surface defect images, perform image enhancement and preprocessing, and construct a PCB surface defect dataset; then divide the PCB surface defect dataset into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0024] Step b. Standardizing or normalizing the training set, validation set, and test set to form a PCB surface defect image dataset that meets the requirements;

[0025] Step c. Configure the training parameters, setting the batch size to 32 and the number of iterations to 200. Train the ESL-YOLO detection model using the training set data. In each training iteration, calculate the loss and update the ESL-YOLO detection model weights. After each training cycle, use the validation set data to evaluate the performance of the ESL-YOLO detection model, monitoring the loss and accuracy on the validation set data. Also, enable Mosaic for data augmentation during ESL-YOLO detection model training and disable Mosaic for the last 10 rounds of training.

[0026] As a preferred embodiment of the present invention, multi-scale prediction paths are optimized simultaneously in the training phase, and bounding box regression and classification prediction are jointly optimized; in the decoding and post-processing phase, a Top-K mechanism is used to select high-confidence boxes, and redundant boxes are removed through non-maximum suppression to generate the final detection results.

[0027] Advantages and beneficial effects of the present invention:

[0028] (1) The present invention replaces the last two C2f modules in the backbone network of the YOLOv10 network model with an Esvg (multi-scale cross-stage fusion module) module. The Esvg module includes multiple ESBlock modules, and the ESBlock module includes a multi-scale convolution (EMSConv) module. The Esvg module can more effectively extract detail features, enhance the model's perception of complex defects, improve the extraction of deep features of images, and capture multi-scale information in input data; its multi-branch structure and residual connection design make the feature extraction process more efficient, and can improve the model's ability to capture details while maintaining computational efficiency.

[0029] (2) The Esvg module in the present invention can effectively process features of different scales by introducing multi-scale convolution (EMSConv), making the model more sensitive to defects of various sizes, enhancing the network model's perception of features of different scales (defects of various sizes), and enabling it to more efficiently capture multi-scale information in the input data; through the multi-branch residual connection of the EMSConv module, the expression ability of fine-grained features is enhanced, making the model more sensitive to the recognition of complex patterns, effectively dealing with subtle defects, helping to improve the overall detection and classification performance, and solving the problem that traditional feature extraction methods are difficult to capture these multi-scale information, especially when the defects are very similar to the background, which is prone to missed detection or false alarms; in addition, when the Esvg module extracts feature information, it adjusts the number of channels through 1×1 convolution and splits the feature map through the Split branch to enhance the efficiency and diversity of feature extraction, ensuring that the model can still accurately detect various defects under complex backgrounds; such a design improves the robustness and accuracy of the model when facing complex and diverse PCB surface defects, and significantly reduces false alarms and missed detections.

[0030] (3) The present invention introduces the Seop (dynamic perception feature information enhancement) module into the neck network of the YOLOv10 network model and applies it to the processing of P4 and P5 feature maps. Seop extracts global information through average pooling, and then generates channel attention weights through convolution operation. At the same time, it uses grouped convolution to extract local features, effectively fusing global context information and local detail information, significantly improving the expressiveness of feature representation and the model's ability to detect defects of different scales. In particular, when dealing with subtle differences and multi-scale targets, it can enhance the model's ability to recognize details, making it easier for the model to understand the context of the image, and solving the problem that traditional convolutional neural networks are insufficient in fusing fine-grained and contextual information when dealing with complex scenes. When there is a lot of background noise and complex circuit layout in the PCB image, it is difficult to distinguish between defects and background, especially when the defects are similar to the background, the model cannot accurately detect PCB defects.

[0031] (4) The adaptive fusion mechanism of the Seop module of the present invention enables the model to retain key detail information while reducing the resolution, significantly improving the model's ability to maintain complex textures and edge features, thereby improving detection accuracy.

[0032] (5) In order to improve the detection accuracy of the model in multi-scale detection tasks, the present invention effectively integrates feature maps of different scales and accurately locates and classifies targets. The detection head of the network model is Lag. The fusion of multi-scale features and the design of the detection head improve the model's ability to detect defects of different sizes and increase the model's flexibility. The number of parameters can be reduced, which makes the model more lightweight and can also improve the performance of the detection head positioning and classification. This solves the problem that traditional detection heads are difficult to handle complex and variable circuit board defects, especially when dealing with defects of different sizes and shapes, which are prone to false detection or missed detection.

[0033] (6) The detection head of the present invention adopts the Lag module to optimize the prediction process of multi-scale features. By introducing the dual-path strategy of "one-to-many prediction" and "one-to-one prediction", the Lag detection head can use the advantages of dense priors and precise positioning characteristics to achieve knowledge complementarity during training, and balance the recall rate and accuracy through logical integration during reasoning; its dynamic scale calibration and multi-dimensional topological sorting mechanism enable the model to locate and classify defects more accurately, and finally output standardized detection results, ensuring the compatibility and efficiency of the model at the deployment end. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] By referring to the following description in conjunction with the accompanying drawings, and with a more complete understanding of the present invention, other objects and results of the present invention will become more clear and easy to understand. In the accompanying drawings:

[0035] Figure 1 Flowchart of the printed circuit board (PCB) surface defect detection method provided by the present invention;

[0036] Figure 2 The framework diagram of the ESL-YOLO detection module designed for the present invention; wherein Conv represents convolution, C2f represents cross-stage dual convolution bottleneck, SCDown represents spatial channel joint downsampling, Esvg represents multi-scale cross-stage fusion, SPPF represents spatial pyramid fast pooling, PSA represents pyramid split attention, Upsample represents upsampling, Concat represents splicing, Seop represents dynamic perception feature information enhancement, C2fCIB represents inverted residual feature fusion, and Lag represents the detection head;

[0037] Figure 3 The following is a schematic diagram of the structure of the Esvg module; Conv stands for convolution, Split stands for segmentation, Concat stands for concatenation, Input stands for input, and EMSConv stands for multi-scale convolution;

[0038] Figure 4 The diagram below shows the structure of the Seop module. Input represents input, AvgPool represents average pooling, Conv represents convolution, Softmax represents the Softmax function, and GroupConv represents group convolution.

[0039] Figure 5 The figure is a schematic diagram of the Lag module structure; P3, P4, and P5 represent feature maps of different scales, GNConv represents group normalization convolution, LoConv represents positioning convolution, ClsConv represents classification convolution, and Sclae represents dynamic scale normalization;

[0040] Figure 6 Flowchart for training the ESL-YOLO detection module. DETAILED DESCRIPTION

[0041] In order to enable those skilled in the art to better understand the technical solutions and advantages of the present invention, the present application is described in detail below with reference to the accompanying drawings, but this is not intended to limit the scope of protection of the present invention.

[0042] Example 1:

[0043] This embodiment provides a printed circuit board (PCB) surface defect detection system based on improved YOLOv10. The technical solution of the present invention is described in detail below with reference to the accompanying drawings.

[0044] like Figures 2 to 4As shown, this embodiment provides a printed circuit board (PCB) surface defect detection system based on improved YOLOv10. The system uses a trained ESL-YOLO detection model. The ESL-YOLO detection model is obtained by improving the YOLOv10 network model. The last two C2f modules in the backbone network of the YOLOv10 network model are replaced with Esvg modules. In the neck network of the YOLOv10 network model, the Seop module is introduced in the P4 and P5 feature map processing links of multi-scale feature fusion. The head network (detection head) of the YOLOv10 network model is replaced with the Lag module.

[0045] Among them, the Esvg module is a multi-scale cross-stage fusion module used to extract deep features of the image and capture multi-scale information in the input data. The Esvg module divides the data into two branches. One branch extracts multi-scale features of the image through n ESBlock modules, and the other branch retains the original features. The original features and the multi-scale features of the image are then spliced ​​together and finally output after convolution processing. The ESBlock module uses a multi-branch structure to process the input feature map.

[0046] The Seop module is a dynamic perception feature information enhancement module that processes the data in two paths. One path is output after passing through the average pooling layer, convolution layer, and Softmax function; the other path is output after multi-branch refinement processing of the features through the grouped convolution module; finally, the two outputs are feature weighted and the weighted downsampled feature map is output;

[0047] The Lag module performs channel adjustment through the group normalization convolution module, and then distributes the data output by the shared convolution group to the positioning convolution module and the classification convolution module through the shared convolution group; among them, the positioning convolution module is responsible for regression, and the classification convolution module is responsible for classification.

[0048] Specifically, the specific construction method of the ESL-YOLO detection model is:

[0049] First, obtain the original YOLOv10 network model. The backbone network of the original YOLOv10 network model has four C2f modules. Replace the last two C2f modules in the backbone network of the YOLOv10 network model with Esvg modules. The design of the Esvg module can improve the model's ability to perceive details.

[0050] The Seop module was innovatively introduced into the neck network design of the YOLOv10 network model and deployed in the P4 and P5 feature map processing stages, which are key levels for multi-scale feature fusion. The YOLOv10 input image size is typically 640×640. After 16x and 32x downsampling, feature maps of 40×40 and 20×20 are obtained, respectively. P4 corresponds to the 40×40 feature map obtained by 16x downsampling, and P5 corresponds to the 20×20 feature map obtained by 32x downsampling.

[0051] The detection head of the YOLOv10 network model is replaced with the Lag module. The fusion of multi-scale features in the Lag module and the design of the detection head enhance the model's ability to detect defects of different sizes, reduce the number of parameters, and increase the model's flexibility, thereby improving the accuracy and robustness of PCB surface defect detection.

[0052] Further, if Figure 3 As shown, in this embodiment, the method of extracting feature information by the Esvg module is as follows: first, the number of channels i of the input feature map is changed to the number of output channels m specified in the model configuration file by using 1×1 convolution. In this implementation, m=512 is set to establish a learnable feature screening mechanism; then, the input is divided into two branches with a channel number of 0.5×m through a Split operation. The first branch extracts multi-scale features of the image through n (specifically 6) ESBlock modules, and the second branch retains the original features to prevent information loss. This branch can improve computational efficiency. The ESBlock module is then spliced ​​with the second branch (original features) through n connections to obtain a feature map with a channel number of (n+2)×0.5m. The number of feature reuses of the residual connection enhances the feature expression capability. Finally, the number of channels is changed to m through a 1×1 convolution to obtain the output.

[0053] In this embodiment, the second branch channel number of ESBlock Calculation formula: ; where c1 is the number of output channels, e is the expansion coefficient, Represents the rounding function.

[0054] For each kernel size ks, the convolution operation It can be expressed as: ; Where x is the input feature map, is the minimum number of channels, ks is the convolution kernel size; 1x1 convolution operation : ; Where x is the input feature map, channel is the number of input and output channels, and 1 represents the convolution kernel size of 1x1.

[0055] In this embodiment, the ESBlock module uses a multi-branch structure to process the input feature map. The ESBlock module includes a first convolution layer and a second convolution layer. The first convolution layer uses a standard convolution layer (convolution kernel size is 3), and performs preliminary processing through the standard convolution layer to extract basic feature information; the second convolution layer uses a multi-scale convolution (EMSConv) module, and the features output by the first convolution layer are passed to the multi-scale convolution (EMSConv) module, which is responsible for processing features of different scales and enhancing sensitivity to defects of various sizes; finally, through the fusion and residual connection of the ESBlock module, a richer and more diverse feature representation is output (when shortcut=True, a residual connection is performed between the input and output data of the ESBlock module).

[0056] Among them, the data processing flow of the multi-scale convolution (EMSConv) module is as follows: first, the input feature map is divided into three parts in proportion through the Split operation, of which 50% of the feature maps are directly retained as the main branch, 25% of the feature maps are processed through a 3×3 convolution layer (a convolution layer with a convolution kernel of 3), and the remaining 25% of the feature maps are processed through a 5×5 convolution layer (a convolution layer with a convolution kernel of 5) for feature extraction; then, the output feature maps of the three branches are spliced ​​in the channel dimension and feature fused through a 1×1 convolution, thereby enhancing the multi-scale feature expression capability and realizing the enhancement and integration of multi-scale features.

[0057] Specifically, in this embodiment, the specific workflow and formula of the ESBlock module are as follows:

[0058] Input feature map: shape is (bs, c1, h, w); where bs (batch size) represents the number of input samples, c1 represents the number of channels of the input feature map, h represents the height of the input feature map, and w represents the width of the input feature map.

[0059] First convolutional layer (cv1): , the output shape is (bs, c3, h, w), c3 = int(c2× e); where x is the input feature map, bs is the batch size (the number of input samples), h is the height of the input feature map, and w is the width of the input feature map. Represents the first layer of convolution operation, the convolution kernel size is k[0], and in this embodiment, 3×3 convolution is used for processing. The output feature map of the first convolution layer, where c2 and e are related parameters. Usually e is a hyperparameter or empirical value used to adjust the number of channels, and c3 is the final output channel number.

[0060] Second convolutional layer (cv2): , the output shape is (bs, c4, h, w); where bs is the batch size, h is the height of the input feature map, w is the width of the input feature map, and c4 is the number of output channels. is the output feature map of the first convolution layer, represents multi-scale convolution, is the output feature map of the second layer of convolution;

[0061] Residual Connection: ; If shortcut=True, the input x and output Add: Otherwise, output directly , where x is the input feature map, The output feature map of the second convolution layer, Represents the output feature map after residual connection, shortcut=True means performing residual connection.

[0062] In this embodiment, the specific working process of the EMSConv module is as follows:

[0063] The EMSConv module receives a feature map of shape (bs, c, h, w); where bs is the batch size, c is the number of channels, h and w are the height and width of the input feature map, and the Split operation divides the input feature map into two parts along the channel dimension. The main branch ( ) and multi-scale branches ( x group ); Among them, the main branch ( ) retains the first 50% of the channels, with a shape of (bs, c / / 2, h, w), It means that the input channel number c is divided into two parts evenly; this part of the feature map is directly retained without any processing to reduce the amount of calculation;

[0064] Multi-scale branches (x group ) is the second part of the feature map after segmentation, retaining the last 50% of the channels, with a shape of (bs, c / / 2, h, w). This part of the feature map will be used for multi-scale feature extraction. The grouping and rearrangement operation is to group the multi-scale branches x group It is further divided into g groups, where g is the number of convolution kernels. The shape of the feature map after grouping is (bs, c / / (2g), h, w, g); where c / / (2g) represents the number of channels of each feature map after grouping, and / / represents integer division.

[0065] After that, multi-scale convolution performs convolution operations of different scales on each group of feature maps. The i-th group of feature maps is processed by a convolution layer with a convolution kernel size of kernels[i]. Multi-scale feature splicing splices the output feature maps of all groups in the channel dimension to obtain a multi-scale fused feature map. The shape of the spliced ​​feature map is (bs, c / / 2, h, w);

[0066] Then, the main branch is fused with the multi-scale branch to perform feature splicing: the main branch and multi-scale branches x group Splice in the channel dimension to get a complete feature map. The shape of the spliced ​​feature map is (bs, c, h, w);

[0067] Finally, 1x1 convolution is used to perform channel fusion on the spliced ​​feature map. The role of 1x1 convolution is to adjust the number of channels to make it consistent with the number of input channels. It enhances the interaction between multi-scale features and further improves the feature expression ability. After the output passes through 1x1 convolution, the shape of the output feature map is still (bs, c, h, w). The final output feature map contains the fusion result of the original features and multi-scale features, has stronger expression ability, and maintains high computational efficiency.

[0068] Specifically, the input feature map x is split along the channel dimension: ; Where x is the input feature map, It means to divide the input channel number c into two parts evenly. It is the first part of the feature map after segmentation. Split represents segmentation. It is the second part of the feature map after segmentation; dim=1 means operation along the channel dimension;

[0069] The second part of the feature map after segmentation Perform grouping and rearrangement, and the feature map after grouping and rearrangement The expression is:

[0070] ;in, Represents the second part of the feature map after segmentation, Indicates that the shape of the feature map is transformed from (bs, c / / 2, h, w) to (bs, ch, h, w, g), where bs is the batch size, c / / 2 represents half the number of channels of the feature map, h is the height of the feature map, w is the width of the feature map, g is the number of groups, ch is the number of channels in each group, ch=c / / (2g), and rearrange represents the rearrangement of the feature map;

[0071] Perform group convolution, feature map after group convolution The expression is:

[0072] ; Among them, self.convs is a list of multi-scale convolutional layers, each convolutional layer corresponds to a kernel size ks, [..., i] represents the The i-th group performs convolution operation, torch.stack represents stacking the convolution results of each group together, Represents stacking tensors in a list together along a new dimension; Represents the number of groups for group convolution;

[0073] Recombined feature map, combined feature map The expression is:

[0074] ;in, is the feature map after group convolution, with a shape of (g, bs, ch, h, w), where g is the number of groups, bs is the batch size, ch is the number of channels, h is the height of the feature map, and w is the width of the feature map. rearrange represents merging the group dimension g and the channel dimension ch to restore the original shape (bs, g × ch, h, w); g × ch represents merging into a new dimension, bs is the batch size, ch is the number of channels, h is the height of the feature map, and w is the width of the feature map.

[0075] Feature splicing, the expression of the spliced ​​feature map X is:

[0076] ;in, is the first part of the feature map after segmentation, The recombined feature map, dim=1 means operation along the channel dimension; torch.cat means concatenating two feature maps along the channel dimension;

[0077] Finally, after 1x1 convolution, the output feature map Y is expressed as:

[0078] ; Where X is the concatenated feature map, Represents the convolutional layer, which is used to fuse information between channels.

[0079] Further, if Figure 4As shown, in this embodiment, the data processing flow of the Seop (dynamic perception feature information enhancement) module is as follows: the input data is first compressed by the AvgPool (average pooling) layer to extract local context information while maintaining the resolution and reducing the amount of calculation; then it enters the Conv (convolution) layer to extract key spatial features and align the channel dimensions; then the features are converted into probability distributions through the Softmax function to obtain the attention weight matrix that characterizes the importance of each sampling point; at the same time, the input data passes through the GroupConv (group convolution) module to perform multi-branch refinement processing on the features to enhance the model's ability to capture different feature patterns; finally, the two outputs are feature weighted to output dimensional features.

[0080] Specifically, the input feature x is first processed in a dual-path manner, where the attention branch uses 3×3 global average pooling (stride=1, padding=1) to maintain the spatial size, and adjusts the number of channels through 1×1 convolution to adapt the features to subsequent calculations; then, a dimension reorganization (rearrange) operation is used to reorganize (bs, ch, 2h, 2w) into (bs, ch, h, w, 4), that is, mapping the 2×2 neighborhood to the fifth dimension to form four candidate downsampling points; in addition, to guide adaptive feature selection, the channel is normalized using the Softmax (dim=-1) function so that the sum of the weights of the four candidate points is 1, ensuring normalized weighted contributions when features are fused;

[0081] The GroupConv branch performs a 2x downsampling through 3×3 convolution (stride=2), halving the size of the feature map and expanding the number of channels to 4 times (bs, 4ch, h / 2, w / 2). Specifically, the feature map is divided into 4 groups, and each group performs an independent convolution operation (the convolution kernel used is 3×3 convolution, the number of channels is controlled to ch=64, and stride=2), halving the size of the feature map (h and h), thereby reducing the amount of computation and memory consumption and maintaining lightweight. In this invention, the grouped convolution maintains a low computational complexity by limiting the number of channels of each group of convolution operations, while making feature extraction more efficient. Then, the output is reshaped into (bs, ch, h, w, 4) through a dimensionality reorganization (rearrange) operation, that is, each pixel corresponds to four candidate feature maps, and is multiplied (point by point) with the weight matrix of the attention branch to achieve feature weighting. Finally, the sum (dim=-1) function is used to sum along the fifth dimension to output the refined (bs, ch, h, w) dimensional features; compared with the traditional fixed-rule downsampling method, this mechanism can dynamically adjust the feature contribution of different spatial positions, and better preserve key texture and edge information while downsampling.

[0082] In this embodiment, the Seop module effectively improves the expressiveness of feature representation by combining global information and local detail features. Global information is extracted through global average pooling. The process is as follows:

[0083] Assume the input feature map is ; Where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively;

[0084] Attention weight calculation:

[0085] The features obtained by global average pooling are ;

[0086] The features obtained using the convolutional layer are ;

[0087] right Apply Softmax operation to get attention weight ; Among them, 4 represents the feature dimension of each position;

[0088] Group Convolution (GroupConv):

[0089] The downsampled features obtained by grouped convolution are ;

[0090] Use rearrange to convert it to , in order to perform weighted operations;

[0091] After weighted summation, the final output is: ;in, ;

[0092] The final output is , represents the weighted down-sampled feature map.

[0093] Further, if Figure 5 As shown, in this embodiment, the Lag module first performs unified representation learning on the multi-scale feature maps from the neck network; performs channel dimension alignment through the layer-by-layer GNConv convolution kernel (self.conv module), and projects the features of different levels into a unified latent space dimension hidc; then the shared convolution group self.share_conv (including dual 3×3 group normalized convolution) performs cross-level contextual information interaction and spatial detail enhancement on the features. This design achieves cross-scale semantic consistency modeling while maintaining the independence of features at each level, providing a highly discriminative feature basis for subsequent multi-task prediction;

[0094] Specifically, the Lag module uses feature maps from the neck network at three different scales: P3, P4, and P5. These are first channel-wise adjusted through a 1×1 GNConv (Grouped Normalized Convolution) module to reduce computation and enhance feature representation. These features are then fused into a 3×3 GNConv module to further extract spatial information and form a shared feature representation. This shared feature is then passed back to the 3×3 GNConv module to enhance the feature representation and assigned to two distinct task branches: the LoConv (Localization Convolution) module and the ClsConv (Classification Convolution) module. The LoConv module is responsible for regression (generating bounding box regression parameters), while the ClsConv module is responsible for classification (predicting class confidence). Furthermore, the LoConv output is adaptively scaled using a Scale function (dynamic scale normalization) to enhance the model's generalization across different object scales.

[0095] In this embodiment, the self.conv module processes input feature maps of different scales through GNConv, corresponding to Figure 5 The GNConv in the model is 1×1; the shared convolution group self.share_conv consists of two 3×3 GNConvs, consistent with the GNConv 3×3 structure. The regression and classification branches are implemented by a 1×1 Conv (LoConv) and a 1×1 Conv (ClsConv), respectively, and are dynamically rescaled using Scale. The classification branch (ClsConv) generates object class probabilities for each position, while the regression branch (LoConv) generates bounding box regression parameters and normalizes the output using a scale factor. During forward propagation, the "one-to-many prediction" and "one-to-one prediction" paths are optimized in parallel during training. During inference, decode_bboxes() decodes the distribution focus using Decentralized Flow (DFL) decoding and uses a post-processing Top-K mechanism to initially filter detection results, ensuring that the model outputs candidate boxes focused on the most likely locations of the object. Based on these candidate boxes, NMS further removes overlapping boxes and retains the optimal one, ultimately resulting in a detection result that achieves accurate detection. Specifically, decoding and post-processing use distributed focus regression (DFL) and a top-k decoding mechanism to generate the final bounding box. Predicted categories are normalized, and non-maximum suppression (NMS) is used to generate the final detection result. This invention introduces an adaptive weight adjustment downsampling operation, enhancing the model's flexibility and information retention during the downsampling process.

[0096] Example 2:

[0097] This embodiment provides a printed circuit board (PCB) surface defect detection method based on improved YOLOv10. The technical solution of the present invention is described in detail below with reference to the accompanying drawings.

[0098] Figure 1 The flow chart of the printed circuit board surface defect detection method provided by the present invention is as follows: Figure 1 As shown, the present invention provides a printed circuit board (PCB) surface defect detection method based on improved YOLOv10, comprising the following steps:

[0099] Step 1. Construct an ESL-YOLO detection model. The ESL-YOLO detection model is obtained by improving the YOLOv10 network model and includes a backbone network, a neck network, and a head network (detection head). The backbone network is used to extract deep features of the image, the neck network is responsible for enhancing the fusion capability of cross-scale features, and the detection head is used to accurately predict the defect location and category based on the extracted features, thereby improving the detection accuracy and robustness of PCB surface defects. The network structure of the ESL-YOLO detection model is as described in Example 1.

[0100] Step 2. Train the ESL-YOLO detection model. Through multiple rounds of iterative optimization, an efficient and accurate trained ESL-YOLO detection model is finally obtained.

[0101] Step 3. Input the PCB image data into the trained ESL-YOLO detection model for surface defect detection, thereby identifying the location and category of surface defects in the PCB image.

[0102] like Figure 6 As shown, in this embodiment, the process of training the ESL-YOLO detection model in step 2 is:

[0103] Step a. Obtain a public dataset of PCB surface defect images (obtained from a publicly available PCB defect detection dataset from Peking University). The dataset includes various types of defect samples, such as open circuits, short circuits, and cold solder joints. Then, perform image enhancement and preprocessing to construct a preprocessed and image-enhanced PCB image dataset. The preprocessed and image-enhanced PCB image dataset is then divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0104] Step b. Standardize or normalize the training set, validation set, and test set to form a PCB surface defect dataset that meets the requirements;

[0105] Step c. Configure the training parameters, set the batch size to 32 and the number of iterations to 200 rounds; train the constructed ESL-YOLO detection model using the training set data. In each training iteration, calculate the loss and update the ESL-YOLO detection model weights. After each training cycle, use the validation set data to evaluate the performance of the ESL-YOLO detection model. Monitor the loss and accuracy on the validation set data so that training can be stopped in time to prevent overfitting. At the same time, enable Mosaic for data augmentation during ESL-YOLO detection model training and disable Mosaic for the last 10 rounds of training to improve the generalization ability of the ESL-YOLO detection model.

[0106] Finally, the performance of the trained ESL-YOLO detection model is evaluated using a test set. Evaluation metrics include detection accuracy, recall rate, and mAP to ensure improved model accuracy and detection accuracy and efficiency. Finally, the detection results are output to obtain the location (positioning) and classification (category information) of defect targets in PCB surface defect images.

[0107] The solution provided in this embodiment ensures the diversity and representativeness of the data set, helps the model to fully learn different types of defect characteristics during training, and at the same time, performs parameter tuning through the validation set. Finally, the generalization ability of the model is evaluated through the test set, which accelerates the model convergence speed, reduces the learning difficulty caused by inconsistent data scales during training, and avoids overfitting.

[0108] In order to verify the effectiveness of the ESL-YOLO model designed by the present invention, the present invention compares it with other existing models. The test results of each model are shown in Table 1.

[0109] Table 1 Test results of each model

[0110] Model Accuracy Recall mAP50 mAP50-90 Paramaters (M) Faster-RCNN 0.874 0.826 0.884 0.526 45.2 SSD 0.851 0.823 0.866 0.614 24.5 YOLOv8 0.922 0.906 0.916 0.497 3.2 YOLOv10 0.926 0.911 0.920 0.507 2.3 ESL-YOLO 0.958 0.926 0.952 0.529 2.3

[0111] Based on the comparative analysis of experimental data in Table 1, the proposed ESL-YOLO model improves upon existing methods like YOLOv10 in key metrics such as precision, recall, and mAP50 for printed circuit board defect detection, while maintaining its lightweight design with a model parameter count of 2.3M. Experimental results demonstrate that the proposed model achieves a good balance between achieving an mAP50-90 of 0.529 for multi-scale defect detection and high computational efficiency, with minimal parameter computation and high detection efficiency.

[0112] The present invention also provides an electronic device, comprising: one or more processors and a memory; wherein the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned printed circuit board (PCB) surface defect detection method based on the improved YOLOv10.

[0113] The present invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned printed circuit board (PCB) surface defect detection method based on the improved YOLOv10.

[0114] Those skilled in the art will appreciate that all or part of the functions of the various methods / modules in the above embodiments may be implemented via hardware or via computer programs. When all or part of the functions in the above embodiments are implemented via computer programs, the program may be stored in a computer-readable storage medium, which may include a read-only memory, random access memory, a magnetic disk, an optical disk, a hard disk, etc., and the program is executed by a computer to implement the above functions. For example, the program may be stored in a memory of a device, and when the program in the memory is executed by a processor, all or part of the above functions may be implemented.

[0115] In addition, when all or part of the functions in the above-mentioned embodiments are implemented by means of a computer program, the program can also be stored in a storage medium such as a server, another computer, a disk, an optical disk, a flash drive or a mobile hard disk, and saved to the memory of a local device by downloading or copying, or the system of the local device is updated. When the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.

[0116] The above description of the present invention using specific examples is intended only to facilitate understanding of the present invention and is not intended to limit the present invention. A person skilled in the art of the present invention may make several simple deductions, modifications, or substitutions based on the principles of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A PCB surface defect detection system based on improved YOLOv10, characterized in that: The system uses a trained ESL-YOLO detection model, which is obtained by improving the YOLOv10 network model. The last two C2f modules in the backbone network of the YOLOv10 network model are replaced with Esvg modules. In the neck network of the YOLOv10 network model, the Seop module is introduced in the P4 and P5 feature map processing links of multi-scale feature fusion. The detection head of the YOLOv10 network model is replaced with the Lag module. Among them, the Esvg module is a multi-scale cross-stage fusion module used to extract deep features of the image and capture multi-scale information in the input data. The Esvg module divides the data into two branches. One branch extracts multi-scale features of the image through n ESBlock modules, and the other branch retains the original features. The original features and the multi-scale features of the image are then spliced ​​together and finally output after convolution processing. The ESBlock module uses a multi-branch structure to process the input feature map. The Seop module is a dynamic perception feature information enhancement module that processes the data in two paths. One path is output after passing through the average pooling layer, convolution layer, and Softmax function; the other path is output after multi-branch refinement processing of the features through the grouped convolution module; finally, the two outputs are feature weighted and the weighted downsampled feature map is output; The Lag module performs channel adjustment through the group normalization convolution module, and then distributes the output data of the shared convolution group to the positioning convolution module and the classification convolution module through the shared convolution group; among them, the positioning convolution module is responsible for regression, and the classification convolution module is responsible for classification; The Esvg module extracts feature information by: changing the number of channels i of the input feature map to the specified number of output channels m by using a 1*1 convolution; then, splitting the input into two branches with a channel number of 0.5*m through a Split operation; wherein the first branch extracts multi-scale features of the image through n ESBlock modules, and the second branch retains the original features; then, the ESBlock module is spliced ​​with the second branch through n connections to obtain a feature map with a channel number of (n+2)*0.5m, and finally, the channel number is changed to m through a 1*1 convolution to obtain the output; the ESBlock module includes a first convolution layer and a second convolution layer; wherein, the first convolution layer is processed using a 3×3 standard convolution, and the second convolution layer uses a multi-scale convolution module; The data processing flow of the Seop module is as follows: the input feature x is processed in two paths, the attention branch uses 3×3 global average pooling to maintain the spatial size, and adjusts the number of channels through 1×1 convolution to make the feature suitable for subsequent calculations; then, the dimension reorganization operation is used to reorganize (bs, ch, 2h, 2w) into (bs, ch, h, w, 4), that is, the 2×2 neighborhood is mapped to the fifth dimension to form four candidate downsampling points; then the Softmax function is used for normalization so that the total weight of the four candidate points is 1. The sum is 1; The GroupConv branch divides the feature map into four groups. Each group is downsampled by a factor of 2 using a 3×3 convolution, halving the size of the feature map and expanding the number of channels to four times. Then, a dimension reshaping operation is performed to reshape the output into (bs, ch, h, w, 4) , where bs is the batch size, ch is the number of channels, h is the height of the feature map, and w is the width of the feature map. That is, each pixel corresponds to four candidate feature maps, which are multiplied by the weight matrix of the attention branch to achieve feature weighting. Finally, the sum function is used to sum along the fifth dimension to output the refined (bs, ch, h, w) dimensional features. The shared convolution group includes two 3×3 grouped normalized convolution modules. The Lag module uses feature maps of three different scales, P3, P4, and P5, in the neck network as input. It first performs channel adjustment through a 1×1 grouped normalized convolution module. After fusion, it enters a 3×3 grouped normalized convolution module to further extract spatial information and form a shared feature representation. Subsequently, the shared features are passed back to the 3×3 grouped normalized convolution module to strengthen the feature representation and are assigned to two different task branches: the localization convolution module and the classification convolution module. The data processing flow of the multi-scale convolution module is as follows: First, the input feature map is divided into three parts proportionally through the Split operation, of which 50% of the feature map is directly retained as the main branch, 25% of the feature map is processed by a 3×3 convolution layer, and the remaining 25% of the feature map is processed by a 5×5 convolution layer for feature extraction; then, the output feature maps of the three branches are spliced ​​in the channel dimension and feature fused through a 1×1 convolution; The grouped convolution module divides the feature map into 4 groups, and each group performs a 3×3 convolution operation independently, with the number of control channels ch=64 and stride=2; The positioning convolution module and the classification convolution module use 1×1 convolution, and the output of the positioning convolution module is scale-adaptively calibrated through the Scale function.

2. A PCB surface defect detection method based on improved YOLOv10, characterized in that: The method comprises the following steps: Step 1. Construct the ESL-YOLO detection model described in claim 1; Step 2. Train the ESL-YOLO detection model to obtain a trained ESL-YOLO detection model; Step 3. Input the PCB image data into the trained ESL-YOLO detection model for surface defect detection, thereby identifying the location and category of surface defects in the PCB image.

3. A PCB surface defect detection method based on improved YOLOv10 according to claim 2, characterized in that: The process of training the ESL-YOLO detection model is as follows: Step a. Obtain a public dataset of PCB surface defect images, perform image enhancement and preprocessing, and construct a PCB surface defect dataset; then divide the PCB surface defect dataset into a training set, a validation set, and a test set in a ratio of 8:1:1; Step b. Standardizing or normalizing the training set, validation set, and test set to form a PCB surface defect image dataset that meets the requirements; Step c. Configure the training parameters, setting the batch size to 32 and the number of iterations to 200. Train the ESL-YOLO detection model using the training set data. In each training iteration, calculate the loss and update the ESL-YOLO detection model weights. After each training cycle, use the validation set data to evaluate the performance of the ESL-YOLO detection model, monitoring the loss and accuracy on the validation set data. Also, enable Mosaic for data augmentation during ESL-YOLO detection model training and disable Mosaic for the last 10 training rounds. Simultaneously optimize multi-scale prediction paths during the training phase, and use bounding box regression and classification prediction for joint optimization; The decoding and post-processing stages use the Top-K mechanism to select high-confidence boxes and remove redundant boxes through non-maximum suppression to generate the final detection results.

Citation Information

Patent Citations

  • Coal mine shaft micro-crack detection method and device and storage medium

    CN119107444A

  • Unmanned aerial vehicle image small target detection method based on improved YOLOv8

    CN119495036A

  • Lightweight mobile phone screen defect detection method based on improved YOLOv10

    CN119904449A