A PCB solder joint defect visual detection method and system
Through the improved FD-YOLO model, the problems of high computational complexity and low detection accuracy in PCB solder joint defect detection are solved, and efficient and accurate solder joint defect detection is achieved, which adapts to multi-scale features and complex backgrounds and meets production needs.
Patent Information
- Application Number
- CN202511037373.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing PCB solder joint defect detection algorithms have problems such as high computational complexity, slow inference speed, high hardware resource requirements, or low detection accuracy, making it difficult to meet actual production needs.
An improved FD-YOLO model is adopted. The feature extraction capability is enhanced by introducing the TripletAT module at the end of the backbone network, the DWAT module is introduced in the neck network to dynamically adjust the local pooling window, and the upsampling module of the neck network is replaced by the Dysample module. Combined with the dynamic window prediction network and the improved loss function, the detection accuracy and efficiency are improved.
It achieves high-precision and high-efficiency PCB solder joint defect detection, reduces computing resource consumption, improves small target resolution, adapts to multi-scale solder joint defect characteristics, and improves the real-time and accuracy of detection.
Smart Images

Figure CN120543549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and in particular to a method and system for visually detecting PCB solder joint defects. Background Art
[0002] Printed circuit boards (PCBs) are essential components in the electronics industry, playing a vital role in sectors such as communications, computing, and electronics. As these industries continue to advance, PCBs are also evolving, increasing in complexity, miniaturization, and precision. However, defects such as solder leaks, voids, and shorts often occur during the manufacturing process. These defects can affect the stability of PCB assemblies and, in turn, the proper functioning of electronic devices. If defective boards with solder defects are not promptly and accurately identified and addressed, they can flow onto subsequent production lines, resulting in high scrap rates, significant financial losses, and damage to the manufacturer's reputation.
[0003] Currently, object detection algorithms can be roughly divided into two categories: one is a two-stage detection algorithm based on region proposal, which has the defects of high computational complexity, slow inference speed, and high hardware resource requirements; the other is a one-stage detection algorithm based on direct regression, which has limited feature fusion capabilities and relatively low detection accuracy.
[0004] In view of this, the present invention proposes a PCB solder joint defect visual detection method and system, which has the advantages of high real-time performance and good precision, and can meet actual production needs. Summary of the Invention
[0005] The purpose of the present invention is to provide a PCB welding defect detection method and system, taking into account both real-time and detection accuracy requirements to meet actual production needs.
[0006] To achieve the above object, the technical solution of the present invention is: a method for visually detecting PCB solder joint defects, specifically comprising the following steps:
[0007] S1, obtain PCB solder joint image and perform preprocessing;
[0008] S2. Build an improved FD-YOLO model based on the YOLO target detection model:
[0009] The TripletAT module is introduced at the end of the backbone network to enhance the correlation between spatial and channel features through rotation operations and cross-dimensional interactions on the input features of the TripletAT module.
[0010] The Upsample layer of the neck network is replaced with the Dysample module. The DWAT module is introduced into the output part of the neck network. The comprehensive complexity score is obtained based on the horizontal and vertical gradient amplitudes and channel variance of the DWAT module input features to dynamically evaluate the regional complexity. The comprehensive complexity score is mapped to the window size through the dynamic window prediction network. Local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features. The multi-scale features are weighted summed to obtain enhanced features.
[0011] S3. Input the preprocessed PCB solder joint image into the trained FD-YOLO model to obtain defect detection results.
[0012] Preferably, the pretreatment comprises the following steps:
[0013] S1.1 Image resizing: The original PCB solder joint image is resized to a fixed input size through bilinear interpolation to meet the network input requirements.
[0014] S1.2 Normalization and Standardization: The pixel values of the resized image are normalized and standardized to eliminate the influence of illumination differences.
[0015] S1.3 Noise suppression: Adaptive median filtering is performed on the weld spot area of the normalized and standardized image to suppress motion blur and high-frequency noise caused by conveyor belt vibration.
[0016] Preferably, the FD-YOLO model is as follows:
[0017] The backbone network includes a first convolutional layer, a second convolutional layer, a first C2F module, a third convolutional layer, a second C2F module, a first SCDown module, a third C2F module, a second SCDown module, a fourth C2F module, an SPPF module, and a TripletAT module connected in sequence;
[0018] The neck network includes 2 Dysample modules, 3 C2F modules, 1 convolutional layer, 1 SCDown module and 3 DWAT modules; the output features of the backbone network TripletAT module are input into the first Dysample module, the output features of the first Dysample module are spliced with the output features of the third C2F module of the backbone network and then input into the fifth C2F module, the output features of the fifth C2F module are input into the second Dysample module, the output features of the second Dysample module are spliced with the output features of the second C2F module of the backbone network and then input into the sixth C2F module, the output features of the sixth C2F module are respectively input into the fourth convolutional layer and the first DWAT module, the output features of the fourth convolutional layer are spliced with the output features of the fifth C2F module and then input into the seventh C2F module, the output features of the seventh C2F module are respectively input into the SCDown module and the second DWAT module, the output features of the SCDown module are spliced with the output features of the backbone network TripletAT module and then input into the third DWAT module;
[0019] The head part includes three groups of detection heads connected to the three DWAT modules of the neck network respectively.
[0020] Preferably, the TripletAT module includes a channel C and spatial H dimension interaction branch, a channel C and spatial W dimension interaction branch, and a spatial H dimension and W dimension interaction branch; wherein the H dimension and the W dimension refer to the height direction dimension and the width direction dimension of the feature map respectively;
[0021] The channel C and spatial H dimension interact with each other to input features to the TripletAT module. Rotate 90 degrees counterclockwise along the H axis to obtain the feature , for features Application Z -pool Operation and convolution operation generate weights , according to the characteristics and weights Calculate output features :
[0022]
[0023] ⊙
[0024] And output features Rotate 90 degrees clockwise along the H axis to obtain the features after rotation recovery ;in is the Sigmoid activation function, represents the channel C and spatial H dimension interactive branch convolution layer, ⊙ represents element-by-element multiplication, Represents Z-pool operate;
[0025] Channel C and spatial W dimension interactive branch, input feature to TripletAT module Rotate 90 degrees counterclockwise along the W axis to obtain the feature , for features Application Z -pool Operation and convolution operation generate weights , according to the characteristics and weights Calculate output features :
[0026]
[0027] ⊙
[0028] And output features Rotate 90 degrees clockwise along the W axis to obtain the features after rotation recovery ;in Represents the channel C and spatial W dimension interactive branch convolution layer;
[0029] The spatial H dimension and W dimension interact with each other to input features to the TripletAT module Application Z -pool Operation and convolution operation generate weights , according to the input features and weights Calculate output features :
[0030]
[0031] ⊙
[0032] in Represents the spatial H-dimensional and W-dimensional interactive branch convolutional layers;
[0033] The output of the three branches is averaged and fused to obtain the TripletAT module output. :
[0034] .
[0035] Preferably, the horizontal and vertical gradient amplitudes and channel variances based on the DWAT module input features are used to obtain a comprehensive complexity score, specifically as follows:
[0036] For DWAT module input characteristics , use Sobel operator to extract input features Horizontal gradient and vertical gradient , generating the gradient magnitude :
[0037]
[0038] Among them, H and W represent the input features respectively The height and width of the input feature map F; i is the pixel index of the input feature map F in the height direction, and 1≤i≤H; j is the pixel index of the input feature map F in the width direction, and 1≤j≤W; is the input feature At the pixel The horizontal gradient at is the input feature At the pixel The vertical gradient at
[0039] For DWAT module input characteristics , calculate the channel variance :
[0040]
[0041] in, Indicates the number of channels; Represents input features c-th channel characteristics; represents the spatial variance of the c-th channel feature map Fc;
[0042] Weighted fusion of gradient amplitude and channel variance to obtain a comprehensive complexity score :
[0043]
[0044] Where, and are the weights of the gradient amplitude and channel variance, respectively.
[0045] Preferably, the comprehensive complexity score is mapped to the window size through a dynamic window prediction network, and the dynamic window prediction network includes two layers Convolution and 1 fully connected layer to generate the window size based on the comprehensive complexity score of the input :
[0046]
[0047] in, represents the comprehensive complexity score, represents the global average pooling operation, and represents two convolutional layers, represents the weight matrix of the fully connected layer, Represents the bias vector of the fully connected layer.
[0048] Preferably, the local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features, and the weighted summation of the multi-scale features is performed to obtain enhanced features, as follows:
[0049] According to the window size Parallel execution of DWAT module input features 3×3, 5×5, and 7×7 local average pooling to obtain multi-scale features :
[0050]
[0051] in, 、 and Respectively represent the input features of the DWAT module The window is 、 、 Local average pooling operation;
[0052] Multi-scale features Perform splicing to obtain local pooling features ;
[0053] Input characteristics to the DWAT module Perform global average pooling to obtain global pooling features ;
[0054] Based on local pooling features and global pooling features The channel attention weight is generated by the Sigmiod function to analyze the multi-scale features. Perform weighted summation to obtain enhanced features :
[0055]
[0056] in is the Sigmoid activation function, and ⊙ represents element-wise multiplication.
[0057] Preferably, the training of the FD-YOLO model adopts the loss function :
[0058]
[0059] Among them, IOU represents the intersection-union ratio of the predicted box and the real box; Represents the center point of the prediction box The center point of the real frame The Euclidean distance of ' represents the diagonal length of the minimum bounding box covering the predicted box and the true box; and Represents the width and height difference between the predicted box and the real box respectively; Represent the predicted box width and the real box width respectively; Represents the predicted box height and the real box height respectively; and Represent the width and height of the minimum bounding box respectively.
[0060] The present invention proposes a PCB solder joint defect visual detection system, comprising a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs any step in the above-mentioned PCB solder joint defect visual detection method.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] The present invention optimizes the YOLO target detection model for PCB solder joint defect detection and proposes an FD-YOLO model. The TripletAT module is introduced at the end of the backbone network to enhance the feature extraction capability of the backbone network for PCB solder joint defect images. The new DWAT module is introduced in the neck network to dynamically adjust the local pooling window, suppress complex background interference, and adapt to multi-scale solder joint defect features. In addition, the original upsampling module in the neck network is replaced with the Dysample module to enhance the features after upsampling, improve the feature fusion effect of the neck part, and solve the problem of important feature information loss in the image during the sampling process. This not only reduces the consumption of computing resources, but also improves the resolution of small targets without adding additional burden. Finally, by replacing the original loss function, the positioning accuracy of solder joint defects and the convergence speed of the model are improved. By performing defect detection on PCB solder joints using the above-mentioned improved FD-YOLO model, high-precision and high-efficiency detection of PCB solder joint defects can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is the framework diagram of the FD-YOLO model of the present invention;
[0064] Figure 2 This is a framework diagram of the TripletAT module of the present invention;
[0065] Figure 3 This is the framework diagram of the DWAT module of the present invention. DETAILED DESCRIPTION
[0066] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0067] refer to Figure 1-3 The present invention proposes a method for visually detecting PCB solder joint defects, which specifically includes the following steps:
[0068] S1. Obtain the original PCB solder joint image captured by the line scan camera and perform preprocessing;
[0069] S2. Build an improved FD-YOLO model based on the YOLO target detection model:
[0070] The TripletAT module is introduced at the end of the backbone network to enhance the correlation between spatial and channel features through rotation operations and cross-dimensional interactions on the input features of the TripletAT module.
[0071] The Upsample layer of the neck network is replaced with the Dysample module. The DWAT module is introduced into the output part of the neck network. The comprehensive complexity score is obtained based on the horizontal and vertical gradient amplitudes and channel variance of the DWAT module input features to dynamically evaluate the regional complexity. The comprehensive complexity score is mapped to the window size through the dynamic window prediction network. Local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features. The multi-scale features are weighted summed to obtain enhanced features.
[0072] S3. Input the preprocessed PCB solder joint image into the trained FD-YOLO model to obtain defect detection results, including marking defective solder joints and determining the type of solder joint defects.
[0073] Preferably, the pretreatment comprises the following steps:
[0074] S1.1 Image resizing: Resize the original PCB solder joint image to a fixed input size (e.g., 640 × 640) through bilinear interpolation to accommodate network input requirements. Use a scaling factor to resize the image to ensure the image aspect ratio remains unchanged, and pad the image edges with zeros.
[0075] S1.2 Normalization and Standardization: The pixel values of the resized image are normalized and standardized to eliminate the influence of illumination differences.
[0076] S1.3 Noise suppression: Adaptive median filtering is performed on the solder joint area of the normalized and standardized image, and the window size is dynamically adjusted (3×3 or 5×5) to suppress motion blur and high-frequency noise caused by conveyor belt vibration.
[0077] Preferably, the FD-YOLO model is as follows:
[0078] The backbone network includes a first convolutional layer, a second convolutional layer, a first C2F module, a third convolutional layer, a second C2F module, a first SCDown module, a third C2F module, a second SCDown module, a fourth C2F module, an SPPF module, and a TripletAT module connected in sequence;
[0079] The neck network includes 2 Dysample modules, 3 C2F modules, 1 convolutional layer, 1 SCDown module and 3 DWAT modules; the output features of the backbone network TripletAT module are input into the first Dysample module, the output features of the first Dysample module are spliced with the output features of the third C2F module of the backbone network and then input into the fifth C2F module, the output features of the fifth C2F module are input into the second Dysample module, the output features of the second Dysample module are spliced with the output features of the second C2F module of the backbone network and then input into the sixth C2F module, the output features of the sixth C2F module are respectively input into the fourth convolutional layer and the first DWAT module, the output features of the fourth convolutional layer are spliced with the output features of the fifth C2F module and then input into the seventh C2F module, the output features of the seventh C2F module are respectively input into the SCDown module and the second DWAT module, the output features of the SCDown module are spliced with the output features of the backbone network TripletAT module and then input into the third DWAT module;
[0080] The head section consists of three detection heads, P3, P4, and P5, connected to the three DWAT modules in the neck network. The heads utilize an anchor-free design and output defect locations and categories (such as cold solder joints and leaky solder joints). Post-processing includes filtering overlapping frames using non-maximum suppression (NMS), retaining predictions with a confidence level greater than 0.5, and annotating the defect location and type.
[0081] Preferably, the TripletAT module includes a channel C and spatial H dimension interaction branch, a channel C and spatial W dimension interaction branch, and a spatial H dimension and W dimension interaction branch; wherein the H dimension and the W dimension refer to the height direction dimension and the width direction dimension of the feature map respectively;
[0082] The channel C and spatial H dimension interact with each other to input features to the TripletAT module. Rotate 90 degrees counterclockwise along the H axis to obtain the feature , for features Application Z -pool Operation and convolution operation generate weights , according to the characteristics and weights Calculate output features :
[0083]
[0084] ⊙
[0085] And output features Rotate 90 degrees clockwise along the H axis to obtain the features after rotation recovery ;in is the Sigmoid activation function, represents the channel C and spatial H dimension interactive branch convolution layer, ⊙ represents element-by-element multiplication, Represents Z -pool operate;
[0086] The Z -pool The specific operation is to perform maximum pooling and average pooling operations along the channel dimension (dimension 0), and then concatenate the pooling results:
[0087]
[0088] in, Indicates Z -pool Characteristics of the operation, Represents the feature Perform maximum pooling along the channel dimension, Represents the feature Perform average pooling along the channel dimension, Indicates concatenation of the two pooling results.
[0089] Channel C and spatial W dimension interactive branch, input feature to TripletAT module Rotate 90 degrees counterclockwise along the W axis to obtain the feature , for features Application Z -pool Operation and convolution operation generate weights , according to the characteristics and weights Calculate output features :
[0090]
[0091] ⊙
[0092] And output features Rotate 90 degrees clockwise along the W axis to obtain the features after rotation recovery ;in Represents the channel C and spatial W dimension interactive branch convolution layer;
[0093] The spatial H dimension and W dimension interact with each other to input features to the TripletAT module Apply Z-pool operation and convolution operation to generate weights , according to the input features and weights Calculate output features :
[0094]
[0095] ⊙
[0096] in Represents the spatial H-dimensional and W-dimensional interactive branch convolutional layers;
[0097] The output of the three branches is averaged and fused to obtain the TripletAT module output. :
[0098] .
[0099] The TripletAT module in this paper is located at the end of the Backbone in the FD-YOLO network and is key to network feature extraction, effectively improving the ability to extract features from small objects. Through rotation operations and residual correction, triple attention establishes the relationship between dimensions, improving the quality of spatial and channel information.
[0100] Preferably, the horizontal and vertical gradient amplitudes and channel variances based on the DWAT module input features are used to obtain a comprehensive complexity score, specifically as follows:
[0101] For DWAT module input characteristics , use Sobel operator to extract input features Horizontal gradient and vertical gradient , generating the gradient magnitude :
[0102]
[0103] Among them, H and W represent the input features respectively The height and width of the input feature map F; i is the pixel index of the input feature map F in the height direction, and 1≤i≤H; j is the pixel index of the input feature map F in the width direction, and 1≤j≤W; is the input feature At the pixel The horizontal gradient at is the input feature At the pixel The vertical gradient at
[0104] For DWAT module input characteristics , calculate the channel variance :
[0105]
[0106] in, Indicates the number of channels; Represents input features c-th channel characteristics; represents the spatial variance of the c-th channel feature map Fc, that is, the average of the squares of the differences between all pixel values and the mean;
[0107] Weighted fusion of gradient amplitude and channel variance to obtain a comprehensive complexity score :
[0108]
[0109] Where, and are the weights of the gradient amplitude and channel variance, respectively.
[0110] Preferably, the comprehensive complexity score is mapped to the window size through a dynamic window prediction network, and the dynamic window prediction network includes two layers Convolution and 1 fully connected layer to generate the window size based on the comprehensive complexity score of the input :
[0111]
[0112] in, represents the comprehensive complexity score, represents the global average pooling operation, and represents two convolutional layers, represents the weight matrix of the fully connected layer, represents the bias vector of the fully connected layer, .
[0113] Preferably, the local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features, and the weighted summation of the multi-scale features is performed to obtain enhanced features, as follows:
[0114] According to the window size Parallel execution of DWAT module input features 3×3, 5×5, and 7×7 local average pooling to obtain multi-scale features :
[0115]
[0116] in, 、 and Respectively represent the input features of the DWAT module The window is 、 、 Local average pooling operation;
[0117] Multi-scale features Perform splicing to obtain local pooling features ;
[0118] Input characteristics to the DWAT module Perform global average pooling to obtain global pooling features ;
[0119] Based on local pooling features and global pooling features The channel attention weight is generated by the Sigmiod function to analyze the multi-scale features. Perform weighted summation to obtain enhanced features :
[0120]
[0121] in is the Sigmoid activation function, and ⊙ represents element-wise multiplication.
[0122] In industrial quality inspection scenarios, PCB solder joint defect detection faces challenges such as the multi-scale distribution of small objects and complex background interference. Traditional channel attention mechanisms (such as SE and ECA) compress channel information through global average pooling (GAP) but ignore the local dependencies of spatial features. Spatial attention mechanisms (such as CBAM), while capable of capturing spatial relationships, suffer from computational redundancy due to their complex dual-branch design. MLCA (Mixed Local Channel Attention) strikes a balance between lightweightness (only 2.99M parameters) and multi-scale information fusion by combining local average pooling (LAP) and global average pooling (GAP). However, its fixed window design (e.g., 5×5) limits its adaptability to multiple scenarios: small objects (such as solder joint defects) require finer-grained local features, while large background areas require larger windows to suppress noise. To address this issue, this paper proposes a dynamic windowed hybrid local channel attention mechanism (DWAT). This mechanism achieves adaptive window adjustment through the following improvements: It dynamically assesses region complexity based on the gradient magnitude (horizontal and vertical gradients extracted using the Sobel operator) and channel variance (a measure of inter-channel distribution differences) of the input feature map. High-gradient regions (edges and textures) are prioritized over 3×3 windows, low-variance regions (background smoothing) are switched to 7×7 windows, and medium-complexity regions are defaulted to 5×5 windows. A lightweight dynamic window prediction network (two layers of 1×1 convolutions plus a fully connected layer) maps complexity scores to window parameters, increasing the total number of parameters by only 0.1M (3.3% of the original module), ensuring hardware friendliness. 3×3, 5×5, and 7×7 local average pooling are performed on the input feature map in parallel to generate multi-scale features. Channel attention weights (SE-like structure) are used to dynamically allocate fusion weights to each branch, and the optimal features are output after weighted summation. This design significantly improves adaptability to multi-scale objects while retaining the original MLCA two-branch processing pipeline.
[0123] Preferably, the training of the FD-YOLO model adopts the loss function :
[0124]
[0125] Among them, IOU represents the intersection-union ratio of the predicted box and the real box; Represents the center point of the prediction box The center point of the real frame The Euclidean distance of ' represents the diagonal length of the minimum bounding box covering the predicted box and the true box; and Represents the width and height difference between the predicted box and the real box respectively; Represent the predicted box width and the real box width respectively; Represents the predicted box height and the real box height respectively; and Represent the width and height of the minimum bounding box respectively.
[0126] EIoU improves the positioning accuracy of solder joint defects by directly optimizing the center point distance, width, and height difference between the detection frame and the true frame:
[0127] (1) Faster model convergence: EIOU eliminates the ambiguity of the aspect ratio penalty term in CIOU by directly optimizing the width-to-height difference (rather than the aspect ratio), making the gradient calculation clearer and accelerating training convergence;
[0128] (2) Improve the detection accuracy of small targets: PCB solder joints are usually small targets. The width-height separation loss term of EIOU can more accurately match the actual size of small targets, reducing missed detection or false detection;
[0129] (3) Enhanced positioning robustness: Under complex backgrounds or interference from multiple solder joints, the explicit loss term of EIOU (center point distance + width and height difference) helps the model distinguish key features and improve positioning stability;
[0130] (4) Adapting to irregular shape defects: Solder defects may have irregular shapes (such as cold solder joints and cracks). The independent width and height optimization of EIOU can better fit non-rectangular targets and improve the matching degree of defect shapes.
[0131] (5) Reduce hyperparameter dependence: EIOU does not require additional balancing of aspect ratio weights, which reduces the difficulty of parameter adjustment and is more suitable for rapid deployment in industrial scenarios.
[0132] The present invention proposes a PCB solder joint defect visual detection system, comprising a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs any step in the above-mentioned PCB solder joint defect visual detection method.
[0133] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.
Claims
1. A visual inspection method for PCB solder joint defects, characterized in that: The specific steps include: S1, obtain PCB solder joint image and perform preprocessing; S2. Build an improved FD-YOLO model based on the YOLO target detection model: The TripletAT module is introduced at the end of the backbone network to enhance the correlation between spatial and channel features through rotation operations and cross-dimensional interactions on the input features of the TripletAT module. The Upsample layer of the neck network is replaced with the Dysample module. The DWAT module is introduced into the output part of the neck network. The comprehensive complexity score is obtained based on the horizontal and vertical gradient amplitudes and channel variance of the DWAT module input features. The comprehensive complexity score is mapped to the window size through the dynamic window prediction network. Local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features. The multi-scale features are weighted summed to obtain enhanced features. S3. Input the pre-processed PCB solder joint image into the trained FD-YOLO model to obtain defect detection results; The TripletAT module includes a channel C and spatial H dimension interaction branch, a channel C and spatial W dimension interaction branch, and a spatial H dimension and W dimension interaction branch; wherein the H dimension and the W dimension refer to the height direction dimension and the width direction dimension of the feature map respectively; The channel C and the spatial H dimension interact with each other, and the input feature x of the TripletAT module is rotated 90 degrees counterclockwise along the H axis to obtain feature x1, and Z is applied to feature x1. -pool The operation and convolution operation generate weight ω1, and the output feature y1 is calculated based on the feature x1 and weight ω1: y1=x1⊙ω1 And rotate the output feature y1 90 degrees clockwise along the H axis to obtain the rotated feature Where σ is the Sigmoid activation function, represents the channel C and spatial H dimension interactive branch convolution layer, ⊙ represents element-by-element multiplication, Z -pool (·) indicates Z -pool operate; Channel C and spatial W dimension interact with each other, and the input feature x of TripletAT module is rotated 90 degrees counterclockwise along the W axis to obtain feature x2, and Z is applied to feature x2. -pool The operation and convolution operation generate weight ω2, and the output feature y2 is calculated based on the feature x2 and weight ω2: y2=x2⊙ω2 And rotate the output feature y2 90 degrees clockwise along the W axis to obtain the rotated feature in Represents the channel C and spatial W dimension interactive branch convolution layer; The spatial H and W dimensions interact and branch, applying Z to the input feature x of the TripletAT module -pool The operation and convolution operation generate weight ω3, and the output feature y3 is calculated based on the input feature x and weight ω3: y3=x⊙ω3 in Represents the spatial H-dimensional and W-dimensional interactive branch convolutional layers; The outputs of the three branches are averaged and fused to obtain the TripletAT module output y:
2. A PCB solder joint defect visual detection method according to claim 1, characterized in that: The pretreatment comprises the following steps: S1.1 Image resizing: The original PCB solder joint image is resized to a fixed input size through bilinear interpolation to meet the network input requirements. S1.2 Normalization and Standardization: The pixel values of the resized image are normalized and standardized to eliminate the influence of illumination differences. S1.3 Noise suppression: Adaptive median filtering is performed on the weld spot area of the normalized and standardized image to suppress motion blur and high-frequency noise caused by conveyor belt vibration.
3. A PCB solder joint defect visual detection method according to claim 1, characterized in that: The FD-YOLO model is as follows: The backbone network includes a first convolutional layer, a second convolutional layer, a first C2F module, a third convolutional layer, a second C2F module, a first SCDown module, a third C2F module, a second SCDown module, a fourth C2F module, an SPPF module, and a TripletAT module, which are connected in sequence. The neck network includes 2 Dysample modules, 3 C2F modules, 1 convolutional layer, 1 SCDown module and 3 DWAT modules; the output features of the backbone network TripletAT module are input into the first Dysample module, the output features of the first Dysample module are spliced with the output features of the third C2F module of the backbone network and then input into the fifth C2F module, the output features of the fifth C2F module are input into the second Dysample module, the output features of the second Dysample module are spliced with the output features of the second C2F module of the backbone network and then input into the sixth C2F module, the output features of the sixth C2F module are respectively input into the fourth convolutional layer and the first DWAT module, the output features of the fourth convolutional layer are spliced with the output features of the fifth C2F module and then input into the seventh C2F module, the output features of the seventh C2F module are respectively input into the SCDown module and the second DWAT module, the output features of the SCDown module are spliced with the output features of the backbone network TripletAT module and then input into the third DWAT module; The head part includes three groups of detection heads connected to the three DWAT modules of the neck network respectively.
4. The method for visually detecting PCB solder joint defects according to claim 1, wherein: The comprehensive complexity score is obtained based on the horizontal and vertical gradient amplitudes and channel variance of the DWAT module input features, as follows: For the DWAT module input feature F, the Sobel operator is used to extract the horizontal gradient of the input feature F and vertical gradient Generate gradient magnitude Gradient: Where H and W represent the height and width of the input feature F respectively; i is the pixel index of the input feature map F in the height direction, and 1≤i≤H; j is the pixel index of the input feature map F in the width direction, and 1≤j≤W; is the horizontal gradient of the input feature F at pixel (i, j), is the vertical gradient of the input feature F at pixel (i, j); For the DWAT module input feature F, calculate the channel variance: Where C represents the number of channels; F c Represents the c-th channel feature of the input feature F; Var(F c ) represents the spatial variance of the c-th channel feature map Fc, the average value of the square of the difference between all pixel values and the mean; Weighted fusion gradient amplitude and channel variance are used to obtain the comprehensive complexity score Complexity: Complexity=α·Gradient+β·Variance Where α and β are the weights of the gradient amplitude and channel variance, respectively.
5. The method for visually detecting PCB solder joint defects according to claim 1, wherein: The comprehensive complexity score is mapped to the window size through the dynamic window prediction network. The dynamic window prediction network includes two layers of 1×1 convolution and one layer of fully connected layer to generate the window size k according to the input comprehensive complexity score: k=argmax(Softmax(W'·GAP(Conv2(Conv1(Complexity)))+b'))×2+3where Complexity represents the comprehensive complexity score, GAP represents the global average pooling operation, Conv1 and Conv2 represent two convolutional layers, W' represents the weight matrix of the fully connected layer, and b' represents the bias vector of the fully connected layer.
6. A visual inspection method for PCB solder joint defects according to claim 1, characterized in that: The local average pooling of the DWAT module input features is performed according to the window size to generate multi-scale features, and the weighted sum of the multi-scale features is performed to obtain enhanced features, as follows: According to the window size k, the 3×3, 5×5, and 7×7 local average pooling of the DWAT module input feature F is performed in parallel to obtain the multi-scale feature F pooled : Among them, AvgPool 3×3 (F), AvgPool 5×5 (F) and AvgPool 7×7 (F) represents the local average pooling operation of the DWAT module input feature F with windows of 3×3, 5×5, and 7×7 respectively; For multi-scale features F pooled Perform splicing to obtain local pooling features F local ; Perform global average pooling on the DWAT module input feature F to obtain the global pooling feature F global ; According to the local pooling feature F local and global pooling feature F global The channel attention weight is generated by the Sigmiod function to analyze the multi-scale features F pooled Perform weighted summation to obtain enhanced features F out : F out =σ(F global +F local )⊙F pooled Where σ is the Sigmoid activation function and ⊙ represents element-wise multiplication.
7. The method for visually detecting PCB solder joint defects according to claim 1, wherein: The FD-YOLO model is trained using the loss function L EIOU : Among them, IOU represents the intersection-union ratio between the predicted box and the real box; ρ(b,b gt ) represents the center point b of the predicted box and the center point b of the real box gt The Euclidean distance of the predicted box and the real box; C' represents the diagonal length of the minimum bounding box covering the predicted box and the real box; ρ(w,w gt ) and ρ(h,h gt ) represent the width and height differences between the predicted box and the real box respectively; w, w gt Represent the predicted box width and the real box width respectively; h, h gt Represents the predicted box height and the real box height respectively; C' w and C' h Represent the width and height of the minimum bounding box respectively.
8. A method for visually detecting PCB solder joint defects, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, the method specifically performs the steps of the method for visually detecting PCB solder joint defects as described in any one of claims 1 to 7.
Citation Information
Patent Citations
PCB plug-in welding spot defect detection method and system and storage medium thereof
CN113724245A
PCB defect detection and identification method based on YOLO-SEE
CN117372339A