Circuit board appearance defect intelligent detection equipment

Through multi-scale feature extraction network, prior knowledge-guided attention calculation and cross-scale feature complementary two-way propagation system, the problems of unbalanced detection accuracy, high computing resources and poor adaptability in PCB solder joint defect detection are solved, and efficient and adaptive solder joint defect detection is achieved.

CN120235874AInactive Publication Date: 2025-07-01SHENZHEN ZHONGYUAN CIRCUIT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510718343.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems in PCB solder joint defect detection, such as unbalanced detection accuracy, fixed feature fusion strategy, high computing resource requirements and lack of effective use of prior knowledge of solder joint structures, resulting in low detection efficiency and poor adaptability.

Method used

A multi-scale feature extraction network, prior knowledge-guided attention calculation, teacher-student network architecture and a two-way communication system with complementary cross-scale features, combined with lightweight optimization, is used to build an intelligent detection device for circuit board appearance defects.

Benefits of technology

The detection accuracy of defects of different sizes, especially the detection rate of small defects, realize adaptive feature fusion, reduce the calculation amount, and improve the detection speed and adaptability on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235874A_ABST
    Figure CN120235874A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and circuit board manufacturing defect detection, and discloses circuit board appearance defect intelligent detection equipment, which comprises an image acquisition and feature extraction module for extracting multi-level features of a circuit board welding spot image; the attention calculation module is used for performing knowledge-guided attention calculation on the multi-scale feature map and focusing on a potential defect area; the knowledge distillation module is used for constructing a teacher-student network architecture and migrating high-precision teacher network knowledge to a lightweight student network; the feature complementation module is used for constructing a cross-scale feature complementation two-way propagation system in the student network; the lightweight optimization module is used for carrying out further lightweight optimization on the student network; key technical problems in PCB welding spot defect detection are solved through a multi-scale feature extraction network, attention calculation guided by priori knowledge, a teacher-student network architecture and a cross-scale feature complementary two-way propagation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and defect detection in printed circuit board manufacturing. More specifically, it relates to an intelligent detection device for printed circuit board appearance defects. Background Art

[0002] With the development of electronic products towards miniaturization, high density, and high reliability, the automatic detection of solder joint defects on printed circuit boards has become particularly important. Especially in consumer electronic products such as smartphones and wearable devices, as well as in fields with high reliability requirements such as automotive electronics and medical devices, the quality of PCB solder joints directly affects product performance and safety.

[0003] Currently, the detection of printed circuit board solder joint defects mainly relies on manual visual inspection, machine vision, and deep learning algorithms. Although manual visual inspection has a relatively high accuracy, it is inefficient and easily affected by human factors; traditional machine vision methods rely on manually designed feature extractors and rules, making it difficult to handle complex and variable solder joint defects; although the deep learning methods that have emerged in recent years have improved in detection accuracy, there are still obvious deficiencies.

[0004] The existing deep learning-based PCB solder joint defect detection technologies have the following main problems: First, the detection accuracy for defects of different sizes is uneven, especially the detection accuracy for small defects (such as tiny virtual soldering and fine cracks) is relatively low; second, the feature fusion strategy is fixed and lacks the ability to adaptively adjust for different defect types; third, the high-precision detection model has a large computational amount and is difficult to run in real time on resource-constrained edge devices on the production line; fourth, the effective utilization of prior knowledge of the solder joint structure is lacking, resulting in a decrease in detection accuracy and poor adaptability on new batches of products.

[0005] These problems seriously restrict the wide application of deep learning methods in the field of PCB solder joint defect detection. There is an urgent need for a high-precision, high-efficiency, and highly adaptable intelligent detection system that can solve the above problems. Summary of the Invention

[0006] The present invention provides an intelligent detection device for printed circuit board appearance defects, which solves the technical problems in the related technologies of uneven detection accuracy for defects of different sizes, fixed feature fusion strategy, high computational resource requirements, and lack of effective utilization of prior knowledge of the solder joint structure.

[0007] The present invention provides an intelligent detection device for printed circuit board appearance defects, including: An image acquisition and feature extraction module, used to construct a multi-scale feature extraction network and a solder joint prior knowledge base, and extract multi-level features of the printed circuit board solder joint image; An attention calculation module, based on the multi-level features output by the multi-scale feature extraction network and the solder joint prior knowledge base, performs knowledge-guided attention calculation on the multi-scale feature map, focusing on potential defect areas; A knowledge distillation module, based on the multi-scale feature extraction network and the attention calculation module, constructs a teacher-student network architecture, and transfers the knowledge of the high-precision teacher network to the lightweight student network; A feature complementary module, used to construct a bidirectional propagation system for cross-scale feature complementarity in the student network, enhancing the information exchange between features of different scales; A lightweight optimization module, used to further optimize the lightweight of the student network so that it can operate efficiently on resource-constrained edge devices.

[0008] Furthermore, the solder joint prior knowledge base contains the geometric shapes, size ratios, surface textures of different types of solder joints, and visual feature information under normal welding states, forming prior knowledge vectors: ; where represents the prior knowledge vector of the th type of solder joint, represents the solder joint type index, represents the dimension of the prior knowledge vector, represents the real number field; The prior knowledge base is expressed as: ; where represents the complete set of solder joint prior knowledge bases, , , respectively represent the prior knowledge vectors of the 1st, 2nd, th type of solder joint, is the total number of solder joint types.

[0009] Furthermore, the attention calculation module includes: Solder joint type recognition and prior knowledge embedding; Attention map calculation.

[0010] Furthermore, the calculation formula for the prior knowledge embedding is: ; where represents the prior knowledge embedding, represents the prior knowledge vector corresponding to the solder joint type, represents the learnable embedding matrix.

[0011] Further, the construction of the teacher-student network architecture includes constructing a multi-level attention distillation module. For the attention maps of each layer of the teacher network and the student network, calculate the multi-level attention distillation loss. The calculation formula is: ; Where represents the multi-level attention distillation loss, and respectively represent the attention maps of the th layer of the teacher network and the student network, represents the separation symbol between two distributions; represents the KL divergence, which is used to measure the difference between two distributions; is the weight coefficient of the th layer, which is used to balance the contribution of the distillation loss of different layers; represents the accumulation starting from the first layer of the network, represents the total number of network layers for attention distillation, represents the accumulation over all layers.

[0012] Further, the steps for constructing the cross-scale feature complementary two-way propagation system include: Construct a feature two-way propagation module for constructing two-way feature propagation from high-level to low-level and from low-level to high-level; Cross-scale feature fusion for fusing feature information at different scales, retaining both high-level semantic information and not losing low-level detail information.

[0013] Further, the cross-scale feature fusion calculation formula is: ; Where represents the fused feature map, is the weight coefficient of the feature of the th layer, represents the index of the feature layer, represents the th layer of the feature map; is the weight coefficient of the feature interaction between the th layer and the th layer, and are the indices of the feature layers and satisfy and ; represents the Hadamard product (element-wise multiplication) operation for implementing feature interaction; represents the total number of feature layers participating in feature fusion; and respectively represent the th layer and the The feature map of the layer Represents the summation symbol.

[0014] Furthermore, the lightweight optimization module includes: Progressive channel pruning based on channel importance measurement, and the channel importance calculation formula is: ; Where Represents the importance score of the th channel, Represents the absolute value symbol, Represents the summation symbol, Represents the weight parameter related to the th channel, And Represents the position index in the weight matrix, Represents the channel index; Performs hardware-aware optimization according to the characteristics of the target hardware platform.

[0015] Furthermore, the hardware-aware optimization according to the characteristics of the target hardware platform includes: Selects an appropriate quantization precision according to the characteristics of the hardware arithmetic unit; Optimizes the parameter storage method according to the hardware memory architecture; Adjusts the network layer structure according to the hardware parallel computing ability.

[0016] A computer-readable storage medium for storing computer-readable instructions that, when read by a computer, can run an intelligent detection device for PCB appearance defects as described above.

[0017] The beneficial effects of the present invention are as follows: By innovative designs such as a multi-scale feature extraction network, prior knowledge-guided attention calculation, a teacher-student network architecture, and a two-way propagation system for cross-scale feature complementarity, the key technical problems in PCB solder joint defect detection are solved; The detection accuracy of defects of different sizes is improved, especially the detection rate of small defects is increased, effectively solving the problem of uneven detection accuracy of defects of different sizes in the prior art; An adaptive feature fusion strategy for different defect types is realized. Through the two-way propagation system for cross-scale feature complementarity, the average detection rate of different types of solder joint defects is improved, effectively solving the problem of fixed feature fusion strategies in the prior art; Through the teacher-student network architecture and model lightweight optimization, the number of model parameters and the model size are reduced, and at the same time, the detection speed on edge devices is improved, effectively solving the problem of large computational complexity of high-precision detection models; Based on the prior knowledge-guided attention calculation mechanism, the system reduces the decline in the detection rate on new batches of products, improves adaptability, and effectively solves the problem of poor adaptability caused by the lack of effective utilization of prior knowledge of solder joint structures in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a module diagram of an intelligent detection device for appearance defects of a circuit board in the present invention; Figure 2 is a flowchart for constructing a multi-scale feature extraction network and a solder joint prior knowledge base in the present invention; Figure 3 is a flowchart of a prior knowledge-guided attention calculation module in the present invention; Figure 4 is a flowchart of a teacher-student network architecture and knowledge distillation in the present invention; Figure 5 is a flowchart of a cross-scale feature complementary bidirectional propagation system in the present invention; Figure 6 is a flowchart of model lightweight optimization in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and the functions and arrangements of the elements discussed can be changed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0020] In at least one embodiment of the present invention, an intelligent detection device for appearance defects of a circuit board is disclosed, as Figures 1 to 6 shown, including: An image acquisition and feature extraction module, configured to construct a multi-scale feature extraction network and a solder joint prior knowledge base, and extract multi-level features of the circuit board solder joint image; Step 1.1, constructing a multi-scale feature extraction network; Construct a multi-layer convolutional neural network structure, and extract multi-scale features of the circuit board solder joint image through convolutional kernels of different scales and a feature pyramid structure. This network includes feature extraction layers, and each layer outputs a feature map: ; wherein represents the feature map of the th layer, ; represents the The number of channels of the layer, and respectively represent the height and width of the feature map of the layer, indicating the real number field.

[0021] The multi-scale feature set is represented as: ; where , , respectively represent the feature maps of the first layer, the second layer, and the layer, indicating the total number of feature extraction layers; represents the lowest-level feature, with a higher spatial resolution but less semantic information; represents the highest-level feature, with a lower spatial resolution but rich semantic information.

[0022] The specific structure of the multi-scale feature extraction network includes two parts: the backbone network and the feature pyramid network; The backbone network adopts an improved ResNet structure, including 5 residual blocks, each residual block contains multiple convolutional layers, a total of 16 convolutional layers. The input of the backbone network is a solder joint RGB image with a size of , and it outputs 5 feature maps with different scales, which are: ; ; ; ; ; where , , , , respectively represent the first, second, third, fourth, and fifth feature maps. The five feature maps are from shallow to deep in sequence, with increasing channel numbers and decreasing spatial resolutions; represents the height of the input image, represents the width of the input image; retains more spatial detail information and is suitable for detecting small defects; contains richer semantic information and is suitable for identifying complex defect patterns.

[0023] This multi-scale feature extraction structure can simultaneously capture local details and global semantic information in the solder joint image, providing a comprehensive feature representation for subsequent defect detection.

[0024] The Feature Pyramid Network fuses features of different scales through a top-down path and lateral connections: The top-down path doubles the size of the high-level feature map to the same as the low-level feature map through 2x upsampling; The lateral connections then adjust the number of channels of the low-level feature map to be the same as that of the high-level feature Figure One through convolution, and then add the two.

[0025] Finally, the Feature Pyramid Network outputs 5 feature maps of the same scale but different levels to , which fuse semantic information and spatial details of different levels.

[0026] In the solder joint defect detection scenario, this multi-scale feature extraction network can handle solder joint defects of different sizes simultaneously. For example, for tiny solder joint voids, the low-level feature maps and can provide sufficient spatial detail information; for larger bridging or short-circuit defects, the high-level feature maps and can provide better semantic recognition ability. Through multi-scale feature extraction, this network effectively solves the problem that traditional single-scale features are difficult to handle different-size defect detections simultaneously.

[0027] Step 1.2, establish a prior knowledge base for solder joint structures; Collect standard geometric feature information of different types of solder joints (including ball grid array solder joints, plug-in solder joints, surface mount solder joints, etc.) to establish a prior knowledge base for solder joint structures; First, for each type of solder joint, extract information such as its geometric shape, size ratio, surface texture, visual features under normal soldering conditions, etc. to form a prior knowledge vector: ; where represents the prior knowledge vector of the k-th type of solder joint, represents the solder joint type index, represents the real number field, represents the dimension of the prior knowledge vector, usually set to a value between 128 and 512 to fully express the geometric and visual features of the solder joint; The prior knowledge base can be expressed as: ; where represents the complete set of prior knowledge bases for solder joints, , , respectively represent the prior knowledge vectors of the 1st, 2nd, th type of solder joint, is the total number of solder joint types; in practical applications, the solder joint types are usually 10 to 30 different types.

[0028] Through this module, a multi-scale feature extraction network and a solder joint prior knowledge base are constructed; the set of multi-level feature maps output by the multi-scale feature extraction network and the prior knowledge base are respectively used as the key inputs for the subsequent modules, providing a basis for subsequent attention calculation, feature fusion, and defect detection; among them, the multi-scale feature maps can capture the visual features of defects of different sizes, while the prior knowledge base provides standard reference information on the solder joint structure. The combination of the two can effectively improve the defect recognition ability of the detection system.

[0029] The attention calculation module, based on the multi-level features output by the multi-scale feature extraction network and the solder joint prior knowledge base, performs knowledge-guided attention calculation on the multi-scale feature maps, focusing on potential defect areas; Based on the solder joint structure prior knowledge base constructed in the image acquisition and feature extraction module and the set of multi-scale feature maps perform knowledge-guided attention calculation on the multi-scale feature maps, focusing on potential defect areas.

[0030] Step 2.1, Solder joint type recognition and prior knowledge embedding; Perform type recognition on the input solder joint image, and obtain the corresponding prior knowledge vector from the prior knowledge base ; Through a learnable embedding matrix: ; where represents the learnable embedding matrix, represents the number of channels of the feature map, represents the dimension of the prior knowledge vector, represents the real number field; Convert the prior knowledge vector into a prior knowledge embedding compatible with the feature map: ; where represents the prior knowledge embedding, represents the prior knowledge vector corresponding to the solder joint type.

[0031] Step 2.2, Attention map calculation; For each layer of feature map , calculate the attention map combined with prior knowledge, and the calculation formula is: ; where Indicates the feature map of the layer at the position of the attention map, indicating the spatial position in the feature map; and respectively represent the weight matrices of the first and second convolutional layers, and , represents the real number field, represents the number of channels of the input feature map, is the intermediate feature dimension (usually set to half to reduce the computational amount); represents the rectified linear unit activation function for introducing non - linear transformation; represents the sigmoid activation function to ensure that the attention weights are between 0 and 1; indicates the layer feature map at the position of the feature vector.

[0032] Through this calculation, the model can automatically focus on the key areas where defects may exist in the solder joint image based on prior knowledge.

[0033] The output of this module is a series of attention maps combined with prior knowledge , corresponding to each layer of the feature map . These attention maps guide the model to focus on the key areas where defects may exist in the solder joint image, improving the detection accuracy, especially the detection ability for small defects. These attention - enhanced features will be used as key components of the teacher - student network architecture in subsequent steps for the knowledge distillation process and feature fusion operations.

[0034] The knowledge distillation module, based on the multi - scale feature extraction network and the attention calculation module, constructs a teacher - student network architecture to transfer the knowledge of the high - precision teacher network to the lightweight student network; Based on the outputs of the image acquisition and feature extraction module and the attention calculation module, namely the multi - scale feature map set and the prior - knowledge - guided attention maps , constructs a teacher - student network architecture, and transfers the knowledge of the high - precision but complex teacher network to the lightweight student network through knowledge distillation technology, achieving model lightweight while maintaining high detection accuracy.

[0035] Step 3.1, construct the teacher network and the student network; The teacher network adopts a complete attention-guided multi-scale feature network, which includes all components and calculation processes described in the image acquisition and feature extraction module and the attention calculation module. It has a large number of parameters but high detection accuracy. Specifically, the teacher network contains 8 convolutional layers, and the number of channels increases layer by layer from 64 to 512. The total number of parameters is approximately 24 million. The teacher network uses a complete prior knowledge-guided attention module, which can accurately locate the possible defect areas in the solder joint image.

[0036] The student network adopts a simplified structure with fewer channels and a reduced network depth, but maintains a similar overall architecture to the teacher network. The student network contains 6 convolutional layers, and the number of channels increases layer by layer from 32 to 256. The total number of parameters is approximately 7 million, which is only about 30% of the teacher network. The student network retains the key components of the prior knowledge-guided attention module but simplifies the calculation process, achieving a reduction in computational complexity while maintaining detection accuracy.

[0037] In the solder joint defect detection scenario, the teacher network can run on a high-performance server for processing complex solder joint defect sample analysis and new model training; while the student network is deployed on the edge computing device of the production line for real-time solder joint defect detection. This architecture enables high-precision solder joint defect detection even with limited computing resources on the production line.

[0038] Step 3.2, construct a multi-level attention distillation module; For the attention maps of each layer of the teacher network and the student network, calculate the multi-level attention distillation loss. The calculation formula is: ; Where represents the multi-level attention distillation loss, and respectively represent the attention maps of the th layer of the teacher network and the student network, represents the separation symbol between two distributions; represents the KL divergence, which is used to measure the difference between two distributions; is the weight coefficient of the th layer, which is used to balance the contribution of the distillation loss of different layers; represents the accumulation starting from the first layer of the network, represents the total number of network layers for attention distillation, represents the accumulation over all layers.

[0039] In this way, the student network can learn an effective attention allocation strategy from the teacher network and focus on key defect areas.

[0040] During the multi-level attention distillation process, the weight coefficients of different layers are configured according to the actual solder joint defect types and sizes. For example, for detecting the defect of small solder joint dry joints, more attention is paid to the attention distillation of low-level features, so the value of the low level is increased; for detecting larger defects such as solder joint bridging, more attention is paid to the attention distillation of high-level features, and the value of the high level is increased accordingly. In practical applications, can be set as a trainable parameter, and the optimal weight configuration is automatically adjusted through training on a specific defect dataset.

[0041] Specifically, an alternating training strategy is adopted: First, train the teacher network using standard detection losses (such as focal loss, IoU loss, etc.) until convergence; Then, fix the parameters of the teacher network and train the student network using the same dataset, where the loss function includes both detection loss and distillation loss; Finally, fine-tune the student network to adapt to a specific deployment environment.

[0042] Experiments show that this training strategy can enable the student network to maintain a detection accuracy close to that of the teacher network while reducing the number of parameters by 70%, especially improving the detection rate of small solder joint dry joint defects.

[0043] The output of this module is a trained teacher network and a student network after knowledge distillation. The teacher network has high accuracy but a large number of parameters and is suitable for use in environments with sufficient resources; the student network maintains a similar detection accuracy but reduces the number of parameters, preparing for subsequent lightweight and practical deployment. The student network inherits the key knowledge of the teacher network, especially the detection ability and attention allocation strategy for different size defects, which will be further optimized in the subsequent steps.

[0044] The feature complementary module is used to construct a two-way propagation system for cross-scale feature complementarity in the student network to enhance the information exchange between different scale features; Based on the student network after knowledge distillation in the knowledge distillation module, a two-way propagation system for cross-scale feature complementarity is constructed in it to enhance the information exchange between different scale features and improve the feature representation ability. This system utilizes the hierarchical relationship of multi-scale features extracted in the image acquisition and feature extraction module, and makes the features of different scales complement and enhance each other through two-way information flow.

[0045] Step 4.1, construct a feature two-way propagation module; Construct a two-way feature propagation module from high level to low level and from low level to high level.

[0046] For the propagation from the high layer to the low layer, the high-level semantic information is transmitted to the low layer through an upsampling operation; For the propagation from the low layer to the high layer, the low-level detailed information is transmitted to the high layer through a learnable pooling operation.

[0047] The bidirectional propagation module is specifically implemented as a multi-level cascaded structure, with each level containing a top-down path and a bottom-up path. The top-down path is implemented by transposed convolution, which upsamples the high-level feature map by a factor of 2 and aligns the number of channels with to obtain the upsampled feature map . The bottom-up path is implemented by a learnable regional pooling layer, which learns the weight coefficients for each region of and then performs weighted pooling for dimensionality reduction to obtain the pooled feature map .

[0048] In the solder joint defect detection scenario, the bidirectional propagation module can effectively handle the feature expressions of different types of solder joint defects. For example, for the subtle scratch defects on the surface of the solder joint, the bottom-up path can transmit the key texture details in the low-level features to the high level, enhancing the sensitivity of the high-level features to subtle texture changes; for the structural defects such as abnormal solder joint shapes, the top-down path can transmit the shape semantic information in the high-level features to the low level, enabling the low-level features to have better shape recognition capabilities.

[0049] Step 4.2, cross-scale feature fusion; Based on the features of bidirectional propagation, the fused features are calculated, and the formula is: ; where represents the fused feature map, is the weight coefficient of the feature map of the th layer, represents the index of the feature layer, represents the th layer of the feature map; is the weight coefficient for the feature interaction between the th layer and the th layer, and are the indices of the feature layers and satisfy and ; represents the Hadamard product (element-wise multiplication) operation for feature interaction; represents the total number of feature layers participating in feature fusion; and respectively represent the feature maps of the th layer and the th layer, Represents the summation symbol.

[0050] In this way, the model can effectively fuse feature information at different scales, retaining both high-level semantic information and not losing low-level detail information.

[0051] In the actual solder joint defect detection application, this cross-scale feature fusion module can adaptively adjust the feature fusion strategy according to different defect types: For example, for the solder joint cold solder defect, which is characterized by the combined features of irregular solder joint edges and abnormal surface gloss, the fusion module will increase the weight of low-level features and to retain edge details, and at the same time increase the interaction weight between low-level and high-level features and to associate the semantic information of abnormal surface gloss; for the solder joint missing defect, it mainly increases the weight of high-level features and to enhance the semantic understanding ability. The experimental results show that compared with the feature fusion with fixed weights, this adaptive fusion strategy improves the average detection rate of various solder joint defects by more than 15%.

[0052] The output of this module is a student network with enhanced cross-scale feature complementarity ability. This network realizes the effective fusion of features at different scales through a two-way propagation system and generates fused features . This fused feature contains both high-level semantic information and does not lose low-level detail information, and can effectively meet the detection requirements of different types and sizes of solder joint defects. The enhanced student network will be further lightweighted in the lightweight optimization module to meet the deployment requirements of edge devices.

[0053] The lightweight optimization module is used to further lightweight optimize the student network so that it can operate efficiently on resource-constrained edge devices; Based on the student network with enhanced cross-scale feature interaction ability in the feature complementarity module, it is further lightweight optimized so that it can operate efficiently on resource-constrained edge devices. The lightweight optimization fully considers the importance of each module in the previous steps and optimizes it specifically to maintain key functions.

[0054] Step 5.1, Progressive channel pruning; Based on the channel importance metric for progressive channel pruning, the channel importance calculation formula is: ; where represents the importance score of the th channel, represents the absolute value symbol, represents the summation symbol, represents the weight parameter related to the th channel, and represents the position index in the weight matrix, represents the channel index.

[0055] Sort the channels according to the importance score, and gradually prune the channels with low importance until the target model size is reached or the performance requirement is met.

[0056] The specific implementation of progressive channel pruning adopts an iterative optimization strategy. Each iteration consists of three stages: importance evaluation, pruning execution, and fine-tuning recovery.

[0057] In the importance evaluation stage, use the trained student network model to infer the validation set, collect the activation value statistics of each channel in each layer, and calculate the channel importance score in combination with the weight parameters ; In the pruning execution stage, sort according to the importance score, and remove the 10% channels with the lowest importance in each layer each time, while adjusting the connection structure of the relevant layers; In the fine-tuning recovery stage, use a small amount of training data (20% of the original training set) to fine-tune the pruned model for 5 epochs to recover the model performance.

[0058] By repeatedly executing these three stages until the model size reaches the target value or the performance drops by more than the preset threshold (usually 5% of the performance of the original model).

[0059] In the application of solder joint defect detection, through progressive channel pruning, the number of parameters of the student network is further reduced from 7 million to 2.1 million, the model size is reduced from 28MB to 8.5MB, and the inference speed is increased by 2.3 times. For specific defect types, such as solder joint voids, more relevant channels can be retained during the pruning process according to the analysis results of the importance of detecting such defects. For example, the convolutional layer channels responsible for extracting low-level texture features are particularly important for void detection, so the pruning ratio of these layers will be reduced accordingly to ensure that the performance of specific defect detection is not affected.

[0060] Step 5.2, hardware-aware optimization; Automatically adjust the network structure and quantization parameters according to the characteristics of the target hardware platform (such as memory limit, computing power, power consumption requirements, etc.). Including: Select the appropriate quantization precision according to the characteristics of the hardware arithmetic unit; Optimize the parameter storage method according to the hardware memory architecture; Adjust the network layer structure according to the hardware parallel computing power.

[0061] Through these optimization measures, the lightweight model achieves the best operating efficiency on a specific hardware platform.

[0062] The specific implementation of hardware-aware optimization is divided into three steps: hardware performance analysis, model structure adaptation, and model quantization conversion. In the hardware performance analysis stage, benchmark testing tools are used to analyze parameters such as the characteristics of the computing units, memory bandwidth, and cache size of the target hardware platform (such as a specific model of ARM processor). In the model structure adaptation stage, the network structure is adjusted according to the hardware analysis results. For example, for a processor with a SIMD unit width of 128 bits, the number of channels in the convolutional layer is adjusted to a multiple of 32 to fully utilize SIMD parallelism; for a processor with a 32KB L1 cache, large convolutional layers are decomposed into multiple small convolutional layers to reduce cache miss rates. In the model quantization conversion stage, the best quantization scheme is selected according to the data precision supported by the hardware. For hardware that supports INT8 operations, the model is quantized to 8-bit integer precision; for hardware that only supports FP16, half-precision floating-point quantization is used.

[0063] In the actual deployment of the solder joint defect detection device, the following hardware-aware optimizations were carried out for the edge computing device (configured with an ARM Cortex-A72 processor and 2GB RAM) used in the circuit board production line: The student network was quantized to INT8 precision, and the model size was further reduced to 2.2MB. The number of channels in the convolutional layer was adjusted to a multiple of 16 to match the NEON SIMD unit of the processor. For the secondary cache size of the processor, the computational graph of the network was reorganized to reduce the number of memory accesses.

[0064] After optimization, the inference speed of the model on this edge device reached 25 frames per second, meeting the real-time detection requirements of the production line. At the same time, the energy consumption was reduced by 65%, extending the battery life of the device. Compared with the model without hardware-aware optimization, the optimized model can handle image resolutions 50% higher and can detect solder joint defects of smaller sizes.

[0065] The final output of this module is a fully optimized intelligent detection model for circuit board appearance defects. While maintaining high detection accuracy, this model reduces the number of parameters and the model size and is deeply optimized for a specific hardware platform. The number of parameters of the final model is reduced from 7 million in the student network of the knowledge distillation module to 2.1 million, and the model size is reduced from 28MB to 2.2MB. It can achieve a real-time detection speed of 25 frames per second on edge computing devices, meeting the actual application requirements of the production line. The optimized model effectively combines all the technological innovations in the image acquisition and feature extraction modules into the feature complementary module, including multi-scale feature extraction, prior knowledge-guided attention mechanism, knowledge distillation, and cross-scale feature complementarity, forming a complete and high-performance intelligent detection solution for PCB appearance defects.

[0066] A computer-readable storage medium for storing computer-readable instructions that, when read by a computer, can run an intelligent detection device for PCB appearance defects as described above.

[0067] Here, the present invention provides an implementation example: This implementation was actually applied on the PCB production line of an electronic component manufacturing enterprise. The main products produced by this enterprise are high-density multilayer PCBs for smartphones and wearable devices, with a large number of densely packed solder joints on the board, small sizes, and various types of solder joints, including ball grid array solder joints, surface mount solder joints, and through-hole solder joints, etc. The original defect detection system had the following problems: the missed detection rate of small-sized solder joint voids was as high as 35%; the generalization ability of the detection model among different batches of products was poor, and a large amount of labeled data needed to be collected and the model retrained every time the product was iterated; the model operation had high requirements for computing power and was difficult to run in real time on the embedded edge detection devices on the production line, resulting in the detection speed not meeting the production line rate requirements; the system had poor adaptability to newly emerging defect types.

[0068] To verify the effectiveness of this implementation, 5000 PCB solder joint images were collected from the production line, including normal solder joints and five types of typical defects: solder joint voids, bridging, insufficient solder, excessive solder, and component misalignment. The image data was divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. Image preprocessing included operations such as standardization, random cropping, and rotation enhancement. The distribution of various solder joint images in the dataset is shown in Table 1: Table 1: Distribution of Solder Joint Image Dataset

[0069] In this application example, the multi-scale feature extraction network was improved based on the ResNet50 architecture, with the input image size of 384×384×3 pixels, and the backbone network extracted feature maps of 5 scales. The prior knowledge base was constructed by analyzing 1000 standard solder joint samples collected during the enterprise's production process, including the geometric and texture features of 3 main types of solder joints. A partial data structure of the prior knowledge base is shown in Table 2: Table 2: Structure of Solder Joint Prior Knowledge Base and Parameters of Teacher-Student Network

[0070] Teacher-Student Network Architecture and Knowledge Distillation Implementation The teacher network adopts a complete attention-guided multi-scale feature network; the student network adopts a simplified structure. During the knowledge distillation process, the weight coefficients of attention distillation for different layers are adjusted according to the detection requirements for defects of different sizes, as shown in Table 3: Table 3: Multi-level attention distillation weights and pruning results

[0071] The distillation training adopts a phased strategy: First, train the teacher network for 50 epochs on 3000 training images; then fix the parameters of the teacher network and train the student network for 100 epochs, where the distillation loss weight is 0.6 and the detection loss weight is 0.4; finally, fine-tune the student network for 20 epochs on 500 images of the enterprise's specific production line.

[0072] Progressive channel pruning and hardware-aware optimization are performed on the student network to enable it to run on edge devices with an ARM Cortex-A72 processor on the production line. The channel pruning adopts 5 rounds of iteration, removing 10% of the least important channels in each round and then performing fine-tuning.

[0073] After INT8 quantization and other hardware-aware optimizations, the final model size is reduced to 2.2MB, achieving a detection speed of 25 frames per second on the target hardware platform, meeting the operation rate requirement of 30 meters per minute on the production line.

[0074] The detection method of this embodiment is compared and tested with the enterprise's original method and two other mainstream methods, and the results are shown in Table 4: Table 4: Performance comparison of different detection methods

[0075] To verify the adaptability of this embodiment to new batches of products, we directly tested it on three batches of new product circuit boards without retraining the model, and the results are shown in Table 5: Table 5: Detection rates and production line application effects of different methods on new batches of products

[0076] The experimental results show that the decline in the detection rate of this embodiment on new batches of products is lower than that of other methods, verifying its excellent adaptability. At the same time, after 6 months of application in the actual production environment, both the product quality and production efficiency have been improved, as shown in Table 6: Table 6: Summary of the core technical indicators of this embodiment

[0077] In summary, in the actual production environment, this embodiment not only improves the detection rate of small defects, but also improves the adaptability of the model to new batches of products. At the same time, it realizes efficient operation on edge devices, meeting the real-time detection requirements of the production line.

[0078] The embodiments of the present invention have been described above. However, these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. An intelligent detection device for the appearance defects of a circuit board, characterized in that, Including: An image acquisition and feature extraction module, which is used to construct a multi-scale feature extraction network and a solder joint prior knowledge base, and extract multi-level features of the circuit board solder joint images; An attention calculation module, which is based on the multi-level features output by the multi-scale feature extraction network and the solder joint prior knowledge base, performs knowledge-guided attention calculation on the multi-scale feature maps, and focuses on potential defect areas; A knowledge distillation module, which is based on the multi-scale feature extraction network and the attention calculation module, constructs a teacher-student network architecture, and transfers the knowledge of the high-precision teacher network to the lightweight student network; A feature complementary module, which is used to construct a bidirectional propagation system for cross-scale feature complementarity in the student network, and enhance the information exchange between features of different scales; A lightweight optimization module, which is used to further optimize the student network in a lightweight manner so that it can operate efficiently on resource-constrained edge devices.

2. The intelligent detection device for appearance defects of a circuit board according to claim 1, characterized in that, The solder joint prior knowledge base contains the geometric shapes, size ratios, surface textures of different types of solder joints, and visual feature information under normal welding conditions, and forms a prior knowledge vector: ; wherein represents the prior knowledge vector of the th solder joint type, represents the solder joint type index, represents the dimension of the prior knowledge vector, represents the real number field; The prior knowledge base is expressed as: ; Among them represents the complete set of prior knowledge bases for solder joints , , respectively represent the prior knowledge vectors of the first, second, and th solder joint types, is the total number of solder joint types 3. An intelligent detection device for the appearance defects of a circuit board according to claim 1, characterized in that, The attention calculation module includes: Solder joint type recognition and prior knowledge embedding; Attention map calculation.

4. An intelligent detection device for the appearance defects of a circuit board according to claim 3, characterized in that, The calculation formula for the prior knowledge embedding is: ; Among them represents the prior knowledge embedding, represents the prior knowledge vector of the th type of solder joint, and represents the learnable embedding matrix.

5. The intelligent detection device for appearance defects of a circuit board according to claim 1, characterized in that, The construction of the teacher-student network architecture includes constructing a multi-level attention distillation module. For the attention maps of each layer of the teacher network and the student network, calculate the multi-level attention distillation loss, and the calculation formula is: ; Among them represents the multi-level attention distillation loss and respectively represent the attention maps of the th layer of the teacher network and the student network represents the separation symbol between two distributions represents the KL divergence, which is used to measure the difference between two distributions is the weight coefficient of the th layer, which is used to balance the contribution of the distillation loss of different layers represents the accumulation starting from the first layer of the network represents the total number of network layers for attention distillation represents the accumulation over all layers 6. The intelligent detection device for the appearance defects of a circuit board according to claim 1, characterized in that, The construction of the bidirectional propagation system for cross-scale feature complementarity includes: Constructing a feature bidirectional propagation module, which is used to construct bidirectional feature propagation from high level to low level and from low level to high level; Cross-scale feature fusion, which is used to fuse feature information of different scales, retain high-level semantic information, and not lose low-level detail information.

7. An intelligent detection device for the appearance defects of a circuit board according to claim 6, characterized in that, The calculation formula for the cross-scale feature fusion is: ; Among them represents the fused feature map, is the weight coefficient of the -th layer of features, represents the index of the feature layer, represents the -th layer of feature map; is the weight coefficient of the interaction between the -th layer and the -th layer of features, and are the indices of the feature layers and satisfy and ; represents the Hadamard product (element-wise multiplication) operation for implementing feature interaction; represents the total number of feature layers participating in feature fusion; and respectively represent the features of the -th layer and the -th layer, represents the summation symbol.

8. An intelligent detection device for the appearance defects of a circuit board according to claim 1, characterized in that, The lightweight optimization module includes: Progressive channel pruning based on channel importance measurement, and the calculation formula for channel importance is: ; Among them represents the importance score of the th channel, represents the absolute value symbol, represents the summation symbol, represents the weight parameter related to the th channel, and represent the position index in the weight matrix, represents the channel index; Performing hardware-aware optimization according to the characteristics of the target hardware platform.

9. An intelligent detection device for the appearance defects of a circuit board according to claim 8, characterized in that, The performing hardware-aware optimization according to the characteristics of the target hardware platform includes: Selecting an appropriate quantization precision according to the characteristics of the hardware arithmetic unit; Optimizing the parameter storage method according to the hardware memory architecture; Adjusting the network layer structure according to the hardware parallel computing ability.

10. A computer-readable storage medium, characterized in that, It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run an intelligent detection device for circuit board appearance defects as described in any one of claims 1-9.

Citation Information

Cited By

  • Intelligent chip array defect detection system based on image processing

    CN120563495A

  • AI appearance defect identification method for multiple acquisition terminals

    CN120894318A