PCB defect automatic grading method based on image recognition

CN122736991APending Publication Date: 2026-09-11BEIHAN (BEIJING) IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610841869.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0006]本发明的一个目的在于提出一种基于图像识别的PCB缺陷自动分级方法,针对现有技术在PCB缺陷的自动分级和精确识别方面存在自动分级精度不足的问题,提出了结合量子图神经网络与视觉Transformer架构的新型方法

Benefits of technology

本发明相较于以卷积神经网络为核心的缺陷识别方法,采用视觉Transformer模型对图像块序列执行自注意力计算与编码器特征聚合,将全局结构关联关系纳入特征学习过程,从而在电路板走线密集、纹理相近、背景干扰较强的场景下,能够对跨区域的形态一致性与结构约束进行统一建模,减少仅依赖局部卷积感受野导致的结构信息割裂问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736991A_ABST
    Figure CN122736991A_ABST
Patent Text Reader

Abstract

The application discloses a PCB defect automatic grading method based on image recognition, which comprises the following steps: collecting and preprocessing PCB image data; constructing a visual Transformer model based on a self-attention mechanism to obtain overall structure information of the image; introducing a quantum graph neural network into the visual Transformer model, optimizing the self-attention weight calculation process, and combining hypergraph convolution to model the spatial correlation between image blocks; extracting global structure information and local defect features based on the optimized model; classifying and regressing defects by using a convolutional neural network to generate defect grading results; and automatically adjusting the visual Transformer model parameters according to the grading results and forming feedback optimization. The application realizes closed-loop processing of defect grading and model updating, and improves the stability and consistency of PCB defect grading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and deep learning technology, and in particular to an automatic PCB defect classification method based on image recognition. Background Technology

[0002] With the widespread application of electronic devices, printed circuit boards (PCBs), as a crucial component of electronic products, directly impact the performance and reliability of these devices through their manufacturing quality. PCB defect detection, a key step in the production process, can effectively improve production efficiency and product quality. However, traditional manual inspection methods, relying on human experience, are easily influenced by subjective factors, resulting in low accuracy and efficiency, failing to meet the modern electronics industry's requirements for high precision, high efficiency, and high reliability. Therefore, the research and application of automated PCB defect detection technology has become an important development direction in the industry.

[0003] In recent years, the combination of image processing technology and deep learning algorithms has provided new approaches to PCB defect detection. Image processing technology can analyze PCB images using computer vision techniques to identify and locate defects. Deep learning, especially models such as convolutional neural networks (CNNs), has been widely applied in tasks such as image classification and object detection due to its powerful feature extraction and pattern recognition capabilities, demonstrating excellent performance. By training deep learning models, defect features in PCB images can be automatically extracted, thereby achieving defect identification and classification.

[0004] However, existing PCB defect detection methods based on models such as Convolutional Neural Networks (CNNs) and Visual Transformers (ViTs), while achieving good results in many applications, still have some problems. First, traditional deep learning models are easily affected by complex background noise or small defects, leading to insufficient recognition accuracy. Second, existing models typically focus only on local features in the image, failing to fully utilize the global contextual information. With the increasing size and complexity of PCBs, the detection of local features alone can no longer meet high-precision requirements; finding a balance between global structure and local defect features has become an urgent problem to be solved.

[0005] Therefore, how to provide an automatic PCB defect classification method based on image recognition is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an automatic PCB defect classification method based on image recognition. Addressing the issue of insufficient accuracy in automatic PCB defect classification and precise identification in existing technologies, this invention proposes a novel method combining quantum graph neural networks and a visual Transformer architecture. By optimizing the attention weight calculation process in the self-attention mechanism through quantum computing and introducing hypergraph convolution to model the higher-order spatial dependencies of defect regions in the image, the model's ability to focus on important features in the image is enhanced. This invention, through the combination of quantum computing and visual Transformer, improves the accuracy and efficiency of PCB defect detection, overcomes the problem of insufficient attention to local features in traditional methods, and exhibits higher accuracy and reliability.

[0007] An automatic PCB defect classification method based on image recognition according to an embodiment of the present invention includes the following steps: S1: Acquire PCB image data and preprocess the image data to obtain preprocessed image data; S2: Based on the preprocessed image data, a visual Transformer model is constructed. The visual Transformer model learns global features of the image through a self-attention mechanism, extracts the overall structural information of the image, and forms a visual Transformer model containing an input layer, a self-attention structure, and an encoder structure. S3: Based on the visual Transformer model, a quantum graph neural network is introduced to optimize the self-attention mechanism of the visual Transformer model. The quantum graph neural network optimizes the calculation process of attention weights through quantum computing, utilizes the quantum superposition and interference properties to enhance the visual Transformer model's attention to important features, and enhances the high-order spatial dependencies between defective regions in the image through hypergraph convolution, thus obtaining the optimized visual Transformer model. S4: Based on the optimized visual Transformer model, feature extraction is performed on the image to generate a feature representation containing global structural information and local defect features; S5: Based on the feature representation of the extracted global structural information and local defect features, a convolutional neural network is used to classify and regress the defects. The classification and regression steps generate the graded results of the defects according to their type, shape and location. S6: Based on the defect classification results, automatically adjust the parameters of the visual Transformer model and further optimize the Transformer model through a feedback mechanism.

[0008] Optionally, S1 specifically includes: S1.1: Acquire PCB image data using imaging equipment; S1.2: The acquired PCB image data is denoised, standardized, and enhanced to generate preprocessed image data.

[0009] Optionally, S2 specifically includes: S2.1: Based on the preprocessed image data, the image data is divided into blocks according to the preset spatial division rules. Each image block is represented as a two-dimensional matrix containing pixel information. S2.2: Generate a position code for each image patch, map the position code into a vector form, and fuse the position code vector with the pixel matrix of the corresponding image patch to form an image patch embedding representation containing spatial position information; S2.3: Input the image patch embedding representations into the input layer of the visual Transformer model in sequence. Perform linear mapping and dimensionality unification processing on the image patch embedding representations in the input layer to generate image patch feature representations for attention calculation. S2.4: Input the image patch feature representation into the multi-layer self-attention structure in the visual Transformer model, calculate the correlation between image patches in the self-attention structure, and perform weighted combination of image patch feature representation based on the correlation to form an intermediate feature representation containing cross-image patch information; S2.5: The intermediate feature representations are sequentially input into multiple Transformer encoders. In the Transformer encoders, through multi-layer information propagation and feature aggregation operations, an image representation that characterizes the overall structural information and defect-related features of the image is generated. S2.6: Form a visual Transformer model that includes an input layer, a self-attention structure, and an encoder structure.

[0010] Optionally, S3 specifically includes: S3.1: In the visual Transformer model, the image patch feature representations output by the input layer are organized into a set of nodes according to the spatial correspondence of the image patches in the original image. The set of nodes is then input into the quantum graph neural network. The quantum graph neural network performs quantum computation processing on the image patch features in the set of nodes to generate intermediate computation results for participating in the self-attention weight calculation of the visual Transformer model. S3.2: In the process of self-attention weight calculation, the quantum graph neural network is based on the quantum parallel computing process to perform parallel calculation on the similarity between the corresponding features of any two image patches in the node set, and the similarity result obtained by parallel calculation is mapped to the attention weight parameter set. The attention weight parameter set is used to replace or correct the original self-attention weight calculation result of the visual Transformer model. S3.3: Based on the attention weight parameter set, the connection relationship between each image block node and the other image block nodes is updated. The updated connection relationship is used to determine the weight allocation method of image block features in the self-attention weighted calculation process, thereby forming a weighted connection relationship between image blocks for subsequent feature update calculation. S3.4: Based on the weighted connectivity between image patches, the quantum graph neural network performs hypergraph convolution operation on the features of multiple interconnected image patches. The hypergraph convolution operation generates a high-order feature representation that characterizes the spatial correlation between image patches by jointly aggregating the features of multiple image patches under the same weighted connectivity. S3.5: Inject the set of attention weight parameters into the multi-layer self-attention calculation process of the visual Transformer model, and introduce the high-order feature representation into the feature update path of the visual Transformer model to update the image patch feature representation, thus forming the optimized visual Transformer model.

[0011] Optionally, S4 specifically includes: S4.1: Based on the optimized visual Transformer model, the image representation is input to the feature extraction layer of the visual Transformer model. The feature extraction layer processes the input image features layer by layer through a multi-layer neural network structure. S4.2: During the feature extraction process, the visual Transformer model uses a self-attention mechanism to dynamically adjust the attention weights and perform weighted aggregation on each region in the image. S4.3: The features of each image patch are aggregated through a multi-layer neural network structure, and global and local features are combined to generate a feature vector containing global structural information of the image. The feature vector represents the global layout and detailed information in the image. S4.4: While extracting global features, the visual Transformer model processes local defect features separately, learns from local features to extract features of defect regions, and generates feature representations containing local defect information. S4.5: The feature representations of global structural information and local defect features are fused to obtain the final image feature representation, which contains detailed information on the overall structure and local defects in the image.

[0012] Optionally, S5 specifically includes: S5.1: Based on the feature representation of the global structural information and local defect features extracted in step S4, the feature representation is input to the input layer of the convolutional neural network structure, which includes multiple convolutional layers, pooling layers and activation functions; S5.2: In the convolutional layer, a convolutional kernel is used to perform a convolution operation on the input image features. The convolution operation includes filtering the image feature map to obtain local pattern features. S5.3: In the pooling layer, a pooling operation is performed to downsample the convolutional feature map. The pooling operation reduces the spatial dimension of the feature map by selecting the maximum or average value in a local region and retains the main features in the image. S5.4: In the activation layer, a non-linear activation function, namely the ReLU function, is applied to the pooled feature map to perform a non-linear transformation on the extracted features, thereby obtaining a feature representation with non-linear expressive power; S5.5: The processed features are flattened through a fully connected layer to obtain a one-dimensional feature vector, which contains global structural information and feature information of defect areas in the image; S5.6: The flattened feature vector is input into the classification and regression network, which generates a classification result of the defect based on the type, shape and location of the defect.

[0013] Optionally, the classification results of the defects specifically include: The defect type information, defect morphology parameters, and defect spatial location information output by the classification and regression network are combined according to the preset grading rules to form a multi-dimensional grading result representation that includes defect category identifier, defect severity level, and defect spatial distribution description.

[0014] Optionally, S6 specifically includes: S6.1: Based on the defect classification results, extract defect classification and grading information, including the type, shape and location of the defect. The defect classification results are the input information for subsequent model adjustments. S6.2: Based on the defect classification results, the hyperparameters of the visual Transformer model are automatically adjusted; S6.3: After automatically adjusting the hyperparameters, train the visual Transformer model using the adjusted hyperparameters and update the model's weights and biases; S6.4: Through the feedback mechanism, the newly trained visual Transformer model is applied to image data, and the performance of the model is evaluated to obtain the evaluation index results; S6.5: Based on the evaluation results, further adjust the hyperparameters of the visual Transformer model.

[0015] The beneficial effects of this invention are: Compared to defect recognition methods based on convolutional neural networks, this invention employs a visual Transformer model to perform self-attention calculation and encoder feature aggregation on image block sequences, incorporating global structural relationships into the feature learning process. This enables unified modeling of cross-regional morphological consistency and structural constraints in scenarios with dense circuit board traces, similar textures, and strong background interference, reducing the problem of fragmented structural information caused by relying solely on local convolutional receptive fields.

[0016] This invention introduces a quantum graph neural network into the self-attention weight calculation path of the visual Transformer model. It uses quantum parallel computing to perform parallel solving of the similarity relationship between image patches and maps the solution results into a set of attention weight parameters to replace or correct the original attention weight calculation results. At the same time, it performs joint aggregation of the high-order correlation relationship of multiple image patches through hypergraph convolution to form a high-order feature representation and participate in the feature update process. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the process of an automatic PCB defect classification method based on image recognition proposed in this invention; Figure 2 This is a schematic diagram of the structure of the quantum graph neural network and the visual Transformer self-attention fusion in the automatic PCB defect classification method based on image recognition proposed in this invention. Detailed Implementation

[0018] Combination Figures 1-2 The present invention will be described in further detail below. These accompanying drawings are simplified schematic diagrams, illustrating only the basic structure of the invention and showing the main components relevant to the invention. Figure 1 and Figure 2 The present invention provides an automatic PCB defect classification method based on image recognition: S1: Acquire PCB image data and preprocess the image data to obtain preprocessed image data; S2: Based on the preprocessed image data, a visual Transformer model is constructed. The visual Transformer model learns global features of the image through a self-attention mechanism, extracts the overall structural information of the image, and forms a visual Transformer model containing an input layer, a self-attention structure, and an encoder structure. S3: Based on the visual Transformer model, a quantum graph neural network is introduced to optimize the self-attention mechanism of the visual Transformer model. The quantum graph neural network optimizes the calculation process of attention weights through quantum computing, utilizes the quantum superposition and interference properties to enhance the visual Transformer model's attention to important features, and enhances the high-order spatial dependencies between defective regions in the image through hypergraph convolution, thus obtaining the optimized visual Transformer model. S4: Based on the optimized visual Transformer model, feature extraction is performed on the image to generate a feature representation containing global structural information and local defect features; S5: Based on the feature representation of the extracted global structural information and local defect features, a convolutional neural network is used to classify and regress the defects. The classification and regression steps generate the graded results of the defects according to their type, shape and location. S6: Based on the defect classification results, automatically adjust the parameters of the visual Transformer model and further optimize the Transformer model through a feedback mechanism.

[0019] In this embodiment, S1 specifically includes: S1.1: Acquire PCB image data using imaging equipment; S1.2: The acquired PCB image data is denoised, standardized, and enhanced to generate preprocessed image data.

[0020] In this embodiment, S2 specifically includes: S2.1: Based on the preprocessed image data, the image data is divided into blocks according to the preset spatial division rules. Each image block is represented as a two-dimensional matrix containing pixel information. S2.2: Generate a position code for each image patch, map the position code into a vector form, and fuse the position code vector with the pixel matrix of the corresponding image patch to form an image patch embedding representation containing spatial position information; S2.3: Input the image patch embedding representations into the input layer of the visual Transformer model in sequence. Perform linear mapping and dimensionality unification processing on the image patch embedding representations in the input layer to generate image patch feature representations for attention calculation. S2.4: Input the image patch feature representation into the multi-layer self-attention structure in the visual Transformer model, calculate the correlation between image patches in the self-attention structure, and perform weighted combination of image patch feature representation based on the correlation to form an intermediate feature representation containing cross-image patch information; S2.5: The intermediate feature representations are sequentially input into multiple Transformer encoders. In the Transformer encoders, through multi-layer information propagation and feature aggregation operations, an image representation that characterizes the overall structural information and defect-related features of the image is generated. S2.6: Form a visual Transformer model that includes an input layer, a self-attention structure, and an encoder structure.

[0021] In this embodiment, S3 specifically includes: S3.1: In the visual Transformer model, the image patch feature representations output by the input layer are organized into a set of nodes according to the spatial correspondence of the image patches in the original image. The set of nodes is then input into the quantum graph neural network. The quantum graph neural network performs quantum computation processing on the image patch features in the set of nodes to generate intermediate computation results for participating in the self-attention weight calculation of the visual Transformer model. S3.2: In the process of self-attention weight calculation, the quantum graph neural network is based on the quantum parallel computing process to perform parallel calculation on the similarity between the corresponding features of any two image patches in the node set, and the similarity result obtained by parallel calculation is mapped to the attention weight parameter set. The attention weight parameter set is used to replace or correct the original self-attention weight calculation result of the visual Transformer model. S3.3: Based on the attention weight parameter set, the connection relationship between each image block node and the other image block nodes is updated. The updated connection relationship is used to determine the weight allocation method of image block features in the self-attention weighted calculation process, thereby forming a weighted connection relationship between image blocks for subsequent feature update calculation. S3.4: Based on the weighted connectivity between image patches, the quantum graph neural network performs hypergraph convolution operation on the features of multiple interconnected image patches. The hypergraph convolution operation generates a high-order feature representation that characterizes the spatial correlation between image patches by jointly aggregating the features of multiple image patches under the same weighted connectivity. S3.5: Inject the set of attention weight parameters into the multi-layer self-attention calculation process of the visual Transformer model, and introduce the high-order feature representation into the feature update path of the visual Transformer model to update the image patch feature representation, thus forming the optimized visual Transformer model.

[0022] In this embodiment, S4 specifically includes: S4.1: Based on the optimized visual Transformer model, the image representation is input to the feature extraction layer of the visual Transformer model. The feature extraction layer processes the input image features layer by layer through a multi-layer neural network structure. S4.2: During the feature extraction process, the visual Transformer model uses a self-attention mechanism to dynamically adjust the attention weights and perform weighted aggregation on each region in the image. S4.3: The features of each image patch are aggregated through a multi-layer neural network structure, and global and local features are combined to generate a feature vector containing global structural information of the image. The feature vector represents the global layout and detailed information in the image. S4.4: While extracting global features, the visual Transformer model processes local defect features separately, learns from local features to extract features of defect regions, and generates feature representations containing local defect information. S4.5: The feature representations of global structural information and local defect features are fused to obtain the final image feature representation, which contains detailed information on the overall structure and local defects in the image.

[0023] In this embodiment, S5 specifically includes: S5.1: Based on the feature representation of the global structural information and local defect features extracted in step S4, the feature representation is input to the input layer of the convolutional neural network structure, which includes multiple convolutional layers, pooling layers and activation functions; S5.2: In the convolutional layer, a convolutional kernel is used to perform a convolution operation on the input image features. The convolution operation includes filtering the image feature map to obtain local pattern features. S5.3: In the pooling layer, a pooling operation is performed to downsample the convolutional feature map. The pooling operation reduces the spatial dimension of the feature map by selecting the maximum or average value in a local region and retains the main features in the image. S5.4: In the activation layer, a non-linear activation function, namely the ReLU function, is applied to the pooled feature map to perform a non-linear transformation on the extracted features, thereby obtaining a feature representation with non-linear expressive power; S5.5: The processed features are flattened through a fully connected layer to obtain a one-dimensional feature vector, which contains global structural information and feature information of defect areas in the image; S5.6: The flattened feature vector is input into the classification and regression network, which generates a classification result of the defect based on the type, shape and location of the defect.

[0024] In this embodiment, the classification result of the defect specifically includes: The defect type information, defect morphology parameters, and defect spatial location information output by the classification and regression network are combined according to the preset grading rules to form a multi-dimensional grading result representation that includes defect category identifier, defect severity level, and defect spatial distribution description.

[0025] In this embodiment, S6 specifically includes: S6.1: Based on the defect classification results, extract defect classification and grading information, including the type, shape and location of the defect. The defect classification results are the input information for subsequent model adjustments. S6.2: Based on the defect classification results, the hyperparameters of the visual Transformer model are automatically adjusted; S6.3: After automatically adjusting the hyperparameters, train the visual Transformer model using the adjusted hyperparameters and update the model's weights and biases; S6.4: Through the feedback mechanism, the newly trained visual Transformer model is applied to image data, and the performance of the model is evaluated to obtain the evaluation index results; S6.5: Based on the evaluation results, further adjust the hyperparameters of the visual Transformer model.

[0026] Example 1: To verify the feasibility of the present invention in practice, this example focuses on an online inspection scenario for quality control in electronic manufacturing processes. In this scenario, the imaging device continuously captures images of the surface of the printed circuit board under test at a fixed working distance, forming a PCB image data stream that can be used for defect classification. The system needs to complete defect type determination, defect morphology parameter fitting, and defect spatial location regression without changing the existing production line cycle time, and output a multi-dimensional classification result representation that meets the preset classification rules, thereby supporting the automatic triggering and recording of subsequent handling strategies.

[0027] During implementation, the PCB image data acquired by the imaging equipment was uniformly converted into a grayscale and three-channel parallel input format, with the original resolution set to 1024 x 1024 pixels. In the preprocessing stage, the images underwent denoising, standardization, and enhancement. Denoising employed a concatenated method of median filtering with a 3x3 window and wavelet thresholding, with two rounds of median filtering iterations. The wavelet decomposition layer was set to three levels, and a soft thresholding strategy was used, with noise intensity estimated based on local variance. Standardization employed linear normalization based on the image mean and standard deviation, converging the pixel mean of each image to 0.5 and the standard deviation to 0.2. Enhancement used a combination strategy of random rotation angles ranging from -5 to +5 degrees, random translation ratios ranging from 0 to 3%, random contrast perturbation coefficients ranging from 0.9 to 1.1, and random brightness perturbation coefficients ranging from 0.95 to 1.05. Four enhanced samples were generated for each image, which, together with the original samples, constituted the training input.

[0028] In constructing the visual Transformer model, the preprocessed image data is divided into blocks, with a spatial partitioning rule of 16x16 pixels per block, resulting in 4096 image blocks. Each image block is represented as a two-dimensional matrix containing pixel information. A positional code is generated for each image block, mapped to a vector of length 768. This vector is then fused with the embedding vector obtained by linear mapping of the corresponding image block's pixel matrix, forming an image block embedding representation containing spatial positional information. This image block embedding representation is input sequentially to the input layer of the visual Transformer model. The input layer performs linear mapping and dimensionality unification on the image block embedding representation, maintaining the sequence feature dimension at 768 and outputting the image block feature representation. The image block feature representation then enters a multi-layer self-attention structure. The self-attention structure has twelve layers and twelve attention heads, each with a dimension of 64. The self-attention structure calculates the relationships between image blocks and weights and combines the image block feature representations to form intermediate feature representations. The intermediate features are sequentially input into multiple Transformer encoders. The Transformer encoders use a feedforward network with a hidden dimension of 3072 and a dropout probability of 0.1. After twelve layers of information propagation and feature aggregation, they output an image representation that characterizes the overall structural information of the image and the defect-related features, forming a visual Transformer model that includes an input layer, a self-attention structure, and an encoder structure.

[0029] When introducing a quantum graph neural network to optimize the self-attention mechanism of the visual Transformer model, the image patch feature representation output from the input layer of the visual Transformer model is organized into a node set according to the spatial correspondence of the image patches in the original image, with a node set size of 4096. The quantum graph neural network performs quantum computation processing on the image patch features in the node set, with the quantum state feature encoding dimension set to eight, the number of parallel channels in the quantum parallel computation process set to 128, and the quantum interference modulation depth set to four rounds. The quantum graph neural network performs parallel computation on the similarity between corresponding features of any two image patches in the node set, using a normalized inner product form for the similarity function. The similarity result is converted into an attention weight parameter set through a mapping function. The attention weight parameter set is replaced or corrected by a weighted fusion method with the original self-attention weight calculation result of the visual Transformer model, with a fusion coefficient set to 0.6. Subsequently, the connection relationship between image patch nodes is updated based on the attention weight parameter set. The connection relationship participates in the weight allocation in the self-attention weighted calculation process in the form of a weighted connection relationship, forming a weighted connection relationship between image patches. The quantum graph neural network constructs a hypergraph structure based on the weighted connectivity between image patches. The hyperedge construction rule is set to be constrained by a spatial neighborhood radius of 24 pixels and a similarity threshold of 0.7. The number of image patch nodes connected by each hyperedge is set to eight to sixteen. The hypergraph convolution operation adopts a two-layer stacked structure. The aggregation weight of the hypergraph convolution is obtained by normalizing the weighted connectivity, resulting in a high-order feature representation that represents the spatial relationship between image patches. The attention weight parameter set is injected into the multi-layer self-attention calculation process of the visual Transformer model. The high-order feature representation is introduced into the feature update path of the visual Transformer model, concatenated with the intermediate feature representation output by the Transformer encoder, and then linearly mapped back to 768 dimensions to complete the update of the image patch feature representation, thus obtaining the optimized visual Transformer model.

[0030] In the feature extraction stage, the optimized visual Transformer model inputs the image representation into the feature extraction layer and processes it layer by layer through a multi-layer neural network structure. A self-attention mechanism weights and aggregates image regions to obtain a feature vector containing global image structural information, with the feature vector dimension set to 1024. To generate local defect features, the feature extraction layer retains a quarter-scale feature map on the corresponding spatial resolution branch and performs local feature learning. The local feature learning window size is set to 7x7 and slides with a stride of 2, outputting a feature representation containing local defect information, with the number of local feature channels set to 256. In the fusion stage, the feature representations of global structural information and local defect features are fused using channel concatenation and 1x1 convolution compression. After compression, the number of channels is set to 512, yielding the final image feature representation. This image feature representation simultaneously preserves the detailed information of both the overall structure and local defects.

[0031] In the defect classification and regression stage, image feature representations are input to the input layer of a convolutional neural network (CNN) structure. The CNN structure uses an alternating configuration of five convolutional layers and three pooling layers. The kernel sizes are set sequentially to 3x3, 3x3, 3x3, 3x3, and 3x3, with the number of channels set sequentially to 64, 128, 256, 256, and 512. The pooling layers use 2x2 downsampling. The activation layer uses the ReLU function and is executed after each convolutional layer to obtain high-dimensional features for classification. The classification and regression network constructs a dual-branch head structure on the output features of the CNN. The classification branch outputs a defect type probability vector, with six categories. The regression branch outputs defect morphological parameters and defect spatial location information. The defect morphological parameters include four continuous variables: defect area, defect perimeter, defect principal axis length, and defect secondary axis length. The defect spatial location information is represented by four continuous variables: the coordinates of the defect center point and the width and height of the defect's circumscribed rectangle. The defect classification results are generated based on preset classification rules. The preset classification rules combine defect type information, defect morphology parameters and defect spatial location information into a multi-dimensional classification result representation that includes defect category identifier, defect severity level and defect spatial distribution description. The defect severity level is determined by a weighted score of defect area and defect principal axis length, with weighting coefficients set to 0.6 and 0.4. The scoring range is divided into five levels and mapped to level one to level five.

[0032] The training and feedback optimization phase adopted a combination strategy of fixed training rounds and feedback triggering. The initial training rounds were set to sixty rounds, and the batch size was set to sixteen. On the test set, the defect type classification accuracy was 0.961, the defect severity level consistency rate was 0.934, the mean error of defect center point position was 1.8 pixels, the mean error of defect area relative error was 4.7%, and the mean error of defect main axis length relative error was 3.9%, which met the requirements of the scene for consistency of hierarchical output and stability of localization regression.

[0033] Table 1 Comparison of Defect Classification Results and Model Predictions

[0034] As shown in Table 1, in the five selected sample groups, the model consistently matched the measured results in both defect category and defect level prediction, with spatial location errors kept within a small range. This indicates that the present invention has good engineering applicability in terms of defect localization and classification stability. During multiple rounds of iterative optimization, the attention weight allocation of the visual Transformer model, with the participation of a quantum graph neural network, gradually converged, and the fluctuation amplitude of the defect classification results significantly decreased, verifying the effectiveness and repeatability of the technical solution of the present invention in practical application scenarios.

Claims

1. An automatic PCB defect classification method based on image recognition, characterized in that, Includes the following steps: S1: Acquire PCB image data and preprocess the image data to obtain preprocessed image data; S2: Based on the preprocessed image data, a visual Transformer model is constructed. The visual Transformer model learns global features of the image through a self-attention mechanism, extracts the overall structural information of the image, and forms a visual Transformer model containing an input layer, a self-attention structure, and an encoder structure. S3: Based on the visual Transformer model, a quantum graph neural network is introduced to optimize the self-attention mechanism of the visual Transformer model. The quantum graph neural network optimizes the calculation process of attention weights through quantum computing, utilizes the quantum superposition and interference properties to enhance the visual Transformer model's attention to important features, and enhances the high-order spatial dependencies between defective regions in the image through hypergraph convolution, thus obtaining the optimized visual Transformer model. S4: Based on the optimized visual Transformer model, feature extraction is performed on the image to generate a feature representation containing global structural information and local defect features; S5: Based on the feature representation of the extracted global structural information and local defect features, a convolutional neural network is used to classify and regress the defects. The classification and regression steps generate the graded results of the defects according to their type, shape and location. S6: Based on the defect classification results, automatically adjust the parameters of the visual Transformer model and further optimize the Transformer model through a feedback mechanism.

2. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S1 specifically includes: S1.1: Acquire PCB image data using imaging equipment; S1.2: The acquired PCB image data is denoised, standardized, and enhanced to generate preprocessed image data.

3. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S2 specifically includes: S2.1: Based on the preprocessed image data, the image data is divided into blocks according to the preset spatial division rules. Each image block is represented as a two-dimensional matrix containing pixel information. S2.2: Generate a position code for each image patch, map the position code into a vector form, and fuse the position code vector with the pixel matrix of the corresponding image patch to form an image patch embedding representation containing spatial position information; S2.3: Input the image patch embedding representations into the input layer of the visual Transformer model in sequence. Perform linear mapping and dimensionality unification processing on the image patch embedding representations in the input layer to generate image patch feature representations for attention calculation. S2.4: Input the image patch feature representation into the multi-layer self-attention structure in the visual Transformer model, calculate the correlation between image patches in the self-attention structure, and perform weighted combination of image patch feature representation based on the correlation to form an intermediate feature representation containing cross-image patch information; S2.5: The intermediate feature representations are sequentially input into multiple Transformer encoders. In the Transformer encoders, through multi-layer information propagation and feature aggregation operations, an image representation that characterizes the overall structural information and defect-related features of the image is generated. S2.6: Form a visual Transformer model that includes an input layer, a self-attention structure, and an encoder structure.

4. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S3 specifically includes: S3.1: In the visual Transformer model, the image patch feature representations output by the input layer are organized into a set of nodes according to the spatial correspondence of the image patches in the original image. The set of nodes is then input into the quantum graph neural network. The quantum graph neural network performs quantum computation processing on the image patch features in the set of nodes to generate intermediate computation results for participating in the self-attention weight calculation of the visual Transformer model. S3.2: In the process of self-attention weight calculation, the quantum graph neural network is based on the quantum parallel computing process to perform parallel calculation on the similarity between the corresponding features of any two image patches in the node set, and the similarity result obtained by parallel calculation is mapped to the attention weight parameter set. The attention weight parameter set is used to replace the original self-attention weight calculation result of the visual Transformer model. S3.3: Based on the attention weight parameter set, the connection relationship between each image block node and the other image block nodes is updated. The updated connection relationship is used to determine the weight allocation method of image block features in the self-attention weighted calculation process, thereby forming a weighted connection relationship between image blocks for subsequent feature update calculation. S3.4: Based on the weighted connectivity between image patches, the quantum graph neural network performs hypergraph convolution operation on the features of multiple interconnected image patches. The hypergraph convolution operation generates a high-order feature representation that characterizes the spatial correlation between image patches by jointly aggregating the features of multiple image patches under the same weighted connectivity. S3.5: Inject the set of attention weight parameters into the multi-layer self-attention calculation process of the visual Transformer model, and introduce the high-order feature representation into the feature update path of the visual Transformer model to update the image patch feature representation, thus forming the optimized visual Transformer model.

5. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S4 specifically includes: S4.1: Based on the optimized visual Transformer model, the image representation is input to the feature extraction layer of the visual Transformer model. The feature extraction layer processes the input image features layer by layer through a multi-layer neural network structure. S4.2: During the feature extraction process, the visual Transformer model uses a self-attention mechanism to dynamically adjust the attention weights and perform weighted aggregation on each region in the image. S4.3: The features of each image patch are aggregated through a multi-layer neural network structure, and global and local features are combined to generate a feature vector containing global structural information of the image. The feature vector represents the global layout and detailed information in the image. S4.4: While extracting global features, the visual Transformer model processes local defect features separately, learns from local features to extract features of defect regions, and generates feature representations containing local defect information. S4.5: The feature representations of global structural information and local defect features are fused to obtain the final image feature representation, which contains detailed information on the overall structure and local defects in the image.

6. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S5 specifically includes: S5.1: Based on the feature representation of the global structural information and local defect features extracted in step S4, the feature representation is input to the input layer of the convolutional neural network structure, which includes multiple convolutional layers, pooling layers and activation functions; S5.2: In the convolutional layer, a convolutional kernel is used to perform a convolution operation on the input image features. The convolution operation includes filtering the image feature map to obtain local pattern features. S5.3: In the pooling layer, a pooling operation is performed to downsample the convolutional feature map. The pooling operation reduces the spatial dimension of the feature map by selecting the maximum value in a local region and retains the main features in the image. S5.4: In the activation layer, a non-linear activation function, namely the ReLU function, is applied to the pooled feature map to perform a non-linear transformation on the extracted features, thereby obtaining a feature representation with non-linear expressive power; S5.5: The processed features are flattened through a fully connected layer to obtain a one-dimensional feature vector, which contains global structural information and feature information of defect areas in the image; S5.6: The flattened feature vector is input into the classification and regression network, which generates a classification result of the defect based on the type, shape and location of the defect.

7. The automatic PCB defect classification method based on image recognition according to claim 6, characterized in that, The specific classification results of the defects include: The defect type information, defect morphology parameters, and defect spatial location information output by the classification and regression network are combined according to the preset grading rules to form a multi-dimensional grading result representation that includes defect category identifier, defect severity level, and defect spatial distribution description.

8. The automatic PCB defect classification method based on image recognition according to claim 1, characterized in that, S6 specifically includes: S6.1: Based on the defect classification results, extract defect classification and grading information, including the type, shape and location of the defect. The defect classification results are the input information for subsequent model adjustments. S6.2: Based on the defect classification results, the hyperparameters of the visual Transformer model are automatically adjusted; S6.3: After automatically adjusting the hyperparameters, train the visual Transformer model using the adjusted hyperparameters and update the model's weights and biases; S6.4: Through the feedback mechanism, the newly trained visual Transformer model is applied to image data, and the performance of the model is evaluated to obtain the evaluation index results; S6.5: Based on the evaluation results, further adjust the hyperparameters of the visual Transformer model.