Industrial sample defect detection method and device, electronic equipment and storage medium

CN118333946BActive Publication Date: 2026-09-29SOUTH CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410348803.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2026-09-29
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

但由于半导体、金属钢材等表面上的缺陷可能仅占了图像中微小部分,因此,目标检测算法中多次特征提取过程中,容易导致微小目标的特征信息丢失,从而导致漏检,最终造成对微小缺陷的检测效果不佳

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333946B_ABST
    Figure CN118333946B_ABST
Patent Text Reader

Abstract

The application relates to an industrial sample defect detection method, comprising the following steps: respectively adopting a sliding window attention and a large convolution kernel to perform multiple times of feature extraction on a to-be-detected image, obtaining a plurality of stage attention feature maps and large receptive field feature maps, and performing one-to-one fusion calculation on the two to obtain integrated feature maps of stages; performing stage combination and fusion on the integrated feature maps of the stages to obtain a plurality of stage fusion integrated feature maps; and respectively encoding the stage fusion integrated feature maps by adopting an attention encoder to obtain a plurality of stage coding feature embedding blocks; and sequentially performing convolution integration and defect recognition on all the stage coding feature embedding blocks to obtain defects in the to-be-detected image. The industrial sample defect detection method has the global advantages of the attention mechanism and the local advantages of the convolution kernel, so that the effectiveness of micro target feature extraction is greatly improved, and the problem of poor detection effect on micro defects is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target inspection, and in particular to a method, apparatus, electronic device, and storage medium for detecting defects in industrial samples. Background Technology

[0002] Surface defect detection is used to identify visible defects in various industrial products such as semiconductors and metal steel. These defects can seriously affect product quality and safety. Therefore, to improve product quality and safety, surface defect detection technology aims to detect defective products as early as possible during the industrial product manufacturing process. This information guides production line adjustments, reduces the production of defective industrial products, improves product performance stability and reliability, and ultimately saves costs and reduces economic losses.

[0003] To address these issues, the initial approach involved manual inspection to detect surface defects. However, manual inspection has significant drawbacks, including slow speed, high labor costs, high false positive and false negative rates, and, most importantly, inconsistent inspection standards among different personnel. Therefore, manual inspection of product surface defects makes it difficult to guarantee consistent product quality.

[0004] Traditional technologies address these shortcomings through computer vision. Generally, computer vision involves designing different algorithms to extract image features based on specific defects, thereby identifying the defects in an image and freeing up human resources. However, designing different algorithms to extract features for specific defects requires a large amount of prior knowledge and is highly dependent on specialized expertise. Furthermore, in industrial production scenarios, most samples are normal, with only a small number of defective samples.

[0005] To address the above issues, existing technologies draw an analogy between the concept of "defect" and "abnormality," treating defects as anomalies in industrial products. This means that the algorithm only needs to distinguish between normal and abnormal samples, transforming a multi-classification problem into a binary classification problem. Therefore, surface defect detection focuses more on detecting abnormal pixels in images.

[0006] Based on this, existing technologies use deep learning-based target detection algorithms to perform multiple feature extractions on images to locate the positions and sizes of abnormal pixels, thereby achieving defect detection. However, since defects on the surfaces of semiconductors, metal steel, etc., may only occupy a small portion of the image, the multiple feature extraction processes in target detection algorithms can easily lead to the loss of feature information of small targets, resulting in missed detections and ultimately poor detection performance for small defects. Summary of the Invention

[0007] Therefore, the purpose of this invention is to provide a method for detecting defects in industrial samples.

[0008] A method for detecting defects in industrial samples includes the following steps: S1: Multiple feature extractions are performed on the image to be detected using sliding window attention and large convolutional kernels, resulting in several stages of attention feature maps and several stages of large receptive field feature maps. The attention feature maps and large receptive field feature maps from each stage are then combined... Figure 1 A one-to-one fusion calculation is performed to obtain the integrated feature map of each stage; S2: Combine and fuse the integrated feature maps of each stage to obtain several stage-fused integrated feature maps; and use an attention encoder to encode each stage-fused integrated feature map to obtain several stage-coded feature embedding blocks. S3: Perform convolution integration and defect identification on the embedded blocks of encoded features at all stages in sequence to obtain the defects in the image to be detected.

[0009] Compared to existing target detection methods, this invention enhances the feature information of minute targets multiple times during both the feature extraction and feature fusion stages, reducing the loss of feature information for these targets. Furthermore, in feature extraction, this invention integrates the global advantages of attention mechanisms and the local advantages of convolution to address the problem of poor detection performance caused by the neglect of minute defects during defect detection.

[0010] Further, step S1 includes the following steps: S11: A sliding window attention method is used to extract features from the image to be detected multiple times, resulting in attention feature maps at four stages. S12: A large convolutional kernel is used to extract features from the image to be detected multiple times to obtain a large receptive field feature map in four stages; S13: Integrate the attention feature maps and large receptive field features of the four stages. Figure 1 A one-to-one fusion calculation is performed to obtain integrated feature maps for stages one, two, three, and four.

[0011] This invention utilizes sliding window attention for feature extraction, maximizing the preservation of feature information and more effectively capturing global dependencies and finer-grained features. Furthermore, the use of large convolutional kernels efficiently enhances the effective receptive field and facilitates contextual connections between local features, while preventing the loss of crucial feature information during downsampling. By fully combining the advantages of both approaches, the problem of feature information loss in minor defects is solved, resulting in improved performance in detecting these defects.

[0012] Furthermore, the specific calculation process of the stage integration feature map in step S13 is as follows:

[0013] in: for Stage-integrated feature map for Stage attention feature map for Stage attention feature map fusion ratio parameter; for Stage-specific large receptive field feature map for Phase-wise large receptive field feature map fusion ratio parameters.

[0014] This invention employs a dynamic adjustment strategy to make the fusion ratio parameter a trainable parameter. This allows the fusion ratio parameter to find the optimal fusion ratio after training, ensuring that the parameter weights of the large receptive field feature map are increased when local features are needed, and the parameter weights of the attention feature map are increased when global features are needed. Ultimately, this allows the characteristics of attention and convolution to be fully utilized.

[0015] Further, step S2 includes the following steps: S21: Perform convolution processing on the four-stage integrated feature map to obtain the four-stage fused integrated feature map; S22: Perform feature fusion between the three-stage integrated feature map and the four-stage fused integrated feature map to obtain the three-stage fused integrated feature map; S23: Perform feature fusion between the two-stage integrated feature map and the three-stage fusion integrated feature map to obtain the two-stage fusion integrated feature map; S24: Merge the two-stage and three-stage integrated feature maps and perform feature texture transfer to obtain a feature texture fusion feature map; S25: Perform feature fusion between the one-stage integrated feature map and the feature texture fusion feature map to obtain a one-stage fusion integrated feature map; S26: An attention encoder is used to encode the first, second, third and fourth stage fused integrated feature maps respectively to obtain the first, second, third and fourth stage encoded feature embedding blocks; Step S3 includes the following steps: S31: Perform convolutional integration on the coded feature embedding blocks of all stages respectively to obtain the feature maps to be identified in stages one, two, three and four; S32: Perform defect identification and detection on the feature maps to be identified in stages one, two, three and four respectively, and put the detection results of each stage into the map to be detected to obtain the defects in the map to be detected.

[0016] In this invention, feature texture transfer is performed between the two-stage and stage-integrated fusion feature maps to prevent overly extreme cases, ensuring versatility while enhancing feature information for small targets.

[0017] In addition, this invention enhances the ability to fuse features at different levels by encoding the fused integrated feature map with an attention encoder before recognition, thereby achieving higher accuracy and effectiveness in recognizing small targets.

[0018] An industrial sample defect detection device includes a feature extraction unit, a feature fusion unit, and a defect detection unit; The feature extraction unit is used to perform multiple feature extractions on the image to be detected using sliding window attention and large convolutional kernels, respectively, to obtain several stages of attention feature maps and several stages of large receptive field feature maps. The attention feature maps and large receptive field feature maps of each stage are then combined. Figure 1 A one-to-one fusion calculation is performed to obtain the integrated feature map of each stage; The feature fusion unit is used to combine and fuse the integrated feature maps of each stage to obtain several stage-fused integrated feature maps; and an attention encoder is used to encode each stage-fused integrated feature map to obtain several stage-coded feature embedding blocks. The defect detection unit is used to sequentially perform convolution integration and defect identification on all stage-coded feature embedding blocks to obtain the defects in the image to be detected.

[0019] Furthermore, the feature extraction unit includes a sliding window attention feature extraction module, a large convolution kernel feature extraction module, and an integrated feature calculation module; The sliding window attention feature extraction module is used to perform multiple feature extractions on the image to be detected using sliding window attention, thereby obtaining attention feature maps in four stages. The large convolutional kernel feature extraction module is used to perform multiple feature extractions on the image to be detected using a large convolutional kernel to obtain a large receptive field feature map in four stages. The integrated feature calculation module is used to combine the attention feature maps and large receptive field features of the four stages. Figure 1 A one-to-one fusion calculation is performed to obtain integrated feature maps for stages one, two, three, and four.

[0020] Furthermore, the specific calculation process of the stage integrated feature map of the integrated feature calculation module is as follows:

[0021] in: for Stage-integrated feature map for Stage attention feature map for Stage attention feature map fusion ratio parameter; for Stage-specific large receptive field feature map for Phase-wise large receptive field feature map fusion ratio parameters.

[0022] Furthermore, the feature fusion unit includes a fourth-stage feature fusion module, a third-stage feature fusion module, a second-stage feature fusion module, a feature texture fusion module, a first-stage feature fusion module, and a feature encoding module; The fourth-stage feature fusion module is used to perform convolution processing on the four-stage integrated feature map to obtain the four-stage fused integrated feature map. The third-stage feature fusion module is used to fuse the three-stage integrated feature map with the four-stage integrated feature map to obtain the three-stage integrated feature map. The second-stage feature fusion module is used to fuse the two-stage integrated feature map with the three-stage integrated feature map to obtain the two-stage integrated feature map; The feature texture fusion module is used to fuse the two-stage and three-stage integrated feature maps for feature texture transfer to obtain a feature texture fusion feature map. The first-stage feature fusion module is used to fuse the first-stage integrated feature map with the feature texture fusion feature map to obtain a first-stage fused integrated feature map; The feature encoding module is used to encode the first, second, third and fourth stage fused integrated feature maps respectively using an attention encoder to obtain first, second, third and fourth stage encoded feature embedding blocks; The defect detection unit includes a feature convolution integration module and an industrial defect detection module; The feature convolution integration module is used to perform convolution integration on the coded feature embedding blocks of all stages respectively to obtain the feature maps to be identified in stages one, two, three and four. The industrial defect detection module is used to identify and detect defects in the feature maps to be identified in stages one, two, three and four, respectively, and to put the detection results of each stage into the map to be detected to obtain the defects in the map to be detected.

[0023] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the industrial sample defect detection device described in this invention; Figure 2 This is a flowchart of the industrial sample defect detection method described in this invention; Figure 3This is a simplified schematic diagram of the sliding window attention feature extraction process described in this invention; Figure 4 This is a simplified schematic diagram of the large convolutional kernel feature extraction process described in this invention; Figure 5 This is a simplified schematic diagram illustrating the specific calculation process of the integrated feature map described in this invention; Figure 6 This is a simplified schematic diagram illustrating the process of integrating feature maps according to the present invention. Figure 7 This is a simplified schematic diagram of the feature texture fusion process described in this invention; Figure 8 This is a simplified schematic diagram of the encoding process of the attention encoder described in this invention. Detailed Implementation

[0025] To address the issue of feature information loss in small targets during multiple feature extraction processes, leading to missed detections and false detections, this invention further investigates the problem of feature information loss. It discovers that when the defects on the surface of the object being detected are extremely small, the pixel connectivity of the defect in the defect sample image is poor. This results in the easy loss of defect feature information during existing feature extraction techniques, ultimately leading to missed detections, false detections, and inability to identify the target.

[0026] Based on this, the present invention envisions using an attention mechanism to capture global dependencies and finer-grained features. Therefore, a backbone network with a sliding window attention mechanism is selected as the feature extractor, namely the Swing Transformer, which has a low feature information loss rate during downsampling. At the same time, due to the characteristics of the sliding window attention mechanism, the computational load is relatively reduced.

[0027] However, the inventors found that the receptive field of the sliding window attention mechanism grew slowly during application. Therefore, they envisioned using a large convolutional kernel for complementation and selected a backbone network with a similar framework as the second feature extractor. After several ablation experiments, RepLkNet was finally selected as the second feature extractor as the integration object.

[0028] Based on this, the present invention complements the attention mechanism and the large receptive field of large convolutional kernels, especially by integrating the feature maps extracted from both into a unified whole, supplemented by a dynamic adjustment strategy, to maximize the preservation of information from both global and local features, thereby enhancing feature representation and significantly reducing feature information loss. Furthermore, the present invention further fuses the integrated feature map with feature texture to improve the accuracy of detecting minute defects or small targets.

[0029] Based on the above research and design, this invention proposes an industrial sample defect detection method and an industrial sample defect detection device based on the method.

[0030] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram of the industrial sample defect detection device described in this invention. Figure 2 This is a flowchart of the industrial sample defect detection method described in this invention.

[0031] The industrial sample defect detection device includes a feature extraction unit 1, a feature fusion unit 2, and a defect detection unit 3.

[0032] The feature extraction unit 1 is used to perform step S1: Multiple feature extractions are performed on the image to be detected using sliding window attention and large convolutional kernels, respectively, to obtain several stages of attention feature maps and several stages of large receptive field feature maps. The attention feature maps and large receptive field feature maps of each stage are then combined... Figure 1 A one-to-one fusion calculation is performed to obtain the integrated feature map of each stage.

[0033] Specifically, the feature extraction unit 1 includes a sliding window attention feature extraction module 11, a large convolution kernel feature extraction module 12, and an integrated feature calculation module 13.

[0034] Please see Figure 3 , Figure 3 This is a simplified schematic diagram of the sliding window attention feature extraction process described in this invention.

[0035] The sliding window attention feature extraction module 11 is used to perform step S11: using sliding window attention to extract features from the image to be detected multiple times to obtain attention feature maps in four stages.

[0036] Specifically, let the dimension of the image to be detected be:

[0037] in, For the height of the image to be detected, The width of the image to be detected. The number of channels in the image to be tested.

[0038] Next, the image to be detected is divided into... The size is The tiles are then reshaped to obtain a dimension of [dimensional value missing]. tiles That is, pressing on the width and height of the image to be detected respectively. The image is divided into sections, with each section having dimensions of [dimensions missing]. .

[0039] After obtaining the image patch, the image patch is linearly transformed and the number of channels is adjusted through a linear embedding layer to obtain an embedding block.

[0040] The embedded block is processed by the first-stage sliding window attention module (Swin Transformer Block) to obtain a first-stage attention feature map. The sliding window attention method restricts attention calculation to a fixed-size window for feature extraction of the embedded block.

[0041] Next, the resolution of the first-stage attention feature map is reduced by a downsampling module (Patch Merging) to obtain the processed first-stage attention feature map.

[0042] The downsampling module (Patch Merging) selects elements from the input feature map at positional intervals of 2, concatenates them into new feature map patches, and then expands all the feature map patches into a tensor to obtain the processed feature map. Based on this, the downsampling module (Patch Merging) achieves the effect of relatively reducing the loss of feature information and focusing on key feature information.

[0043] The first-stage attention feature map is processed through the sliding window attention module to obtain the second-stage attention feature map.

[0044] After obtaining the two-stage attention feature map, the resolution of the two-stage attention feature map is reduced by the downsampling module and then processed by the sliding window attention module to obtain the three-stage attention feature map.

[0045] After obtaining the three-stage attention feature map, the resolution of the three-stage attention feature map is reduced by the downsampling module and then processed by the sliding window attention module to obtain the four-stage attention feature map.

[0046] The sliding window attention module and downsampling module mentioned above are both components of the existing SwinTransformer, and their specific details will not be elaborated or limited here.

[0047] Feature extraction using sliding window attention maximizes the preservation of feature information. Furthermore, the sliding window attention mechanism, by sliding across the feature map, enables interaction between adjacent windows, thereby establishing contextual connections and increasing the continuity between pixels in the feature map. This ultimately approaches a global modeling capability, preserving global feature information to the greatest extent possible. In addition, by confining attention computation to a fixed-size window, sliding window attention reduces the massive computational cost (FLOPs) required by traditional attention methods, thus lowering model complexity.

[0048] Please see Figure 4 , Figure 4 This is a simplified schematic diagram of the large convolution kernel feature extraction process described in this invention.

[0049] The large convolutional kernel feature extraction module 12 is used to perform step 12: using a large convolutional kernel to extract features from the image to be detected multiple times to obtain a large receptive field feature map in four stages.

[0050] Specifically, the image to be detected is processed by convolution and depthwise convolution to obtain a preprocessed feature map, and the preprocessed feature map is then used to extract features through a large convolution kernel to obtain a first-stage large receptive field feature map.

[0051] The first-stage large receptive field feature map is downsampled sequentially through convolution and depthwise convolution to obtain the processed first-stage large receptive field feature map; features are extracted from the processed first-stage large receptive field feature map using a large convolution kernel to obtain the second-stage large receptive field feature map.

[0052] The two-stage large receptive field feature map is downsampled sequentially through convolution and depthwise convolution to obtain the processed two-stage large receptive field feature map; features are extracted from the processed two-stage large receptive field feature map using a large convolution kernel to obtain the three-stage large receptive field feature map.

[0053] The three-stage large receptive field feature maps are downsampled sequentially through convolution and depthwise convolution to obtain the processed three-stage large receptive field feature maps; features are extracted from the processed three-stage large receptive field feature maps using large convolution kernels to obtain the four-stage large receptive field feature maps.

[0054] In particular, the use of large convolutional kernels can more efficiently improve the effective receptive field and help the contextual connection of local feature information. At the same time, it can ensure that the loss of feature information is greatly reduced during downsampling.

[0055] Since large convolutional kernels tend to increase computational cost and parameter count, this invention further employs small convolutional kernels for structural reparameterization to compress these costs. Specifically, this invention replaces the multi-headed self-attention mechanism in the sliding window attention mechanism with a large convolutional kernel (31x31 Depth-Wise CNN). For detailed information, please refer to the existing technology RepLKNet (A Large-Kernel Architecture). This invention will not elaborate further or limit its scope here.

[0056] Increasing the receptive field of the feature map can effectively preserve the feature information of small targets, ensuring the connectivity of their feature information among pixels in the feature map. At the same time, a large receptive field allows the network to reach the predetermined effective receptive domain with fewer layers, thus avoiding the optimization problems associated with increasing depth.

[0057] Please see Figure 5 , Figure 5 This is a simplified schematic diagram illustrating the specific calculation process of the integrated feature map described in this invention.

[0058] The integrated feature calculation module 13 is used to perform step S13: combining the attention feature maps and large receptive field features of the four stages. Figure 1 A one-to-one fusion calculation is performed to obtain integrated feature maps for stages one, two, three, and four.

[0059] Specifically, for The specific fusion calculation process of the staged integrated feature map is as follows:

[0060] in: for Stage-integrated feature map for Stage attention feature map for Stage attention feature map fusion ratio parameter; for Stage-specific large receptive field feature map for Phase-wise large receptive field feature map fusion ratio parameters.

[0061] For example, the specific fusion calculation process for a one-stage integrated feature map is as follows:

[0062] in, For a one-stage integrated feature map, This is a one-stage attention feature map. The fusion ratio parameter for the first-stage attention feature map; This is a feature map of a large receptive field in one stage. The fusion ratio parameter for a large receptive field feature map in one stage; The specific fusion calculation process for the two-stage integrated feature map is as follows:

[0063] in, For two-stage integrated feature maps, This is a two-stage attention feature map. The fusion ratio parameter for the two-stage attention feature map; This is a feature map of the large receptive field in the second stage. The fusion ratio parameter for the two-stage large receptive field feature map; The specific fusion calculation process for the three-stage integrated feature map is as follows:

[0064] in, For three-stage integrated feature maps, This is a three-stage attention feature map. The fusion ratio parameter for the three-stage attention feature maps; This is a feature map of the three-stage large receptive field. The fusion ratio parameter for the three-stage large receptive field feature maps; The specific fusion calculation process for the four-stage integrated feature map is as follows:

[0065] in, For four-stage integrated feature maps, This is a four-stage attention feature map. The fusion ratio parameter for the four-stage attention feature map; This is a feature map of the four-stage large receptive field. The fusion ratio parameters for the four-stage large receptive field feature maps.

[0066] Since the feature information of tiny targets accounts for a very small proportion, most of the background will be regarded as noise interference. Therefore, this invention uses additive fusion on the feature map while adding a fusion ratio parameter. Additive fusion can better reduce the noise interference generated by the background and enhance the tiny feature information to a certain extent, while having relatively less computation and fewer parameters.

[0067] In addition, this invention sets fusion ratio parameters for the attention feature map and the large receptive field feature map corresponding to each stage. This allows the parameters to amplify the feature extraction characteristics for small targets during the learning process. Since the attention feature map focuses more on the global picture, while the large receptive field feature map focuses more on the local picture, if the parameters are optimized through learning, different fusion ratio parameters can be obtained for different stages. For example, when the feature map of stage one has more positional feature information and a larger resolution, it is necessary to focus more on local feature extraction for small target features, thus requiring an increase in the parameter weight of the large receptive field feature map. Conversely, the feature map of stage four has more semantic information and a smaller resolution, thus requiring a greater focus on global feature extraction, thus requiring an increase in the parameter weight of the attention feature map. Therefore, this invention significantly improves the detection and recognition accuracy of small targets by continuously learning the fusion ratio parameters to find the relatively optimal fusion ratio between feature maps.

[0068] The parameter of the fusion ratio has a certain upper limit threshold. When the upper limit threshold is reached, the weight is normalized. However, the upper limit threshold can be arbitrarily limited based on the user's needs. This invention does not specifically limit it here.

[0069] Please see Figure 6 , Figure 6 This is a simplified schematic diagram of the process of integrating feature maps according to the present invention.

[0070] The feature fusion unit is used to perform step S2: to combine and fuse the integrated feature maps of each stage to obtain several stage fused integrated feature maps; and to use an attention encoder to encode each stage fused integrated feature map to obtain several stage encoded feature embedding blocks.

[0071] Specifically, the feature fusion unit includes a fourth-stage feature fusion module 21, a third-stage feature fusion module 22, a second-stage feature fusion module 23, a feature texture fusion module 24, a first-stage feature fusion module 25, and a feature encoding module 26.

[0072] The fourth-stage feature fusion module 21 is used to perform step S21: convolution processing on the four-stage integrated feature map to obtain the four-stage fused integrated feature map.

[0073] Specifically, the four-stage fused integrated feature map is obtained by adjusting the number of channels to a preset number of fusion channels through convolution. .

[0074] The number of fusion channels refers to the number of channels used uniformly during the fusion feature process, and this invention does not impose a specific limitation on it.

[0075] The third-stage feature fusion module 22 is used to perform step S22: to fuse the three-stage integrated feature map with the four-stage fused integrated feature map to obtain the three-stage fused integrated feature map.

[0076] Specifically, the three-stage integrated feature map is adjusted to the preset fusion channel number through convolution to obtain the processed three-stage integrated feature map. Simultaneously, the four-stage fusion integrated feature map is... Upsampling is performed to enlarge the resolution size to match the integrated features of the three-stage post-processing. Figure 1 The resulting four-stage fusion and integration feature map is obtained.

[0077] The processed three-stage integrated feature map is fused with the processed four-stage fused integrated feature map to obtain the three-stage fused integrated feature map. .

[0078] The above operations are conventional feature fusion processes, and the fusion details and limitations will not be elaborated here.

[0079] The second-stage feature fusion module 23 is used to perform step S23: to fuse the two-stage integrated feature map with the three-stage fusion integrated feature map to obtain the two-stage fusion integrated feature map.

[0080] Specifically, the number of channels in the two-stage integrated feature map is adjusted to the preset number of fusion channels through convolution to obtain the processed two-stage integrated feature map. Simultaneously, the three-stage fusion integrated feature map is... Upsampling is performed to enlarge the resolution size to match the resolution size of the processed two-stage integrated feature map, thereby obtaining the processed three-stage fused integrated feature map.

[0081] The processed two-stage integrated feature map is fused with the processed three-stage fused integrated feature map to obtain a two-stage fused integrated feature map. .

[0082] The above operations are conventional feature fusion processes, and the fusion details and limitations will not be elaborated here.

[0083] Please see Figure 7 , Figure 7 This is a simplified schematic diagram of the feature texture fusion process described in this invention.

[0084] The feature texture fusion module 24 is used to perform step S24: to merge the two-stage and three-stage integrated feature map and perform feature texture transfer to obtain a feature texture fusion feature map.

[0085] Specifically, the three-stage fusion and integration feature map Content features are extracted using a content extractor, and then sub-pixel convolution is used to increase the resolution of these content features to match the two-stage fusion feature map. With the same resolution and size, the expanded content features are obtained.

[0086] The two-stage fusion integrated feature map is overlapped with the expanded content features. The overlapped features are then processed by a texture extractor to select reliable regions from the two overlapping features to obtain texture features.

[0087] The texture features and the expanded content features are fused using a residual connection to obtain a feature-texture fusion feature map. The formula is expressed as follows:

[0088] in, For texture extractor, For content extractor, To increase the size of the current resolution by a factor of two, Representation of features With features Interconnected, This is a two-stage fusion and integration feature map. This is a three-stage fusion and integration feature map. Feature maps are fused to feature textures.

[0089] This invention performs feature texture transfer between the two-stage and three-stage integrated feature maps, that is, it selects the intermediate level feature map for processing to prevent overly extreme cases. It comprehensively considers the strong semantic information in the low-resolution feature map and the key local details in the high-resolution feature map, ensuring versatility. At the same time, it accurately extracts the feature information of small defects or small targets, which greatly improves the effect of identifying small defects or small targets.

[0090] The first-stage feature fusion module 25 is used to perform step S25: to fuse the first-stage integrated feature map with the feature texture fusion feature map to obtain a first-stage fused integrated feature map.

[0091] Specifically, the first-stage integrated feature map is convolved with a number of channels up to the preset number of fusion channels to obtain the processed first-stage integrated feature map. Simultaneously, the feature texture fusion feature map is... Upsampling is performed to enlarge the resolution size to match the resolution size of the processed two-stage integrated feature map, thereby obtaining the processed feature texture fusion feature map.

[0092] The processed first-stage integrated feature map is fused with the processed feature texture fusion feature map to obtain a first-stage fused integrated feature map. .

[0093] Among them, since the first-stage fusion integrated feature map has the characteristics of feature texture fusion feature map, and also has more position feature information and high resolution, the defect feature information will be further enhanced by position feature information and the continuity of pixels (feature points) in the feature map in the first-stage fusion integrated feature map.

[0094] Please see Figure 8 , Figure 8 This is a simplified schematic diagram of the encoding process of the attention encoder described in this invention.

[0095] The feature encoding module 26 is used to perform step S26: using an attention encoder to encode the first, second, third and fourth stage fused integrated feature maps respectively to obtain the first, second, third and fourth stage encoded feature embedding blocks.

[0096] Specifically, firstly, all fused integrated feature maps are divided into blocks to obtain patches of the first, second, third, and fourth stage fused integrated feature maps, and then all patches are mapped to obtain embedded patches of the first, second, third, and fourth stages.

[0097] Next, all embedded blocks are transformed into processed embedded blocks of sequence structure through layer normalization.

[0098] Then, attention calculations were performed on all processed embedding blocks using a multi-head attention mechanism to obtain attention feature embedding blocks for stages one, two, three, and four, respectively.

[0099] Residual connections are made between all attention feature embedding blocks and their corresponding embedding blocks to obtain attention feature embedding blocks after the first, second, third, and fourth stages of processing, respectively.

[0100] Then, all the processed attention feature embedding blocks are processed by layer normalization and output to a multilayer perceptron for further processing. The processed feature embedding blocks are then residually connected with all the attention feature embedding blocks to obtain the first, second, third and fourth stage encoded feature embedding blocks, respectively.

[0101] By associatively connecting all feature embedding blocks through attention encoding, the continuity of the context is improved, focusing important key information on each feature element, thereby better understanding the overall semantics and contextual relationships of all embedding blocks. Furthermore, due to sliding window attention and a large receptive field, feature information of minor defects is preserved. Therefore, attention encoding further ensures that the feature information of minor defects or small targets is more significant, thus improving the effectiveness and accuracy of identifying minor defects or small targets.

[0102] The attention encoding process uses the existing Transformer Encoder structure, which is not specifically limited here.

[0103] The defect detection unit 3 is used to perform step S3: sequentially convolutional integration and defect identification on all stage encoded feature embedding blocks to obtain the defects in the image to be detected.

[0104] Specifically, the defect detection unit 3 includes a feature convolution integration module 31 and an industrial defect detection module 32.

[0105] The feature convolution integration module 31 is used to perform step S31: convolution integration of all stage encoded feature embedding blocks to obtain the first, second, third and fourth stage feature maps to be identified.

[0106] Convolution can integrate embedded blocks into feature maps, ensuring that the feature maps can be directly used for subsequent detection.

[0107] The industrial defect detection module 32 is used to perform step S32: perform defect identification and detection on the feature maps to be identified in stages one, two, three and four respectively, and put the detection results of each stage into the map to be detected to obtain the defects in the map to be detected.

[0108] Specifically, the corresponding features in the feature maps to be identified in stages one, two, three and four are identified and detected respectively, and the detection results of the feature maps in stages one, two, three and four are obtained respectively. All detection results are placed into the map to be detected, and the defects in the map to be detected are significantly represented, thereby obtaining the defects in the map to be detected.

[0109] The above describes the general functions of a probe, namely regression and classification. This invention will not elaborate on or limit these functions here.

[0110] To ensure the accuracy and effectiveness of the industrial sample defect detection method described in this invention, the trainable parameters are further trained using a training dataset.

[0111] Due to the diversity of defect types, and the influence of various factors such as different industrial product manufacturing equipment and environment, the manifestation of defects varies. Furthermore, users have different defect error thresholds. Therefore, the training data can be collected and labeled independently, thereby improving the relevance and effectiveness of the industrial sample defect detection method described in this invention. This invention will not be specifically elaborated or limited herein.

[0112] Among them, the scarcity of defects in industrial samples can easily lead to imbalance in training data, i.e., normal samples, which can result in overfitting, i.e., insufficient generalization ability, and easy to miss detections.

[0113] In this embodiment, the present invention performs data augmentation on the training data, namely random flipping, Mosaic data augmentation, Mixup data augmentation, and copypaste data augmentation, to obtain an enhanced training dataset.

[0114] The above data augmentation methods are all existing technologies. Furthermore, the required augmentation methods vary depending on the application scenario. Therefore, this invention does not specifically limit the data augmentation methods.

[0115] After obtaining the enhanced training dataset, the detection results are obtained by inputting the training data in the enhanced training dataset into the industrial sample detection device.

[0116] The detection results are compared with the corresponding standard results in the training dataset using a loss function to calculate the loss value.

[0117] Based on the feedback from the loss value, the weights of the trainable parameters are updated.

[0118] Repeat the above steps of inputting training data, calculating loss, and updating weights until the trainable parameters reach a relatively optimal level.

[0119] In this embodiment, the specific calculation formula for the loss is as follows:

[0120] in, The EIOU loss function is used. This is the ratio of the intersection and union of the predicted bounding box and the ground truth bounding box. These represent the center points of the predicted bounding box and the ground truth bounding box, respectively. To calculate the Euclidean distance between two center points, It is the diagonal distance of the smallest closure region that simultaneously contains both the predicted and ground truth boxes. The widths of the predicted bounding box and the ground truth bounding box. The actual height of the bounding box. and These represent the minimum box width and height that cover the ground truth box and the predicted box, respectively.

[0121] Therefore, because the causes of defects vary, their shapes can be irregular, easily affecting the width and height of the predicted bounding box. Thus, the EIOU loss function additionally calculates the height and width of the detection box separately to ensure that irregularities in the defect shape are accounted for.

[0122] Compared to existing technologies, this invention improves upon the YOLOv5 architecture by employing a sliding window attention mechanism and large convolutional kernels for feature extraction. The feature maps from both methods are then integrated, and the integrated feature maps are fused and subjected to feature texture transfer. Finally, all fused feature maps are attention-encoded for defect detection. This invention enhances the feature information of small targets multiple times in both the feature extraction and feature fusion stages to prevent the loss of such information. Furthermore, in feature extraction, this invention combines the global advantages of the attention mechanism with the local advantages of the convolutional kernel, significantly improving the effectiveness of small target feature extraction and solving the problem of poor detection performance for small defects.

[0123] Based on the same inventive concept, this application also provides an electronic device, which can be a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.) or other terminal device. The device includes one or more processors and a memory, wherein the processor is used to execute a program to implement the industrial sample defect detection method of the embodiments of the present invention; the memory is used to store a computer program executable by the processor.

[0124] Based on the same inventive concept, this application also provides a computer-readable storage medium corresponding to the aforementioned embodiment of an industrial sample defect detection method. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the industrial sample defect detection method described in any of the above embodiments.

[0125] This application may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0126] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and the present invention also intends to include these modifications and variations.

Claims

1. A method for detecting defects in industrial samples, characterized in that, Includes the following steps: S1: Multiple feature extractions are performed on the image to be detected using sliding window attention and large convolution kernel respectively, resulting in several stages of attention feature maps and several stages of large receptive field feature maps. The attention feature maps and large receptive field feature maps of each stage are fused one by one to obtain the integrated feature map of each stage. S2: Combine and fuse the integrated feature maps of each stage to obtain several stage-fused integrated feature maps; and use an attention encoder to encode each stage-fused integrated feature map to obtain several stage-coded feature embedding blocks. S3: Perform convolution integration and defect identification on all stage encoded feature embedding blocks in sequence to obtain the defects in the image to be detected; Step S1 includes the following steps: S11: A sliding window attention method is used to extract features from the image to be detected multiple times, resulting in attention feature maps at four stages. S12: A large convolutional kernel is used to extract features from the image to be detected multiple times to obtain a large receptive field feature map in four stages; S13: The attention feature maps and large receptive field feature maps of the four stages are fused and calculated one by one to obtain the integrated feature maps of the first, second, third and fourth stages. Step S2 includes the following steps: S21: Perform convolution processing on the four-stage integrated feature map to obtain the four-stage fused integrated feature map; S22: Perform feature fusion between the three-stage integrated feature map and the four-stage fused integrated feature map to obtain the three-stage fused integrated feature map; S23: Perform feature fusion between the two-stage integrated feature map and the three-stage fusion integrated feature map to obtain the two-stage fusion integrated feature map; S24: Merge the two-stage and three-stage integrated feature maps and perform feature texture transfer to obtain a feature texture fusion feature map; S25: Perform feature fusion between the one-stage integrated feature map and the feature texture fusion feature map to obtain a one-stage fusion integrated feature map; S26: An attention encoder is used to encode the first, second, third and fourth stage fused integrated feature maps respectively to obtain the first, second, third and fourth stage encoded feature embedding blocks; Step S3 includes the following steps: S31: Perform convolutional integration on the coded feature embedding blocks of all stages respectively to obtain the feature maps to be identified in stages one, two, three and four; S32: Perform defect identification and detection on the feature maps to be identified in stages one, two, three and four respectively, and put the detection results of each stage into the map to be detected to obtain the defects in the map to be detected.

2. The industrial sample defect detection method according to claim 1, characterized in that, The specific calculation process of the stage integration feature map in step S13 is as follows: in: for Stage-integrated feature map for Stage attention feature map for Stage attention feature map fusion ratio parameter; for Stage-specific large receptive field feature map for Phase-wise large receptive field feature map fusion ratio parameters.

3. An industrial sample defect detection device, characterized in that, It includes a feature extraction unit, a feature fusion unit, and a defect detection unit; The feature extraction unit is used to perform multiple feature extractions on the image to be detected using sliding window attention and large convolution kernels respectively, to obtain several stages of attention feature maps and several stages of large receptive field feature maps. The attention feature maps and large receptive field feature maps of each stage are fused and calculated one-to-one to obtain the integrated feature map of each stage. The feature fusion unit is used to combine and fuse the integrated feature maps of each stage to obtain several stage-fused integrated feature maps; and an attention encoder is used to encode each stage-fused integrated feature map to obtain several stage-coded feature embedding blocks. The defect detection unit is used to sequentially perform convolution integration and defect identification on all stage encoded feature embedding blocks to obtain defects in the image to be detected. in, The feature extraction unit includes a sliding window attention feature extraction module, a large convolutional kernel feature extraction module, and an integrated feature calculation module. The sliding window attention feature extraction module is used to perform multiple feature extractions on the image to be detected using sliding window attention to obtain four stages of attention feature maps. The large convolutional kernel feature extraction module is used to perform multiple feature extractions on the image to be detected using large convolutional kernels to obtain four stages of large receptive field feature maps. The integrated feature calculation module is used to fuse the four stages of attention feature maps and large receptive field feature maps one-to-one to obtain integrated feature maps of stages one, two, three, and four. The feature fusion unit includes a fourth-stage feature fusion module, a third-stage feature fusion module, a second-stage feature fusion module, a feature texture fusion module, a first-stage feature fusion module, and a feature encoding module. The fourth-stage feature fusion module performs convolution processing on the four-stage integrated feature map to obtain a four-stage fused integrated feature map. The third-stage feature fusion module fuses the three-stage integrated feature map with the four-stage fused integrated feature map to obtain a three-stage fused integrated feature map. The second-stage feature fusion module fuses the two-stage integrated feature map with the three-stage fused integrated feature map to obtain a two-stage fused integrated feature map. The feature texture fusion module performs feature texture transfer on the two-stage and three-stage fused integrated feature maps to obtain a feature texture fused feature map. The first-stage feature fusion module fuses the one-stage integrated feature map with the feature texture fused feature map to obtain a one-stage fused integrated feature map. The feature encoding module uses an attention encoder to encode the one-stage, two-stage, three-stage, and four-stage fused integrated feature maps respectively to obtain one-stage, two-stage, three-stage, and four-stage encoded feature embedding blocks. The defect detection unit includes a feature convolution integration module and an industrial defect detection module. The feature convolution integration module is used to perform convolution integration on the coded feature embedding blocks of all stages respectively to obtain the feature maps to be identified in stages one, two, three and four. The industrial defect detection module is used to identify and detect defects in the feature maps to be identified in stages one, two, three and four respectively, and put the detection results of each stage into the image to be detected to obtain the defects in the image to be detected.

4. The industrial sample defect detection device according to claim 3, characterized in that, The specific calculation process of the stage integrated feature map of the integrated feature calculation module is as follows: in: for Stage-integrated feature map for Stage attention feature map for Stage attention feature map fusion ratio parameter; for Stage-specific large receptive field feature map for Phase-wise large receptive field feature map fusion ratio parameters.

5. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements an industrial sample defect detection method as described in any one of claims 1 to 2.

6. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are used in an industrial sample defect detection method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • PCB defect image detection method based on improved deep learning algorithm

    CN115409797A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1