Two-stage metal surface defect detection method based on computing power perception

By combining the improved YOLOv8 and U-Net models, an adaptive confidence threshold mechanism, and a channel-spatial attention module, the shortcomings of the metal product surface defect detection system in terms of detection accuracy, real-timeness, and user-friendliness are addressed, achieving efficient and reliable detection results.

CN120807888APending Publication Date: 2025-10-17NANYANG NORMAL UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510931551.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing metal product surface defect detection systems have shortcomings in detection accuracy, real-time performance, generalization capability, system scalability and user-friendliness, making them difficult to be used stably and efficiently in complex industrial environments.

Method used

By fusing an improved YOLOv8 model with an optimized U-Net model, combined with an adaptive confidence threshold mechanism and a channel-spatial attention module, two-stage detection is implemented, dynamically balancing detection accuracy and computing resources to improve the accuracy of small target defect detection.

Benefits of technology

It significantly improves the accuracy and robustness of small target defect detection, optimizes computing resource utilization, adapts to detection needs in complex environments, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807888A_ABST
    Figure CN120807888A_ABST
Patent Text Reader

Abstract

The invention provides a two-stage metal surface defect detection method based on computing power perception, which is characterized in that improved U-Net and YOLOv8 are cascaded, the method is mainly used for metal surface defect detection, and aims to solve the problems of low precision, insufficient computing resource utilization rate, high detection efficiency and the like in small target defect detection in the prior art. And the detection efficiency and the accuracy are difficult to balance in a complex environment. The operation process of the method comprises the following steps: firstly, carrying out rough detection on a target by utilizing a YOLOv8 model; and then, according to the confidence of the coarse detection result and the current computing power of the equipment, designing a self-adaptive high-low confidence threshold mechanism, and grading the coarse detection result according to the mechanism. Experimental results show that the method remarkably improves the detection precision of the small target defect sample, optimizes the allocation of computing resources, effectively balances the detection efficiency and accuracy, and can be well suitable for a complex industrial detection environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of target detection and defect detection, and particularly relates to a two-stage metal surface defect detection method based on computing power perception. BACKGROUND

[0002] In the process of industrial automation production, product defect detection is a core link to ensure product quality, and its importance is self-evident. The traditional manual detection method has obvious disadvantages, not only the detection efficiency is low and the labor cost is high, but also the detection result is easily affected by the subjective factors of the detection personnel, which is difficult to meet the actual needs of product quality detection in modern large-scale production.

[0003] Although the existing industrial visual detection system has realized automation to some extent, there are still many problems to be solved. On the one hand, it is difficult to effectively balance the detection accuracy and real-time performance. Single-stage detection algorithm has high real-time performance, but the detection accuracy is often difficult to meet the production requirements; multi-stage algorithm has relatively high accuracy, but consumes a large amount of computing resources, resulting in slow detection speed. On the other hand, the system has poor adaptability to complex industrial scenes. When facing different light conditions, various product types and changing defect forms, it is difficult to play a stable and accurate role. In addition, model updating and training are highly dependent on a large amount of labeled data, while in the industrial field, data acquisition costs are high, the labeling process is tedious and inefficient, which greatly limits the updating and optimization of the model. At the same time, the system architecture has high coupling degree, and the hardware compatibility and algorithm scalability are insufficient, which makes it difficult to quickly adapt to new detection requirements and meet the requirements of the rapid development of industrial production.

[0004] In the production process of metal products, surface defect detection is also a key link to ensure product quality. The traditional manual detection method also has problems such as low efficiency, strong subjectivity, high omission rate, etc. With the rapid development of computer vision and deep learning technology, defect detection methods based on deep learning have gradually become the mainstream. However, the existing detection system has many deficiencies in user interaction, such as inconvenience of user uploading image or video files, non-intuitive and clear display of detection results, complex and tedious model parameter adjustment process, etc. These problems seriously affect the practicality and user experience of the detection system, limiting the application effect of the detection method based on deep learning in actual production.

[0005] In summary, the existing technology has certain deficiencies in the field of metal product surface defect detection in terms of detection accuracy, real-time performance, generalization ability, system scalability and user friendliness, which makes it difficult to be applied stably and efficiently in complex industrial production environment. Therefore, designing an efficient and easy-to-use user interface has important practical significance and urgent needs for improving the application value of the metal product surface defect detection system.

[0006] To solve the above problems, the application provides a self-adaptive cascade detection method and system based on computing power perception, which fuses an improved YOLOv8 model and an optimized U-Net model and applies them to metal surface defect detection. In the coarse detection stage, the YOLOv8 model is used to extract the multi-scale features of the image and generate initial defect candidate boxes, and then the detection results are divided into high, medium and low three grades through adaptive high-low confidence thresholds (τ, τ'). The computing power perception dynamic adjustment mechanism can dynamically adjust the thresholds τ and τ' according to the real-time computing power of the device. When the device computing power is high, the medium confidence interval is expanded to improve the detection accuracy; when the device computing power is low, the interval is reduced to ensure the real-time performance of the detection, so as to realize the dynamic balance between the computing resources and the detection performance. In the fine detection stage, the ROI of the medium confidence sample is cropped and input into the improved U-Net network. The network effectively enhances the expression of small target features with the help of the channel-space attention module, and optimizes the positioning of the defect boundary by combining the Bayesian fusion technology. Finally, the results of coarse detection and fine detection are fused, and the final defect classification and pixel-level segmentation results are output through coordinate correction and confidence weighting. The application significantly improves the detection accuracy of small target defects and provides an efficient and reliable solution for metal product surface defect detection. SUMMARY

[0007] To solve the core challenges of large detection difficulty, limited computing resources and difficult balance between detection accuracy and efficiency in the process of small target defect detection on metal surface, the application provides a two-stage metal surface defect detection method based on computing power perception.

[0008] The first stage is coarse detection. The YOLOv8 model is used for fast coarse detection of the input image to generate multi-scale feature maps and candidate detection boxes, and the detection results are divided into high, medium and low confidence levels through an adaptive high-low confidence threshold mechanism. Among them, the high confidence and low confidence targets are directly output, and the medium confidence target is introduced into the second stage for further fine segmentation.

[0009] The second stage is fine segmentation. The ROI of the medium confidence target region is cropped and input into the improved U-Net model integrated with the channel-space attention module for semantic segmentation. The module effectively reduces background noise interference and enhances target feature expression through the double attention mechanism (channel attention and spatial attention), thereby significantly improving the detection accuracy of small target defects in complex background. Finally, the confidence of the detection box is updated through the Bayesian fusion strategy to ensure the boundary accuracy and reliability of the detection results.

[0010] Adaptive confidence threshold mechanism. An adaptive confidence threshold mechanism is introduced, which can dynamically adjust the threshold according to the real-time computing capacity of the detection device. When the device computing capacity is sufficient, the medium confidence interval is expanded, allowing more samples to enter the fine segmentation stage to improve detection accuracy; when the device computing capacity is limited, the medium confidence interval is narrowed to ensure real-time processing performance.

[0011] In this way, the present application can achieve dynamic balance between detection accuracy and efficiency in complex environments, optimizing the use of computing resources. Experimental results show that this method exhibits significant effectiveness and superiority in small target defect detection on metal surfaces.

[0012] Advantages

[0013] The multi-model two-stage cascade detection method proposed by the present application combines YOLOv8 model and improved U-Net organically, fully utilizing the high efficiency of YOLOv8 in target detection and the fineness of U-Net in image segmentation. This combination significantly improves the detection accuracy of small target defects, especially in the presence of occlusion and blur. At the same time, by introducing an adaptive confidence threshold mechanism, the present application can dynamically adjust the detection strategy according to the real-time computing capacity of the detection device, ensuring efficient balance between detection accuracy and processing speed under limited computing resources, and achieving optimal utilization of computing resources. In addition, by integrating channel-spatial attention modules in U-Net and adopting Bayesian fusion strategy, the present application effectively enhances the robustness of the model in complex background and noise interference, significantly improving the accuracy and reliability of the detection results. Therefore, the present application provides an efficient, accurate and adaptable solution for metal surface defect detection in industrial scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 : Based on adaptive double confidence threshold, the cascade decision logic framework diagram.

[0015] Figure 2 : mAP50 curve diagram of each algorithm.

[0016] Figure 3 : YOLOv8 algorithm framework diagram.

[0017] Figure 4 : U-Net algorithm architecture diagram.

[0018] Figure 5 : Channel-spatial attention diagram

[0019] Figure 6 : Two-stage detection framework diagram of YOLOv8 cascade improved U-Net (I).

[0020] Figure 7 : Two-stage detection flowchart of YOLOv8 cascade improved U-Net (II).

[0021] Figure 8 : Chart 1 dataset statistics and enhancement strategy.

[0022] Figure 9 : Chart 2 comparison of different object detection algorithms. DETAILED DESCRIPTION

[0023] The application will be further described below with reference to the accompanying drawings. As shown in Figures 1 to 9 .

[0024] First stage: rough detection (YOLOv8) (as shown in Figure 1 ).

[0025] 1. Input image preprocessing: Normalize the original image and adjust the resolution of the pixels. At the same time, in order to enhance the robustness of the model, perform various data enhancement operations on the image to simulate various interference situations in the industrial environment.

[0026] 2. Backbone network feature extraction: Use the backbone network of YOLOv8 to extract features from the preprocessed image. The backbone network contains C2f modules and Spatial Pyramid Pooling Fast (SPPF) modules. C2f modules have more branches and can better perform gradient backpropagation; SPPF modules are located at the end of the backbone network and can convert objects of different sizes into fixed-size feature representations, improving the adaptability of the model. Through these modules, multi-scale feature maps (P3, P4, P5) are generated, which correspond to different resolution detection heads.

[0027] 3. Neck network feature fusion: Use the neck structure composed of Feature Pyramid Network (FPN) and Path Aggregation Network (PAN) to fuse the features extracted by the backbone network. FPN performs top-down feature transmission, PAN performs bottom-up feature transmission, and the combination of the two realizes effective fusion of features of different scales, providing more rich feature information for subsequent detection.

[0028] 4. Detection head generates candidate boxes: The detection head uses a "decoupling" structure to independently extract class and position features. Based on the features fused by the neck network, generate candidate detection boxes, each candidate box contains the position information and confidence score of the target.

[0029] 5. Confidence classification and preliminary processing:

[0030] 1. Dynamic threshold calculation

[0031] The dynamic adjustment formula of the high confidence threshold τ and the low confidence threshold τ' is:

[0032]

[0033] Wherein:

[0034] (1)T low , T high are the initial low and high confidence thresholds, respectively, determined by the high / low confidence cluster centers μ high , μ low and the offset δ obtained by K-means++ clustering, calculated as follows:

[0035] T high = μ high - δ,

[0036] T low = μ low - δ.

[0037] (2) t is the real-time delay of the device, t0=100ms is the reference delay, Δt is the delay fluctuation coefficient, and k=0.1 is the smoothing coefficient.

[0038] 2. Confidence classification logic

[0039] The classification condition of the candidate box confidence S is:

[0040] High confidence target: S>τ, directly output;

[0041] Low confidence target: S<τ', directly output;

[0042] Medium confidence target: τ'≤S≤τ, passed to the second stage.

[0043] 3. Threshold adjustment mechanism with computing power awareness

[0044] (1) When the device computing power is sufficient t<t0, tanh(t-t0) / △t<0,

[0045] At this time: τ increases, τ' decreases The medium confidence interval [τ',τ] expands.

[0046] (2) When the device computing power is insufficient t≥t0, tanh(t-t0) / △t≥0,

[0047] At this time: τ increases, τ' decreases The medium confidence interval [τ',τ] shrinks.

[0048] Threshold range: τ∈[T low ,T high], τ' ∈ [0, T high -T low ] can ensure that the medium confidence interval is not empty and does not overlap with the high confidence interval and the low confidence interval. Through the reverse adjustment mechanism of τ and τ', the dynamic balance of "computing power, precision and real-time performance" is intuitively demonstrated, reducing the calculation amount of the second stage to ensure real-time performance.

[0049] Second stage: fine segmentation

[0050] I. ROI cropping and original mask generation

[0051] The ROI (Region of Interest) of the medium confidence target region output by the first stage YOLOv8 is cropped, and the cropping is extended by 20% on the basis of the detection frame to avoid loss of target boundary information. The cropped region is used as input to generate the original mask, providing a basic region range for subsequent segmentation.

[0052] II. Improved U-Net encoder processing (as shown in Figure 4

[0053] 1. Feature extraction: input the cropped ROI image into the improved U-Net encoder, extract features through 4 groups of down-sampling convolutional layers, each group containing a ReLU activation function, gradually reducing the image size and extracting deep semantic features. In the bottleneck layer, generate deep features through double convolution.

[0054] 2. Attention enhancement: integrate channel-spatial attention modules in the encoder. Generate channel statistics through global average pooling and global maximum pooling, get channel weights through shared MLP layers and Sigmoid activation, enhance the expression of target feature channels; at the same time, aggregate the mean and maximum values of the channel dimension of the feature map, generate spatial weights through 7x7 convolution and Sigmoid activation, strengthen the target boundary features and reduce background noise interference.

[0055] III. Decoder processing and feature fusion

[0056] 1. Up-sampling and feature recovery: the decoder performs up-sampling through transpose convolutional layers, gradually recovering the resolution of the feature map. After each up-sampling, the feature map of the corresponding scale in the encoder is spliced through a jump connection to preserve the shallow detail information.

[0057] 2. Multi-scale feature fusion: through the symmetrical encoder-decoder structure, combine deep semantic information with shallow spatial details to ensure the accuracy and integrity of the segmentation results. In each decoding stage, use the channel-spatial attention module to further optimize feature expression and improve the detection accuracy of small targets and fuzzy targets.

[0058] ​Four, confidence update and result fusion

[0059] 1. Bayesian fusion strategy: Bayesian fusion strategy C final = aC YOLO + (1-a)C U-Net Update the confidence of the detection box, a e [0, 1]. Where a is the weight coefficient optimized by cross-validation, realizing the reliability fusion of YOLOv8 coarse detection result and U-Net segmentation result.

[0060] 2. Multi-confidence result merging: merge the detection results of high-confidence targets, low-confidence targets, and medium-confidence targets after fine segmentation and confidence update, process the conflict detection boxes through NMS (non-maximum suppression), and generate the final defect detection result.

[0061] Five, channel-space attention module implementation (as shown in Figure 5 )

[0062] 1. Channel attention construction: global average pooling and global maximum pooling are performed on the input feature map respectively to generate channel statistics Z avg and Z max , and the channel weight M c (f) is obtained after processing by shared MLP network and Sigmoid activation, and the feature map is weighted in channel dimension.

[0063] 2. Spatial attention construction: mean pooling and maximum pooling are performed on the feature map in channel dimension to generate f avg and f max , and the spatial weight M s (f) is obtained after splicing, 7x7 convolution and Sigmoid activation, and the feature map is weighted in spatial dimension to strengthen the target boundary region.

[0064] 3. Experiment and analysis

[0065] The experimental method adopts a two-stage cascading framework, cascading YOLOv8 and improved U-Net. In the initial stage, lightweight YOLOv8 helps to quickly detect objects, thus producing high-confidence and low-confidence samples. Subsequently, in the second stage, the remaining medium-confidence samples (which may include occluded targets) are processed by ROI cropping and input into the improved U-Net for detailed segmentation. In order to improve the depiction of target boundaries, a spatial attention module is integrated in the U-Net decoder. In addition, a dynamic threshold mechanism is also implemented to adaptively manage computing resources: devices with higher computing power prioritize accuracy by lowering the threshold, while devices with limited computing power focus on maintaining real-time performance by raising the threshold. The detailed contents of the experiment are as follows:

[0066] The experimental study adopted the publicly available NU-DEC (Northeastern University Defect Evaluation and Testing) dataset, which includes six common metal surface defect categories: crack, inclusion, patch, pitting, scratch, and oxide. The dataset contains 1800 high-resolution images (2000x2000 pixels) and 4326 annotated bounding boxes. The data partitioning follows an 8:1:1 ratio, resulting in training (1,440 images, 80%), validation (180 images, 10%), and test sets (180 images, 10%). To address the metal surface reflectivity issue and improve the robustness of the model, a multi-stage preprocessing pipeline was implemented: 1) image normalization to 640x640 resolution; 2) geometric transformation with random horizontal / vertical flipping; 3) CutOut occlusion (maximum 20% coverage); 4) mosaic augmentation (4 image synthesis); 5) adaptive brightness adjustment (±15%); 6) Gaussian noise injection (σ=0.05) to simulate industrial disturbances. The experiments were conducted on an Intel Core i5-13500 CPU (13th generation) equipped with NVIDIA GeForce RTX 3060 GPU using PyTorch 1.10. The training parameters included 640x640 input resolution, 300 epoch training duration, and 3 epoch learning rate warm-up. Dataset statistics and Settings specific to the augmentation strategy are shown in Table 1. Figure 2

[0067] The above examples are only used to illustrate the technical solutions of the present application and do not limit it. Other modifications or equivalent replacements of the technical solutions of the present application made by those skilled in the art should be covered within the scope of the claims of the present application, as long as they do not deviate from the spirit and scope of the technical solutions of the present application.​

Claims

1. A two-stage metal surface defect detection method based on computing power perception, and its detection scheme includes the following steps: Step 1: Use the YOLOv8 model to perform a rough detection on the input metal surface image, and generate candidate boxes containing defect positions and confidence levels; Step 2: Based on the real-time computing resource status of the device, dynamically calculate the high confidence threshold τ and the low confidence threshold τ', and divide the candidate boxes into the following three categories: High confidence (confidence ≥ τ): directly output the detection result; Medium confidence (τ' < confidence < τ): enter the fine detection process; Low confidence (confidence ≤ τ'): directly filter; Step 3: Perform regional expansion and cropping on the low confidence region to obtain the ROI region image, and input it into the improved U-Net network for pixel-level segmentation; Step 4: Update the confidence level through the Bayesian fusion strategy for the segmentation mask output by U-Net and the detection box of YOLOv8, and output the final defect classification and localization result.

2. The method according to claim 1, characterized in that The classification of candidate boxes and dynamic thresholds in Step 2 can be specifically divided into two steps: Step 2.1 Candidate box classification: K-means++ algorithm is used to cluster the confidence distribution of all candidate boxes output by YOLOv8. Based on the sample distribution characteristics, the confidence is divided into high confidence clusters and low confidence clusters. The centroids of the two clusters are determined to be the high confidence benchmark threshold μ high and the low confidence benchmark threshold μ low Then, based on the preliminary threshold, the confidence of the candidate boxes is divided into three categories: high confidence class, medium confidence class, and low confidence class; Step 2.2 Adaptively adjust the threshold based on the device computing power. First, define the device delay-related parameters, including the device real-time computing delay time t, the preset reference delay t0, and the delay fluctuation coefficient △t that can be dynamically adjusted through real-time computing ability monitoring; adjust the threshold according to the comparison result between the device delay t and the preset reference delay t0; if t < t0, it means that the device computing power is sufficient, and at this time τ increases and τ' decreases to expand the medium confidence interval and let more medium confidence samples enter U-Net for refined detection; if t ≥ t0, it means that the device computing power is insufficient, and at this time τ decreases and τ' increases to narrow the medium confidence interval and reduce the U-Net detection sample volume to ensure detection real-time performance; through the reverse adjustment mechanism of τ and τ', achieve the dynamic balance of "prioritizing accuracy when the computing power is sufficient and prioritizing real-time performance when the computing power is insufficient", and optimize the device resource utilization rate.

3. The method according to claim 1, characterized in that The steps of the improved U-Net network in Step 3 are: Step 3.1 Replace the backbone network of the traditional U-Net with the U-Net V2 network architecture, and perform spatial attention and channel attention operations on each level of features generated by the encoder using the semantic and detail fusion module to enhance the feature representation; Step 3.2 At each level of the decoder, adjust the size of the feature maps of other levels to be the same as the current level feature map, and smooth the adjusted feature maps with a 3×3 convolution, and then perform further operations on all processed feature maps, and input the obtained fusion features into the decoder of the current level, so as to achieve cross-level feature fusion and enhance the semantic and detail information of the features.

4. The method according to claim 1, wherein The specific steps of the Bayesian fusion strategy in Step 4 are: Step 4.1 Calculate the fusion confidence for the overlapping region of the YOLOv8 detection box and the U-Net segmentation mask.

5. The method according to claim 1, wherein Steps 1-4 include: Step 1.1 Cascade detection engine: including: YOLOv8 rough detection unit, configured with a CSPDarknet backbone network and an FPN-PAN feature fusion structure; Step 1.2: Adaptive threshold generation module: Adaptively generate τ and τ' according to the dynamic threshold method of claim 2; Step 1.3: Improve the U-Net fine detection unit and integrate the channel-spatial attention module; Step 1.4: Result fusion module: executing the Bayesian fusion algorithm described in claim 5; Step 1.5 Visualization interface: Output an overlay image with defect classification labels, positioning boxes, and segmentation masks.

Citation Information

Cited By

  • Target detection method and system based on scene prior and hierarchical detection mechanism

    CN121170276A

  • Target detection method and system based on scene prior and hierarchical detection mechanism

    CN121170276B