High-robustness multi-scale defect detection method and system for metal matrix surface

By introducing the EMA attention mechanism and Shape-IoU loss function into the YOLOv8 network, the problems of missing small targets and background interference in the detection of defects on the surface of metal substrates are solved, and efficient and accurate detection of multi-scale defects is achieved, which is suitable for complex industrial environments.

CN121707972APending Publication Date: 2026-03-20SOUTHWEST TECHNICAL ENGINEERING RESEARCH INSTITUTE OF CHINA SOUTH IND GROUP +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511902033.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies for detecting defects on metal substrate surfaces suffer from problems such as easy omission of small targets, strong background interference, and inaccurate localization of irregular defects. They are particularly inefficient and inaccurate in complex industrial environments.

Method used

An EMA attention mechanism is introduced to suppress background noise, a P2 micro-target detection head is constructed to retain shallow features, and Shape-IoU is combined to optimize localization accuracy. The YOLOv8 network architecture is improved, and the accuracy and robustness of detection are enhanced through multi-scale feature fusion.

Benefits of technology

It significantly reduces the false negative rate of early-stage minor defects, optimizes the positioning accuracy of defects with extreme aspect ratios and irregularities, achieves highly robust detection of multi-scale defects, and provides efficient detection support in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707972A_ABST
    Figure CN121707972A_ABST
Patent Text Reader

Abstract

The invention discloses a high-robustness multi-scale defect detection method for a metal matrix surface, and relates to the technical field of machine vision and surface defect detection. In order to solve the problems of large scale difference of metal matrix surface defects, strong background noise and easy missing detection of tiny targets in a complex service environment, an EMA-YOLO deep network proposed by the invention is characterized in that an efficient multi-scale attention (EMA) mechanism is introduced into a YOLOv8 backbone network, and the influence of background interferences such as surface water stains and oil stains on defect identification is inhibited through a cross-space learning strategy; a P2 layer-based high-resolution tiny target detection head is constructed, and shallow geometric details are reserved to accurately capture early pitting corrosion and microcracks; a Shape-IoU loss function is introduced, and through shape perception and adaptive weight adjustment, the positioning precision of cracks and irregular spalling defects with extreme length-width ratios is optimized. The EMA-YOLO deep network provided by the invention can effectively improve the accuracy, stability and robustness of multi-scale defect detection under a complex texture background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a surface defect detection method, and more particularly to a robust multi-scale defect detection method and system for metal substrate surfaces. Background Technology

[0002] Metal substrates are widely used in buildings, bridges, and large industrial facilities. During long-term service, coating surfaces are prone to defects such as corrosion, cracks, peeling, and blistering. Traditional inspection methods mainly rely on manual visual inspection, which is inefficient and highly subjective. In recent years, deep learning-based target detection technologies (such as the YOLO series) have been introduced into this field. However, in real-world, complex industrial environments, existing technologies have the following shortcomings:

[0003] 1) Small targets are easy to miss: Early pitting corrosion and microcracks are extremely small in scale, and their geometric features are easily lost during multiple downsampling processes of general networks (such as YOLOv8).

[0004] 2) Strong background interference: In the in-service scenario, background noise such as water stains and oil stains on the surface of the metal substrate is similar to the defect characteristics, which can easily lead to false detection.

[0005] Inaccurate localization of irregular defects: Existing bounding box regression loss functions (such as CIoU) are poorly adapted to cracks or irregular spalling defects with extreme aspect ratios. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, this invention provides a robust multi-scale defect detection method and system for metal substrate surfaces. This method introduces an EMA attention mechanism to suppress background noise, constructs a P2 micro-target detection head to retain shallow features, and combines Shape-IoU to optimize positioning accuracy, thereby improving the accuracy, stability, and robustness of detection results in complex industrial environments.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] A robust multi-scale defect detection method for metal substrates is proposed. This method incorporates an efficient multi-scale attention (EMA) mechanism into the feature extraction stage of the general YOLOv8 backbone network. The EMA module is embedded at the end of the C2f module structure of the YOLOv8 network backbone. Through grouped convolution and cross-spatial learning strategies, feature weights are recalibrated at the pixel level to suppress background noise interference from water stains and oil contaminants on the metal substrate surface. Simultaneously, the conventional three-head detection architecture is improved by introducing a shallow C2 feature layer in the backbone network to construct a high-resolution P2 micro-target detection head. This head utilizes rich geometric details to capture early pitting corrosion and micro-cracks, ensuring that micro-defect features are not lost during multiple downsampling processes. Furthermore, a single-stage detection technique based on shape perception and multi-scale feature fusion is employed to improve the accuracy, stability, and robustness of detecting irregular defects such as corrosion, cracks, and spalling on the metal substrate surface under complex service environments.

[0009] As a further improvement to the present invention, the specific steps are as follows:

[0010] S1: Build an image acquisition system to obtain images of metal substrate surfaces under real service conditions, and construct a high-resolution dataset containing four typical defects: corrosion, cracks, spalling, and blistering.

[0011] S2: Preprocess the dataset by uniformly scaling the images to 640×640 pixels and dividing it into training and validation sets to meet the resolution requirements of the input images for network model training.

[0012] S3: Feed the preprocessed image into the improved YOLO v8 backbone network for training; in layers P2 to P5 of the backbone network, replace the original C2f module with the C2f_EMA module, extract multi-scale spatial semantic information through parallel substructures and aggregate it to output a clean feature map.

[0013] S4: The features extracted from the backbone network are fed into the neck network; a feature branch is drawn from the second layer of the backbone network, and after upsampling, it is fused with the neck features to generate a P2 feature map with a resolution of 160×160. A four-scale detection head structure containing P2, P3, P4 and P5 is constructed to complete the feature preservation and enhancement of sub-pixel-level small targets.

[0014] S5: In the bounding box regression stage of model training, the Shape-IoU loss function is introduced to replace the traditional CIoU loss; by calculating the shape and scale weight coefficients of the predicted box and the real box in the horizontal and vertical directions, the penalty term is dynamically adjusted to achieve accurate localization of cracks and irregular peeling defects with extreme aspect ratios.

[0015] S6: Perform non-maximum suppression (NMS) processing to output the final defect category, confidence level, and location coordinates, achieving efficient identification of multi-scale coating defects.

[0016] As a further improvement of this invention, the main purpose of constructing the dataset and preprocessing in step S1 is to resolve the contradiction between the extremely high resolution of the original images collected in the industrial field and the limited input size of the detection model, to prevent the loss of texture features of minor defects (such as pitting and microcracks) due to direct scaling, and to solve the problem of unbalanced distribution of sample categories; specifically:

[0017] S-A1: Collect images of metal substrate surfaces under real service conditions. The objects collected cover four typical defects: corrosion, peeling, cracks, and blistering. The background of the images must include water stains, oil stains, and uneven lighting interference to ensure the diversity and authenticity of the data.

[0018] S-A2: The sliding window segmentation algorithm is used to process the high-resolution original image. The segmentation window size is set to 1200 × 1200 pixels, and the sliding step size is set to be smaller than the window size to generate overlapping areas, preventing defects from being truncated, and cropping the large-format original image into multiple local sub-images.

[0019] S-A3: The segmented sub-images were manually screened and cleaned to remove blurry or incorrectly labeled samples. A total of 8193 valid defective samples were selected and randomly divided into training and validation sets in an 8:2 ratio to ensure the independence of training and validation data.

[0020] S-A4: Data augmentation is performed on the training set samples. The Mosaic augmentation strategy is used to enrich the background information, and the target scale distribution in the matching dataset is calculated using adaptive anchor boxes.

[0021] S-A5: The processed image is uniformly scaled to a fixed size of 640 × 640 pixels using a bilinear interpolation algorithm, and the pixel values ​​are normalized to the [0,1] interval, transforming it into a tensor format that meets the input requirements of neural networks.

[0022] As a further improvement of the present invention, the processing of the efficient multi-scale attention mechanism (EMA) includes: grouping the input feature map into multiple sub-feature sets representing different spatial levels in the channel dimension; constructing three parallel branches to extract attention weights: the first two branches perform one-dimensional global average pooling and convolutional encoding in the X and Y directions respectively, and the third branch performs 3×3 convolution to capture cross-channel feature interactions; using the cross-spatial learning module, the outputs of the parallel branches are fused, and pixel-level attention weight allocation is achieved through the Softmax function and matrix inner product operation.

[0023] As a further improvement of the present invention, the specific method for constructing the P2 micro-target detection head in step S4 is as follows: a feature branch is drawn from the second layer (C2 layer) of the backbone network; the feature of this branch is upsampled through the neck network and fused with the deep feature; a P2 feature map with a resolution of 160 × 160 is generated, the P2 feature map corresponds to a 4-fold downsampling of the input image, and is used to detect micro-cracks and early pitting defects.

[0024] As a further improvement of the present invention, parameter calculation is performed in step S5 to obtain the bounding box regression loss, specifically as follows:

[0025] S-B1: Calculate the horizontal weighting coefficient and vertical weighting coefficient :

[0026]

[0027]

[0028] in, This is the width of the ground truth box. The actual height of the bounding box; The scaling factor is related to the target scale of the dataset; the scaling factor is used to adjust the weight attention on different axes using the above formula.

[0029] S-B2: Establish a coordinate system and calculate the shape distance between the center point of the predicted bounding box and the center point of the ground truth bounding box. :

[0030]

[0031] in, The x-coordinate of the center point of the prediction box. The ordinate of the center point of the prediction box; The x-coordinate of the center point of the true bounding box. The ordinate of the center point of the true bounding box; The diagonal distance between the minimum bounding box that covers the predicted box and the ground truth box; by introducing and This makes distance calculations direction-sensitive;

[0032] S-B3: Calculate the shape term that reflects differences in geometric form. :

[0033]

[0034] in, Variables representing the relative deviation in the width and height directions; The base of the natural logarithm; To control the hyperparameters of shape attention, this step quantifies the geometric fit between the predicted bounding box and the ground truth bounding box through exponential calculation.

[0035] S-B4: Combining the above parameters, calculate the final Shape-IoU bounding box regression loss. :

[0036]

[0037] in, Intersection over Union (IUCN) is the intersection-over-union ratio between the predicted bounding box and the ground truth bounding box, representing the ratio of the area of ​​the overlapping region of the two bounding boxes to the area of ​​the joint region. The above formula is a weighted sum of the IUCN loss, the center point distance loss considering the orientation weight, and the shape geometry loss, which serves as the basis for the backpropagation optimization of the network.

[0038] As a further improvement of the present invention, the surface defects of the metal substrate in step S1 include four types: corrosion, peeling, cracks and blistering; and the preprocessing includes uniformly scaling the image to a size of 640×640 pixels.

[0039] This invention also provides a highly robust multi-scale defect detection system for metal substrate surfaces, comprising:

[0040] Image acquisition module, used to acquire high-resolution images of the surface of the metal substrate;

[0041] A processor, configured with memory and a computer program, which executes the computer program to implement the above-described method;

[0042] The output module is used to display or output information about the type, location, and confidence level of defects.

[0043] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0044] The technical advantages of this invention are as follows: By improving the YOLOv8 network architecture and introducing the C2f_EMA module in the feature extraction stage, this invention utilizes the EMA mechanism to aggregate cross-spatial information, effectively suppressing background noise from water stains and oil contaminants on the coating surface. By constructing a high-resolution P2 micro-target detection head, rich geometric details of the 4x downsampling layer are preserved, significantly reducing the false negative rate of early-stage micro-defects. Furthermore, by introducing the Shape-IoU loss function, the localization accuracy for cracks and irregular defects with extreme aspect ratios is optimized. This method achieves robust detection of multi-scale defects while maintaining high detection accuracy, providing effective technical support for early warning and maintenance of steel structure corrosion. Attached Figure Description

[0045] Figure 1 This is a flowchart of a highly robust multi-scale defect detection method for metal substrate surfaces according to the present invention;

[0046] Figure 2 Corrosion diagram of steel structure;

[0047] Figure 3 This is a schematic diagram of the C2f_EMA module structure;

[0048] Figure 4 A schematic diagram of the improved YOLO model;

[0049] Figure 5 This is a schematic diagram of feature fusion in the P2 micro-target detection head;

[0050] Figure 6 This is a diagram showing the effect of the model detecting defects. Detailed Implementation

[0051] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0052] A robust multi-scale defect detection method for metal substrate surfaces is proposed. This method uses an improved YOLO model and adopts a technical approach based on shape perception and multi-scale feature fusion to help the system accurately locate defects such as corrosion and cracks.

[0053] I. Environmental Configuration

[0054] 1. Hardware Environment Configuration

[0055] Central Processing Unit (CPU): Intel Core i7-12700F processor;

[0056] Graphics Processing Unit (GPU): It adopts an NVIDIA GeForce RTX 4090 graphics card with 24GB of video memory to provide sufficient computing power to support the parallel computing of deep learning models;

[0057] Memory: Equipped with 64GB of memory to meet data loading requirements.

[0058] 2. Software Environment Configuration

[0059] Development tools: Visual Studio Code 2023 is used as the integrated development environment, with Anaconda3 for environment management;

[0060] Programming language and framework: Based on Python 3.8, with PyTorch 2.1.2 as the deep learning framework;

[0061] Acceleration Library: Configure CUDA 12.1 to utilize GPU for accelerated computation.

[0062] 3. Training hyperparameter settings

[0063] Input image size: uniformly adjusted to 640 × 640 pixels;

[0064] Batch Size: Set to 16 to balance memory usage and gradient descent stability;

[0065] Training epochs: A total of 400 training epochs are conducted to ensure that the model fully converges;

[0066] Optimizer: The AdamW optimizer is used for parameter updates;

[0067] Learning Rate: The initial learning rate is set to 0.001;

[0068] Momentum: The parameter is set to 0.98 to accelerate model convergence and suppress oscillations;

[0069] Weight Decay: The coefficient is set to 0.0005 to prevent the model from overfitting;

[0070] Early Stopping Strategy: Automatically terminate training when the validation set loss no longer decreases within 50 consecutive epochs;

[0071] Loss weights: Classification loss weights Set to 1.0, bounding box regression loss weights Set it to 7.5 to guide the model to focus more on positioning accuracy.

[0072] II. Specific Implementation Steps

[0073] As shown in Figure 1, a robust multi-scale defect detection method for metal substrate surfaces mainly includes the following steps:

[0074] S1: Construct a high-resolution coating defect dataset by collecting surface images of steel structures under real service conditions, covering four typical defects: corrosion, peeling, cracking, and blistering, as shown in Figure 2. The acquisition environment included water stains, oil stains, and uneven lighting interference.

[0075] S2: Dataset Preprocessing. Due to the extremely high resolution of the original images, a sliding window strategy was adopted: the segmentation window was set to 1200×1200 pixels, and the stride was set smaller than the window size to generate overlapping regions and prevent defects from being truncated. After screening, a total of 8193 valid samples were obtained, and the training and validation sets were divided in an 8:2 ratio. Before training, the images were uniformly scaled to 640×640 pixels.

[0076] S3: Feed the image into the improved backbone network for feature extraction. For example... Figure 3 , Figure 4 As shown, in layers P2 to P5 of the YOLOv8 backbone network, the C2f_EMA module is used to replace the original C2f module.

[0077] Processing procedure: The input feature map is processed by C2f and then enters the EMA module. Through grouped convolution and cross-spatial learning strategies, pixel-level attention weights are generated using matrix inner product and softmax operations.

[0078] Results: Effectively suppresses water stains and oil noise in the background during the feature extraction stage.

[0079] S4: Construct the P2 small target detection head and perform feature fusion. For example... Figure 5 As shown, to address the problem of early microcracks (only 1-2 pixels wide) being easily missed, a high-resolution branch (160×160, 4x downsampling) is drawn from the second layer (C2 layer) of the backbone network.

[0080] Specifically, this branch is upsampled by the neck network and fused with deep features, then input into a dedicated P2 detection head to form a four-head detection architecture (P2, P3, P4, P5). The EMA module groups the feature maps along the channel dimension.

[0081] Results: It preserves rich shallow geometric details and significantly reduces the rate of missed detection of minute defects.

[0082] S5: Apply the Shape-IoU loss function for model training. In the bounding box regression stage, Shape-IoU is introduced to replace CIoU.

[0083] Principle: By introducing shape and scale factors, the weight coefficients of the predicted bounding box and the ground truth bounding box in the horizontal and vertical directions are calculated, with a focus on the consistency of aspect ratio. The feature map resolution is 160 × 160 (corresponding to 4 times downsampling of the input image);

[0084] Results: Improved the localization accuracy for cracks and irregular spalling defects with extreme aspect ratios. In this step, the model was trained for 400 rounds using the dataset constructed in S1 and the aforementioned training parameters, and the optimal weights were saved.

[0085] S6: Inference and Output (NMS Processing). Input the image to be tested, and the model outputs the defect category, confidence score, and coordinates. Finally, Non-Maximum Suppression (NMS) processing is performed to remove overlapping redundant boxes, and the final detection result is output (e.g., ...). Figure 6 (As shown).

[0086] III. Verification of Experimental Results

[0087] Under the experimental conditions described above, the performance of the EMA-YOLO model proposed in this embodiment on the test set is as follows:

[0088] Precision: 81.7%, a 3.0% improvement over the baseline model YOLOv8m;

[0089] The mean accuracy (mAP50) reached 66.6%, which is superior to the benchmark model and the latest YOLOv11m (61.4%). Experiments have demonstrated that the method of this invention achieves robust detection of multi-scale defects while maintaining high detection accuracy, and can effectively meet the needs of early warning of steel structure corrosion in industrial sites.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A robust multi-scale defect detection method for metal substrate surfaces, characterized in that: An efficient multi-scale attention (EMA) mechanism is added to the feature extraction stage of the general YOLOv8 backbone network. The EMA module is embedded into the C2f module of the YOLOv8 network backbone. Through grouped convolution and cross-space learning strategies, feature weights are recalibrated at the pixel level to suppress background noise interference from water stains and oil stains on the metal substrate surface. At the same time, the conventional three-head detection architecture is broken, and a shallow C2 feature layer of the backbone network is introduced to construct a high-resolution P2 micro-target detection head. The rich geometric details of the shallow layer are used to capture early pitting corrosion and micro-cracks, ensuring that micro-defect features are not lost during multiple downsampling processes. Meanwhile, a single-stage detection technology route based on shape perception and multi-scale feature fusion is adopted to improve the accuracy, stability and robustness of metal substrate surface texture for detecting irregular defects such as corrosion, cracks and spalling in complex service environments.

2. The method for detecting highly robust multi-scale defects on a metal substrate surface according to claim 1, characterized in that, The specific steps are as follows: S1: Build an image acquisition system to acquire images of the surface of metal substrates in service scenarios, and construct a high-resolution dataset containing four typical defects: corrosion, cracks, spalling and blistering. S2: Preprocess the dataset by uniformly scaling the images to 640×640 pixels and dividing it into training and validation sets to meet the resolution requirements of the input images for training the network model. S3: The preprocessed image is fed into the improved YOLOv8 backbone network for training; in layers P2 to P5 of the backbone network, the original C2f module is replaced by the C2f_EMA module, and multi-scale spatial semantic information is extracted and aggregated through parallel substructures to output a clean feature map. S4: The features extracted from the backbone network are fed into the neck network. Feature branches are drawn from the second layer of the backbone network. After upsampling, they are fused with the neck features to generate a P2 feature map with a resolution of 160×160. A four-scale detection head structure containing P2, P3, P4 and P5 is constructed to complete the feature preservation and enhancement of sub-pixel-level micro targets. S5: In the bounding box regression stage of model training, the Shape-IoU loss function is introduced to replace the traditional CIoU loss; by calculating the shape and scale weight coefficients of the predicted box and the real box in the horizontal and vertical directions, the penalty term is dynamically adjusted to achieve accurate localization of cracks and irregular peeling defects with extreme aspect ratios. S6: Perform non-maximum suppression (NMS) processing to output the final defect category, confidence level, and location coordinates, achieving efficient identification of multi-scale coating defects.

3. The method for detecting highly robust multi-scale defects on a metal substrate surface according to claim 2, characterized in that, In step S1, the primary purpose of constructing the dataset and preprocessing is to resolve the contradiction between the extremely high resolution of the original images acquired in complex in-service scenes and the limited input size of the detection model, to prevent the loss of texture features of minor defects due to direct scaling, and to address the problem of imbalanced sample class distribution; specifically: S-A1: Collect images of metal substrate surfaces in real-world service environments. The collected objects cover four typical defects: corrosion, spalling, cracks, and blistering. The image background must include water stains, oil stains, and uneven lighting interference to ensure the diversity and authenticity of the data. S-A2: The sliding window segmentation algorithm is used to process high-resolution original images. The segmentation window size is set to 1200×1200 pixels, and the sliding step size is set to be smaller than the window size to generate overlapping areas, preventing defects from being truncated, and cropping the large-format original image into multiple local sub-images. S-A3: The segmented sub-images were manually screened and cleaned to remove blurry or incorrectly labeled samples. A total of 8193 valid defective samples were selected and randomly divided into training and validation sets in an 8:2 ratio to ensure the independence of training and validation data. S-A4: Data augmentation is performed on the training set samples. The Mosaic augmentation strategy is used to enrich the background information, and the target scale distribution in the matching dataset is calculated using adaptive anchor boxes. S-A5: The processed image is uniformly scaled to a fixed size of 640×640 pixels using a bilinear interpolation algorithm, and the pixel values ​​are normalized to the [0,1] interval, converting it into a tensor format that meets the input requirements of a deep convolutional neural network.

4. The highly robust multi-scale defect detection method for a metal substrate surface according to claim 2, characterized in that, The efficient multi-scale attention mechanism (EMA) process includes: grouping the input feature map into multiple sub-feature sets representing different spatial levels along the channel dimension; constructing three parallel branches to extract attention weights: the first two branches perform one-dimensional global average pooling and convolutional encoding in the X and Y directions respectively, and the third branch performs 3×3 convolution to capture cross-channel feature interactions; using the cross-spatial learning module, the outputs of the parallel branches are fused, and pixel-level attention weight allocation is achieved through the Softmax function and matrix inner product operation.

5. The method for detecting highly robust multi-scale defects on a metal substrate surface according to claim 2, characterized in that, The specific method for constructing the P2 micro-target detection head in step S4 is as follows: a feature branch is drawn from the second layer (C2 layer) of the backbone network; the features of this branch are upsampled through the neck network and fused with the deep features; a P2 feature map with a resolution of 160×160 is generated. The P2 feature map corresponds to a 4-fold downsampling of the input image and is used to detect micro-cracks and early pitting defects.

6. The method for detecting highly robust multi-scale defects on a metal substrate surface according to claim 2, characterized in that, In step S5, parameter calculation is performed to obtain the bounding box regression loss, specifically as follows: S-B1: Calculate the horizontal weighting coefficient and vertical weighting coefficient : , , in, This is the width of the ground truth box. The actual height of the bounding box; This is a scaling factor related to the target scale of the dataset; using the above formula, Factor adjustment of the weighting of different axes; S-B2: Establish a coordinate system and calculate the shape distance between the center point of the predicted bounding box and the center point of the ground truth bounding box. : , in, The x-coordinate of the center point of the prediction box. The ordinate of the center point of the prediction box; The x-coordinate of the center point of the true bounding box. The ordinate of the center point of the true bounding box; The diagonal distance between the minimum bounding box that covers the predicted box and the ground truth box; by introducing and This makes distance calculations direction-sensitive; S-B3: Calculate the shape term that reflects differences in geometric form. : , in, Variables representing the relative deviation in the width and height directions; Hyperparameters for controlling shape attention; The base of the natural logarithm is represented by an exponential function; this step quantifies the geometric fit between the predicted and actual bounding boxes through exponential operations. S-B4: Based on the above parameters, calculate the final Shape-IoU bounding box regression loss. : , in, Intersection over Union (IUCN) is the intersection-over-union ratio between the predicted bounding box and the ground truth bounding box, representing the ratio of the area of ​​the overlapping region of the two bounding boxes to the area of ​​the joint region. The above formula is a weighted sum of the IUCN loss, the center point distance loss considering the orientation weight, and the shape geometry loss, which serves as the basis for the backpropagation optimization of the network.

7. The method for detecting highly robust multi-scale defects on a metal substrate surface according to claim 2, characterized in that, The surface defects of the metal substrate in step S1 include four types: corrosion, peeling, cracks, and blistering; and the preprocessing includes uniformly scaling the image to a size of 640×640 pixels.

8. A highly robust multi-scale defect detection system for metal substrate surfaces, characterized in that, include: Image acquisition module, used to acquire high-resolution images of the surface of the metal substrate; A processor, configured with memory and a computer program, wherein the processor executes the computer program to implement the method as described in any one of claims 1 to 7; The output module is used to display or output information about the type, location, and confidence level of defects.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Drainage pipeline defect detection method and device based on cost-sensitive risk decision and scale perception optimization, and computer readable storage medium

    CN122049873A