Metal surface defect detection method based on multi-scale dynamic convolution and space attention

By improving the YOLOv8 network structure and combining multi-scale dynamic convolution and spatial attention modules, the problems of multi-scale feature capture and background interference suppression in surface defect detection of petal-type metal components are solved, and high-precision, low-power real-time detection is achieved, which is suitable for complex industrial equipment.

CN120634972APending Publication Date: 2025-09-12SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510679147.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies for surface defect detection of petal-type metal components have problems such as difficulty in capturing multi-scale features, insufficient background interference suppression, and low efficiency of mobile terminal deployment, making it difficult to achieve high-precision, low-power real-time detection.

Method used

Multi-scale dynamic convolution and spatial attention are used to improve the YOLOv8 network structure. The multi-scale dynamic deformable convolution module (MDConv) and the channel-spatial layered attention module (CPMS) work together, combined with data enhancement and loss function optimization to improve the model's adaptability to cross-scale defects and background interference suppression capabilities. Hardware optimization is also performed to adapt to industrial equipment.

Benefits of technology

It significantly improves detection accuracy and robustness, reduces false detection rate and power consumption, achieves efficient detection of cracks from sub-pixel to millimeter level, supports the real-time and low power consumption requirements of industrial equipment, reduces annual maintenance costs and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634972A_ABST
    Figure CN120634972A_ABST
Patent Text Reader

Abstract

The invention discloses a metal surface defect detection method based on multi-scale dynamic convolution and space attention, and the method comprises the steps: carrying out the following improvement on the basis of a YOLOv8 frame for solving the problems of strong background interference and low multi-scale defect detection precision in split equipment: introducing multi-scale dynamic deformable convolution MDConv, a dynamic expansion rate adjusting mechanism realizes 11 * 11 to 33 * 33 multi-scale receptive field adaptive matching, and the sub-pixel-level crack detection capability is improved; and embedding an attention module CPMS guided by a geometric mask in a feature fusion path, and inhibiting joint texture false detection. According to the method, the technical bottlenecks of high micro crack omission ratio and low hardware deployment efficiency in a complex industrial scene are solved, and the method is suitable for high-precision metal surface quality inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image detection and machine learning technology. Specifically, it is a metal surface defect detection method based on multi-scale dynamic convolution and spatial attention. It is particularly suitable for real-time detection of sub-pixel to millimeter-level cracks in industrial equipment with complex geometric structures such as split-type water-cooled crucibles and semiconductor single crystal growth furnaces. Background Art

[0002] The development of metal surface defect detection technology has gone through three major stages: traditional image processing, machine vision, and deep learning. Its core challenge has always revolved around balancing accuracy and efficiency in complex industrial scenarios. Traditional image processing methods rely on threshold segmentation and morphological operations, and can achieve a detection accuracy of 85% under simple backgrounds. However, when faced with sub-pixel differences in seam texture (0.2-0.5mm) and microcracks (<50μm) in split-flap equipment, manual feature engineering has difficulty effectively distinguishing the slight contrast changes (0.12-0.15) in the grayscale co-occurrence matrix, and the false detection rate is generally higher than 25%. Such methods are not robust enough to dynamic interference such as lighting fluctuations and metal surface reflections, and the efficiency of high-resolution image processing cannot meet real-time requirements.

[0003] While the introduction of machine vision technology has advanced detection capabilities through multispectral imaging and 3D reconstruction, its high hardware cost (over $50,000 per system) and environmental sensitivity (false positive rates fluctuating 15%-20% due to water mist and oil stains) have limited its widespread adoption. Deep learning models, such as the YOLO series, have significantly improved detection efficiency through multi-scale feature pyramids (FPN / PANet). Improved models like YOLOv8-EDGE achieve a mean average precision (mAP@0.5) of 90-93% on industrial datasets. However, significant bottlenecks remain. First, the model's adaptability to defects across scales is limited. The recall rate difference between a 50μm microcrack and a 2mm extended crack exceeds 35%, making it difficult for a single convolution kernel to capture dynamic features from the subpixel to millimeter scale. Second, the geometric texture of the seam area of ​​​​split-flap equipment highly overlaps with the crack morphology (with a similarity of up to 65%). Traditional attention mechanisms (such as CBAM) have limited ability to suppress background interference, resulting in a persistent false positive rate exceeding 10%.

[0004] Existing technologies are further limited by the specificities of industrial scenarios: deep learning models rely on large-scale annotated data, but the scarcity of microcrack samples (less than 5%) and the cost of annotation (over $3.5 per sample) significantly reduce model generalization capabilities. At the hardware deployment level, mainstream models (such as YOLOv8n) experience inference delays exceeding 20ms and power consumption exceeding 2W on mobile NPUs, making them difficult to adapt to the real-time and low-power requirements of industrial equipment. In multispectral imaging, issues such as feature conflicts between UV and IR channels and loss of edge gradients due to quantization distortion increase the risk of missed detection of minor defects. These limitations severely restrict the application of intelligent inspection for key equipment such as split-flap water-cooled crucibles, necessitating an innovative solution that balances accuracy, efficiency, and engineering applicability. Summary of the Invention

[0005] In response to the core problems existing in the surface defect detection of petal-type metal components, such as the difficulty in capturing multi-scale features, insufficient background interference suppression, and inefficient mobile terminal deployment, the present invention proposes a metal surface defect detection method based on multi-scale dynamic convolution and spatial attention.

[0006] The technical solution adopted by the present invention is: a metal surface defect detection method based on multi-scale dynamic convolution and spatial attention, characterized by comprising the following steps:

[0007] S1, collect metal surface images, pre-process and create a dataset for training;

[0008] S2 combines the multi-scale dynamic deformable convolution module MDConv and the channel-spatial layered attention module CPMS with geometric masks to improve the YOLOv8 network structure and loss function. The improved network is iteratively trained using dataset images, and the network parameters are optimized to obtain an ideal model for metal surface defect detection.

[0009] S3. Actual metal surface images are collected and pre-processed before being input into the ideal model, and the defect and defect location regression boxes are automatically input.

[0010] The pretreatment includes:

[0011] Generate geometric mask G based on the petal structure CAD model mask , mark the coordinate set of the seam area;

[0012] The standard deviation and kernel size are set, Gaussian noise and motion blur are randomly added, and images simulating interference in industrial scenes are generated.

[0013] The metal surface images include defects of various crack widths and images without crack defects, and are labeled using binary classification.

[0014] The Cn2f module of the YOLOv8 network structure is improved into a multi-scale dynamic deformable convolution module MDConv. By integrating the deformable convolution kernel with the dynamic expansion rate adjustment mechanism of the void convolution, it achieves adaptive matching of multi-scale receptive fields from 11×11 to 33×33.

[0015] The mathematical expression of MDConv is:

[0016] Among them, the dynamic expansion rate d k It is generated through global average pooling and lightweight convolution, and the offset Δp is generated using a single-layer convolution and Tanh activation function.

[0017] The channel-spatial hierarchical attention module (CPMS) is embedded in the PANet feature fusion path of the YOLOv8 network structure, and a multi-step pyramid compression-stimulated channel attention and geometric mask-guided spatial attention collaborative working mechanism are adopted. The operations of the CPMS module include:

[0018] Features of different scales are extracted through multi-scale pooling layers with unequal step sizes and Sigmoid gating mechanism;

[0019] The geometric mask G generated based on the petal structure CAD model is used mask Fusion layer with spatial attention realizes invalid feature filtering through matrix dot multiplication to suppress false detection of seam texture.

[0020] The method adopts a phased training strategy: in the initial stage, the original YOLOv8 parameters are frozen and only the MDConv and CPMS modules are trained; in the fine-tuning stage, a cosine annealing learning rate is used for end-to-end optimization.

[0021] In the loss function, the classification loss weight is increased to 1.5 times, and the mask IoU loss constraint is added;

[0022] The expression is: L mask =1-IoU(G mask ,S att )

[0023] Among them, G mask Geometric mask, S att is the spatial attention output.

[0024] The present invention has the following beneficial effects and advantages:

[0025] 1. Breakthrough in detection accuracy and robustness:

[0026] The synergy between multi-scale dynamic deformable convolution (MDConv) and channel-spatial layered attention (CPMS) significantly improves the model's adaptability to cross-scale defects. The MDConv module uses a dynamic expansion rate adjustment mechanism to adaptively expand the receptive field of the convolution kernel according to the crack scale (11×11 to 33×33). On the split-flap crucible dataset, the characteristic response intensity of 50μm microcracks increased by 42%, and the recall rate of millimeter-scale expansion cracks reached 98.2%, an improvement of 23 percentage points compared to the traditional FPN structure. By fusing the geometric mask prior of the split-flap seams with multi-step frequency domain feature screening, the CPMS module reduces the false detection rate in the seam area from 12.7% to 1.5%, effectively suppressing 89% of false positive samples caused by dynamic interference such as cooling pipe shadows and oxide scale shedding.

[0027] 2. Innovation in hardware deployment efficiency:

[0028] In response to the stringent computing power constraints of industrial mobile devices, this solution adopts a parameter sharing strategy and integer shift operation optimization to achieve lightweight deployment while maintaining model accuracy. The improved YOLOv8 model has compressed the number of parameters to 24.7M (the original model was 36.4M), and the computational complexity has been reduced to 7.0GFLOPs. Combined with TensorRT INT8 quantization technology, the model size is only 6.7MB. The single-frame inference time is shortened to 8ms, supporting 35FPS real-time detection, with a typical power consumption of 1.2W, which is 52% lower than the original model. The model also supports a wide temperature range of -20℃ to 60℃ and a humidity range of 10% to 95%, adapting to the high temperature and high humidity environment of metallurgical workshops. The false alarm rate after 720 hours of continuous operation is less than 0.1%.

[0029] 3. Engineering applicability and economic benefits:

[0030] The technical achievements of this invention have been successfully applied to high-end manufacturing fields such as semiconductor single crystal growth furnaces and aerospace titanium alloy melting equipment. Actual cases have shown that the system's detection sensitivity for microcracks in split crucibles reaches 30μm (traditional methods are 100μm), annual maintenance costs are reduced by 72%, and the equipment's mean time between failures (MTBF) has increased from 400 hours to 2000 hours. In addition, the model supports offline deployment and multispectral data fusion, and can be seamlessly connected to the Industrial Internet of Things platform, achieving full process automation for defect location, dimensional measurement, and risk warning, which is 17 times more efficient than manual inspections. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Flow chart of the method of the present invention;

[0032] Figure 2 This is a diagram of the improved YOLOv8 model structure of the method of the present invention;

[0033] Figure 3 This is a diagram of the multi-scale dynamic deformable convolution structure of the method of the present invention;

[0034] Figure 4 This is a channel-space hierarchical attention structure diagram of the method of the present invention;

[0035] Figure 5 This is a crack detection effect diagram of the method of the present invention. DETAILED DESCRIPTION

[0036] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] like Figure 1 As shown, the present invention includes the following steps:

[0038] Step 1: Data preparation and preprocessing

[0039] 1. Collect grayscale images of the surface of the split-type water-cooled crucible (resolution 2048×2048, 16-bit depth);

[0040] 2. Generate geometric mask G based on CAD model mask , mark the seam area coordinate set

[0041] 3. Data augmentation: Randomly add Gaussian noise (σ = 0.01) and motion blur (kernel size 5 × 5) to simulate industrial site interference.

[0042] Step 2: Model building and training

[0043] 1. Module deployment:

[0044] In the Backbone layer (Stage 1-4) of YOLOv8, the original C2f module is replaced with the MDConv module, where:

[0045] The shallow layer (Stage 1-2) adopts a lightweight design with a channel compression ratio of 1:4;

[0046] The deep layer (Stage3-4) retains all channel parameters to maintain the integrity of semantic features;

[0047] Insert the CPMS module into the upsampling and downsampling branches of PANet, and perform matrix multiplication between the spatial attention weight and the predefined geometric mask (Gmask);

[0048] 2. Training configuration:

[0049] In the initial stage, the original YOLOv8 parameters are frozen and only the MDConv and CPMS modules are trained (lr = 0.001);

[0050] The fine-tuning phase uses end-to-end optimization with a cosine annealing learning rate (0.01→0.0001);

[0051] 3. Loss function:

[0052] The classification loss weight is increased to 1.5 times to enhance microcrack identification;

[0053] Add mask IoU loss (L_mask = 1-IoU(Gmask, Satt)) to constrain the attention mechanism.

[0054] Step 3: Mobile deployment optimization

[0055] Model quantization: Using TensorRT INT8 calibration, the model size is compressed to 6.7MB;

[0056] NPU adaptation: MDConv integer operations are optimized by the processor instruction set, reducing memory usage to 256MB;

[0057] Real-time inference: Achieve 35FPS detection on the CSK6-MIXAI chip (power consumption 1.2W), supporting offline deployment and WiFi data transmission.

[0058] Step 4: Industrial Scenario Verification

[0059] Accuracy Verification: On the split crucible test set, mAP@0.5 reached 96.7%, and the recall rate of 50μm microcracks was improved by 42% (compared to traditional methods);

[0060] False detection rate test: In scenarios with seam texture interference, the false detection rate is stable below 1.5%, meeting the ASTM E2478 industry standard.

[0061] Power consumption and latency: During 8 hours of continuous operation, the NPU consumed an average of 1.2W of power, and single-frame detection took 8ms±0.3ms, with no performance degradation.

[0062] Data preparation and preprocessing in step 1 include:

[0063] Collect metal surface images and pre-process them to create a dataset for training; the metal surface images include defects with various crack widths and images without cracks;

[0064] Generate geometric mask G based on the petal structure CAD model mask , mark the coordinate set of the joint area; the marking adopts the binary classification of error defects.

[0065] The standard deviation and kernel size are set, Gaussian noise and motion blur are randomly added, and images simulating interference in industrial scenes are generated.

[0066] Model building and training in step 2 include:

[0067] YOLOv8 model: To address the core issues of multi-scale feature capture difficulties, insufficient background interference suppression, and inefficient mobile deployment in the detection of surface defects of split-type metal components, YOLOv8, as the leading real-time target detection framework in the current industrial inspection field, has advanced technology in three core dimensions: multi-scale feature fusion efficiency, hardware adaptability, and modular scalability. It has become the base framework selected for this patented method. This framework uses an improved C2f (CrossStage Partial-fusion) module to reconstruct the backbone network. Through cross-stage partial connections and dual convolutional branch design, it reduces the computational redundancy of Stage 3-4 layers by 32% while maintaining the resolution capability of high-resolution feature maps (such as 2048×2048). Compared with the previous generation YOLOv5, its parameter size has been compressed to 28.1M (YOLOv8s version), and the initial mAP@0.5 in the split-type crucible detection task has reached 93.1%, providing a high starting point benchmark for subsequent algorithm optimization. At the feature fusion level, YOLOv8's PAN-FPN structure achieves dynamic weighted fusion of multi-level semantic features through bidirectional cross-scale connections. Specifically, the 1024×1024 feature map output by the deep network (Stage4) and the 512×512 feature map of the shallow layer (Stage1) are adaptively up-sampled and channel-stitched, which increases the detail retention rate of microcracks (<50μm) to 89%, significantly better than the 72% of the traditional FPN structure. At the same time, its Anchor-free detection head design abandons the predefined anchor box constraints and directly predicts the bounding box coordinates through center point regression. The positioning error in the petal seam area (average width 0.3mm) is compressed from YOLOv7's ±15μm to ±7μm, meeting the detection accuracy requirements of sub-pixel defects. The core technical motivation for choosing YOLOv8 lies in its open modular architecture and industrial-grade deployment efficiency. The framework supports seamless docking with mainstream inference engines such as TensorRT and ONNX, and achieves a single-frame processing speed of 8ms on the NPU side through layer fusion and operator optimization, laying the foundation for the hardware collaborative optimization of MDConv and CPMS modules. In addition, its dynamic label assignment strategy (Task-Aligned Assigner) calculates the joint weight of classification score and regression IoU, which increases the positive sample matching accuracy to 91% (82% of the original model), effectively alleviating the training bias problem caused by the extreme imbalance of positive and negative samples in industrial scenarios (microcracks account for <5%). Experiments show that the improved model based on the YOLOv8 base framework pushes the detection accuracy boundary to 96.7% mAP@0.5 while maintaining 35FPS real-time performance, becoming the optimal technical carrier for intelligent detection of complex industrial equipment. Therefore, the present invention proposes a collaborative optimization detection method based on improved YOLOv8, which realizes high-precision, low-power industrial-grade real-time detection through algorithm-hardware joint design.

[0068] like Figure 2 As shown in FIG, the structure diagram of the improved YOLOv8 model of the method of the present invention is shown. The core improvements of the present invention include the following three parts:

[0069] 1. Backbone Reconstruction (Multi-scale Dynamic Deformable Convolution MDConv)

[0070] like Figure 3 As shown in the figure, Multi-scale Dynamic Deformable Convolution (MDConv) is an innovative convolutional structure for industrial defect detection. Its core achieves accurate capture of cross-scale defects in complex scenarios by fusing the geometric adaptability of deformable convolution with the multi-scale characteristics of dilated convolution. This module adopts a dynamic parameter generation mechanism: the dilation rate (d) of the convolution kernel is adaptively adjusted (range 1-3) by a lightweight sub-network based on the input features, and a spatial offset (Δp) is generated through a single-layer convolution, so that a single convolution kernel can cover a multi-scale receptive field from 1×1 to 3×3. This design enables the model to focus on the local texture of sub-pixel microcracks (such as 50μm cracks) while perceiving the global semantic features of millimeter-scale expansion defects.

[0071] At the hardware optimization level, MDConv reduces parameter usage by 23% through a dynamic parameter sharing mechanism (such as the joint generation of expansion rate and offset), and designs integer shift operations specifically for mobile NPUs, reducing the latency of single convolution calculations to 8ms. Experiments show that this module improves the characteristic response strength of microcracks in a split-petal crucible inspection task by 42%, while also reducing the computational workload from 28.5GFLOPs to 7.0GFLOPs, meeting the dual requirements of real-time performance and low power consumption for industrial equipment.

[0072] Structural replacement: The original YOLOv8's C2f standard convolution unit is replaced with an independently designed MDConv module;

[0073] The mathematical expression of MDConv is:

[0074] Among them, p o Indicates the current position of the output feature map, p k represents the fixed position offset of the k-th branch, w k represents the weight parameter of the kth branch, x represents the input function, and the dynamic expansion rate d k Generated by global average pooling and lightweight convolution, offset Δp k Generated using single-layer convolution and Tanh activation function;

[0075] The mathematical expression of its dynamic expansion rate adjustment mechanism is:

[0076] d k =Softmax(Wd ·GAP(x))·D max

[0077] Among them D max is the maximum expansion rate, W d A learnable parameter matrix is ​​used to achieve adaptive matching of multi-scale receptive fields from 11×11 to 33×33, and improve sub-pixel crack detection capabilities.

[0078] Hardware adaptation: By adopting a parameter sharing strategy, the number of parameters in the MDConv module is reduced by 32% compared to the original C2f (from 3.7M to 2.5M), and integer shift operations are used instead of floating-point calculations on the NPU side, which increases the inference speed by 41%.

[0079] 2. Neck Enhancement (Channel-Spatial Hierarchical Attention CPMS)

[0080] like Figure 4 As shown in Figure 2, Channel-Spatial Hierarchical Attention (CPMS) is a dual-path attention mechanism designed specifically for complex industrial scenarios. It significantly improves the model's ability to perceive small defects and suppress background interference by synergizing multi-scale frequency domain analysis in the channel dimension with geometric prior guidance in the spatial dimension. Its core design consists of two parts: a channel attention path and a spatial attention path, which achieve end-to-end optimization through feature fusion.

[0081] In the channel attention path, CPMS innovatively introduces a multi-step pyramid compression excitation structure, which extracts feature information in different frequency domains through pooling operations of steps 1, 2, and 3. The global average pooling of step 1 captures the overall semantic features of the image, such as the macroscopic direction of millimeter-level cracks; the local pooling of step 2 focuses on the mid-frequency texture features, which is used to identify medium-scale defects such as sub-surface micropores; the high-density pooling of step 3 strengthens the high-frequency edge details by calculating the standard deviation of the local area, effectively enhancing the gradient mutation response of 50μm-level microcracks. The outputs of these three branches are compressed by independent fully connected layers, and the channel weight vectors are generated by the Sigmoid function, and finally the weighted fusion forms the channel attention map. Experiments show that this multi-scale frequency domain fusion strategy increases the channel activation strength of microcracks by 37%, while reducing the number of parameters by 33% compared with the traditional SE module.

[0082] The spatial attention path deeply integrates the structural prior knowledge of the equipment and generates a geometric mask (Gmask) based on the 3D CAD model of the petal-type component. By parsing the surface parameterized data in IGES format, the 3D seam is projected onto the 2D detection plane, and a morphological dilation algorithm (kernel size 0.3mm) is used to construct a safety buffer to cover the ±0.1mm position deviation that may exist in actual assembly. During the attention calculation, the spatial attention weight of the feature map is processed through a 3×3 convolution layer to process the average-max pooling splicing result, and then a matrix dot multiplication operation is performed with Gmask to directly suppress invalid feature activation in the seam area. In the petal-type crucible detection task, this mechanism successfully compressed the false detection rate of the seam area from 12.7% to 1.5%, especially in the cooling pipe shadow interference scenario, the false positive samples were reduced by 89%.

[0083] To adapt to the demanding deployment environments of industrial equipment, CPMS has undergone in-depth hardware optimization. Channel attention utilizes a parameter sharing strategy, with the three-layer MLP network reusing the first-layer weights, reducing the number of parameters from 768KB in the traditional design to 512KB. Spatial masks are accelerated through binary storage (0 / 1 matrices) and bitwise operations, enabling high-speed filtering using AND / OR instructions in the NPU, reducing the computational load per frame by 40%. Combined with TensorRT INT8 quantization technology, the model size is compressed to 6.7MB, achieving 8ms single-frame inference on the CSK6-MIXAI chip, maintaining a stable power consumption of 1.2W, and supporting operation over a wide temperature range of -20°C to 60°C. Actual industrial verification has shown that this module has increased the detection rate of 30μm microcracks to 98.2% in semiconductor single crystal furnace inspections, reducing annual maintenance costs by 72%, and providing a reliable intelligent quality inspection solution for high-end equipment manufacturing.

[0084] Module embedding: The CPMS module is embedded in the PANet feature fusion path. Its dual-path working mechanism is as follows:

[0085] Channel optimization: Multi-scale pooling with a step size of 1 / 2 / 3 is used to extract high-frequency defect features, reducing the number of parameters by 33% compared to traditional SE modules;

[0086] Spatial suppression: Generate geometric mask G based on the petal seam CAD model mask , 40% invalid feature filtering is achieved through bit operations, and the mathematical expression is:

[0087]

[0088] Among them, F in represents the input features, F out Represents the output features.

[0089] Interference suppression: The false detection rate in the flap seam area (average width 0.3mm) is reduced from 12.7% to 1.5%.

[0090] 3. End-to-end optimization strategy

[0091] Training in stages:

[0092] 1) Frozen pre-training: Initially, only the MDConv and CPMS modules are trained for 100 epochs with a learning rate of 0.001 (Cosine decay).

[0093] 2) Joint fine-tuning: Unfreeze all parameters for end-to-end optimization, introducing the EMA model sliding average and RandAugment dynamic enhancement;

[0094] Loss function design:

[0095] L total =1.5L cls +0.8L box +1.2L mask

[0096] Among them L mask =1-IoU(S att ,G mask ) is the mask alignment loss, which constrains the consistency of spatial attention and geometric prior, L cls represents the classification loss, L box represents the position regression box loss.

[0097] like Figure 5 As shown in FIG. , this is a crack detection effect diagram of the method of the present invention, and the present invention obtains the following results:

[0098] Accuracy Verification: On the split crucible test set, mAP@0.5 reached 96.7%, and the recall rate of 50μm microcracks was improved by 42% (compared to traditional methods);

[0099] False detection rate test: In scenarios with seam texture interference, the false detection rate is stable below 1.5%, meeting the ASTM E2478 industry standard.

[0100] Power consumption and latency: During 8 hours of continuous operation, the NPU consumed an average of 1.2W of power, and single-frame detection took 8ms±0.3ms, with no performance degradation.

[0101] The foregoing is merely a detailed description of specific embodiments of the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to be accorded the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A metal surface defect detection method based on multi-scale dynamic convolution and spatial attention, characterized in that: The following steps are involved: S1, collect metal surface images, pre-process and create a dataset for training; S2 combines the multi-scale dynamic deformable convolution module MDConv and the channel-spatial layered attention module CPMS with geometric masks to improve the YOLOv8 network structure and loss function. The improved network is iteratively trained using dataset images, and the network parameters are optimized to obtain an ideal model for metal surface defect detection. S3. Actual metal surface images are collected and pre-processed before being input into the ideal model, and the defect and defect location regression boxes are automatically input.

2. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 1 is characterized in that: The pretreatment includes: Generate geometric mask G based on the petal structure CAD model mask , mark the coordinate set of the seam area; The standard deviation and kernel size are set, Gaussian noise and motion blur are randomly added, and images simulating interference in industrial scenes are generated.

3. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 2 is characterized in that: The metal surface images include defects of various crack widths and images without crack defects, and are labeled using binary classification.

4. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 1 is characterized in that: The Cn2f module of the YOLOv8 network structure is improved into a multi-scale dynamic deformable convolution module MDConv. By integrating the deformable convolution kernel with the dynamic expansion rate adjustment mechanism of the void convolution, it achieves adaptive matching of multi-scale receptive fields from 11×11 to 33×33. The mathematical expression of MDConv is: Among them, the dynamic expansion rate d k It is generated through global average pooling and lightweight convolution, and the offset Δp is generated using a single-layer convolution and Tanh activation function.

5. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 1 is characterized in that: The channel-spatial hierarchical attention module (CPMS) is embedded in the PANet feature fusion path of the YOLOv8 network structure, and a multi-step pyramid compression-stimulated channel attention and geometric mask-guided spatial attention collaborative working mechanism are adopted. The operations of the CPMS module include: Features of different scales are extracted through multi-scale pooling layers with unequal step sizes and Sigmoid gating mechanism; The geometric mask G generated based on the petal structure CAD model is used mask Fusion layer with spatial attention realizes invalid feature filtering through matrix dot multiplication to suppress false detection of seam texture.

6. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 1 is characterized in that: The method adopts a phased training strategy: in the initial stage, the original YOLOv8 parameters are frozen and only the MDConv and CPMS modules are trained; The fine-tuning phase uses cosine annealing learning rate for end-to-end optimization.

7. The metal surface defect detection method based on multi-scale dynamic convolution and spatial attention according to claim 1 is characterized in that: In the loss function, the classification loss weight is increased to 1.5 times, and the mask IoU loss constraint is added; The expression is: L mask =1-IoU(G mask ,S att ) Among them, G mask Geometric mask, S att is the spatial attention output.