Metal defect identification method based on YOLOv8

By improving the YOLOv8 network structure and combining multi-angle image acquisition, data enhancement and feature fusion modules, the accuracy and robustness issues of metal defect detection in complex industrial scenarios are solved, and efficient and accurate metal defect recognition is achieved.

CN120612318APending Publication Date: 2025-09-09SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510759251.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies for metal defect detection in complex industrial scenarios suffer from insufficient detection accuracy, high false detection rate, and high computational complexity, making it difficult to meet the high efficiency and high accuracy requirements of modern industrial production.

Method used

Improve the YOLOv8 network structure and enhance feature extraction and recognition capabilities through multi-angle image acquisition, data enhancement, feature fusion module improvement, graph convolutional network and attention mechanism enhancement.

Benefits of technology

It significantly improves the accuracy and robustness of metal defect recognition, can efficiently detect metal defects in complex industrial scenarios, and meet the real-time requirements of industrial detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612318A_ABST
    Figure CN120612318A_ABST
Patent Text Reader

Abstract

The invention discloses a metal defect identification method based on YOLOv8, and the method comprises the steps: carrying out the semi-automatic labeling of a collected sample data set through a LabelMe semi-automatic labeling tool, and carrying out the detection feedback through an existing YOLOv8 model. Carrying out manual labeling correction on a labeling sample with a relatively poor detection result through a LabelImage labeling tool; replacing a feature fusion part in an original YOLOv8 network with a combination of a weighted splicing operation, a feature enhancement layer and a gating mechanism based on a bidirectional feature pyramid network idea, and performing adaptive fusion on features of different scales; inserting a GCNECA module into the last layer of the backbone network of the original YOLOV8 network; the method comprises the following steps of: inserting a TransformerSE module behind a P3 layer in a neck network layer of an original YOLOV8 network; a MultiScaleLSTMSimAM module is inserted behind a P5 layer in a neck network layer of an original YOLOV8 network, so that the detection effect of large target metal defects is enhanced. According to the method, the metal defect identification capability in a natural scene is enhanced, and higher robustness, accuracy and efficiency are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial quality inspection technology, and in particular to a metal defect recognition method based on YOLOv8, which is particularly suitable for automatic detection and recognition of metal defects in complex industrial scenarios. Background Art

[0002] Metal materials, as essential building blocks for industrial production and construction, have a quality that directly impacts product safety and service life. However, during the production, processing, and use of metal materials, various defects, such as cracks, pores, and inclusions, inevitably arise. If these defects are not discovered and addressed promptly, they can lead to serious safety incidents and economic losses. Therefore, the identification and detection of metal defects has become a critical step in industrial quality control.

[0003] Early metal defect detection mainly relied on machine vision and image processing technologies, such as the new surface defect detection method based on convolutional neural networks proposed by the Yun team (Yun JP, Shin WC, Koo G, et al. Automated defect inspection system for metal surfaces based on deep learning and data augmentation [J]. Journal of Manufacturing Systems, 2020, 55: 317-324.). This method trains deep networks of different depths and node layers to directly obtain feature vectors from the convolutional layers, thereby achieving automatic visual inspection and identifying defects such as dirt, scratches, burrs, and wear on the surface. However, when applied to a variety of different surfaces, insufficient sample data or increased image size can lead to overfitting problems, while increased image size increases computational costs, thereby affecting the reliability of detection. Detection methods based on traditional classifiers, such as SVM and AdaBoost, are prone to misjudgment in complex textured backgrounds. Zhang (Xue-Wu Z, Yan-Qiong D, Yan-Yun L, et al. A vision inspection system for the surface defects of strongly reflected metal based on multi-class SVM. Expert Systems with Applications, 2011, 38(5): 5930-5939.) and his team developed a visual inspection system based on a multi-class support vector machine to detect defects on strongly reflective metal surfaces. However, the SVM parameters obtained through cross-validation of this method are only optimized under specific experimental conditions. The classification accuracy is significantly reduced when faced with dynamic interference such as production line lighting fluctuations and changes in metal reflective intensity. Furthermore, the multi-stage serial processing architecture leads to cumulative computational delays, making it difficult to meet the millisecond-level real-time requirements of modern industrial inspection.

[0004] With the breakthrough progress of deep learning, the solution based on single-stage target detection has significantly improved the detection efficiency. As a representative method in the field of target detection, the YOLO series of algorithms has attracted widespread attention for its high efficiency and accuracy. Based on the improved YOLOv5 model, Zhao Hui's team conducted experimental research on metal surface defect detection, aiming to solve the problem that traditional methods are difficult to effectively detect non-standardized complex defects. (Zhao Hui, Chen Zhifeng, Zhang Jiawei, et al. Experimental research on metal surface defect detection based on the improved YOLOv5 model. Experimental Technology and Management, 2025, 42(01): 66-74. DOI: 10.16791 / j.cnki.sjg.2025.01.009.) It is worth noting that the new generation YOLOv8 algorithm has achieved significant breakthroughs in the following dimensions: higher detection accuracy, stronger multi-scale feature fusion capabilities, and significantly improved robustness for small targets and complex scenes compared to YOLOv5. At the same time, it achieves higher reasoning efficiency and computing resource utilization through deep optimization of lightweight network structure.

[0005] In industrial production environments, metal defect identification faces numerous challenges. First, defects come in a variety of shapes, including point defects, linear defects, and surface defects, and their scales vary widely. Second, metal surfaces often have complex textures and reflective properties, which can easily interfere with defect detection. Furthermore, the complex environments of industrial sites may present problems such as uneven lighting and dust interference, further increasing the difficulty of detection. Traditional detection methods, which primarily rely on manual visual inspection or rule-based image processing techniques, are not only inefficient but also susceptible to subjective factors, making them difficult to meet the needs of modern industrial production.

[0006] The YOLOv8 model, by incorporating advanced deep learning techniques, provides a solution to these problems. Its improved backbone network and head structure more effectively extract metal surface feature information while improving detection accuracy. To address the practical needs of metal defect detection, researchers can make targeted improvements based on YOLOv8, such as adding an attention mechanism and optimizing feature fusion strategies, to better address metal defect recognition tasks in complex industrial environments.

[0007] However, metal defect recognition technology still faces several challenges. For example, detection accuracy for defects with extremely irregular shapes or those that are highly similar to the background needs to be improved. The model's generalization capabilities under extreme environmental conditions need to be further enhanced. Furthermore, reducing the model's computational complexity to enable deployment on resource-constrained embedded devices is another area of ​​research that requires significant attention.

[0008] In summary, while deep learning technology has made significant progress in metal surface defect detection, its practical application in industrial scenarios still faces the following core challenges: First, insufficient feature recognition accuracy for small defects and complex backgrounds; second, a high false detection rate caused by optical interference from highly reflective surfaces. Therefore, improvements are needed to better meet the demands for high-efficiency and high-accuracy metal surface defect detection, bringing more efficient and intelligent operations and management to industrial production, and promoting the development and improvement of the entire industry. Summary of the Invention

[0009] The present invention aims to overcome the shortcomings of the existing technology and proposes a metal defect recognition method based on deep learning to improve the efficiency, robustness and accuracy of metal defect detection in complex industrial scenarios.

[0010] In order to achieve the above invention, the present invention provides a metal defect recognition method based on an improved YOLOv8 network, comprising the following steps:

[0011] (1) Collecting image samples of metal defects: Using a multi-angle industrial camera array, three cameras are arranged in a 120° ring to capture images of the metal surface. Polarization filters are used to suppress reflection interference, and an adaptive exposure algorithm is used to ensure that the defect features are clearly visible. OpenCV is used to normalize the images to a uniform size of 640×640, enhance them with CLAHE, and perform non-local mean denoising.

[0012] (2) Preprocess and label the images: Use the OpenCV library to resize the input images to a uniform size of 640 × 640. Use the LabelMe semi-automatic labeling tool to semi-automatically label the captured images. Use the existing YOLOv8 model to detect and provide feedback. For labeled samples with poor detection results, use the LabelImage labeling tool supplemented by manual labeling correction.

[0013] (3) Data augmentation-based training set creation: Mosaic-8 data augmentation, color perturbation, and brightness adjustment methods were used to create a training dataset. Eight images were grouped together and spliced ​​using random scaling, random cropping, and random arrangement. The RGB channel values ​​in the image were randomly changed, and a random value was added or subtracted to adjust the image brightness to simulate changes under different lighting conditions. Eight new images were obtained, and random occlusion processing was performed on the new images.

[0014] (4) Improve the feature fusion module in the YOLOv8 network: replace the feature fusion part in the original YOLOv8 network with a combination of weighted concat operation, feature enhancement layer and gating mechanism based on the BiFPN idea. This module adaptively fuses features of different scales through learnable weights, and introduces a feature enhancement layer to improve the expressiveness of features. At the same time, a gating mechanism is used to dynamically adjust the ratio of original features and enhanced features to achieve more effective feature fusion. This improvement helps to improve the network's detection ability for metal defects, especially for the extraction and fusion of defect features of different scales and complexities. This module consists of a weight learning unit, a feature concat operation, a convolution-batch normalization-activation feature enhancement layer and a gating unit based on the sigmoid function, which can improve the model's recognition accuracy for metal defects while maintaining computational efficiency;

[0015] (5) Enhance the basic feature extraction capability of the YOLOv8 network: Insert the GCNECA module into the last layer of the backbone of the YOLOv8 network to enhance the basic feature extraction capability. This module integrates multi-scale feature extraction, graph convolutional network (GCN) and channel attention mechanism, which can greatly improve the efficiency and accuracy of the model in extracting metal defect features. Multi-scale convolution captures feature information of different scales through convolution kernels of different sizes, which enhances the adaptability of the model to defects of various sizes. The graph convolutional network layer effectively models the non-local relationship between features by constructing the adjacency matrix of the feature graph, and improves the ability to understand complex defect morphology. Among them, the GCN module uses the graph structure to model features and captures the long-range dependency between features through the message passing mechanism, which enhances the model's perception of the overall structure of the metal surface. The channel attention mechanism adaptively adjusts the weight of each channel feature to highlight key information. These improvements work together to improve the backbone network's ability to extract metal defect features, laying a solid foundation for subsequent detection tasks.

[0016] (6) Improve the detection effect of the YOLOv8 network on small target metal defects: Insert the TransformerSE module after the P3 layer in the neck layer of the YOLOv8 network to enhance the detection effect of small target metal defects. The TransformerSE module contains spatial attention, channel attention and multi-head self-attention mechanisms. It generates spatial and channel attention features respectively through adaptive average pooling, and combines the multi-head self-attention layer for feature enhancement. This module is introduced after the P3 layer and can effectively capture the detailed information of small-scale metal defects. The feedforward network and layer normalization in the module will further process the features and improve the model's ability to recognize complex metal defects. The attention weight is generated through the Sigmoid activation function and applied to the input features to achieve dynamic feature weighting. This improvement not only enhances the Neck layer's ability to extract and fuse small target features, but also improves the model's recognition accuracy for metal defects of different scales, while also maintaining computational efficiency;

[0017] (7) Improve the detection effect of YOLOv8 network on large target metal defects: The MultiScaleLSTMSimAM module is inserted after the P5 layer in the neck layer of the YOLOv8 network to enhance the detection effect of large target metal defects. This module integrates the SimAM attention mechanism, LSTM temporal modeling and multi-scale feature extraction, which can significantly improve the model's recognition ability for large-scale metal defects. Among them, the SimAM attention mechanism adaptively adjusts the feature weight by calculating the feature map, and the LSTM layer captures the long-range dependency of the features and effectively models the context information of large targets. At the same time, multi-scale convolution (1x1, 3x3, 5x5) is used to extract features of different receptive fields, which enhances the model's adaptability to defects of different sizes. These features are integrated through the fusion convolution layer and combined with the original input through the residual connection, which not only retains the original information but also enhances the feature expression. It not only improves the extraction and fusion ability of the Neck layer for large target features, but also enhances the recognition accuracy of complex morphological metal defects through the LSTM temporal modeling.

[0018] By improving the YOLOv8 network structure, the present invention significantly improves the accuracy and robustness of metal defect recognition, providing an efficient technical solution for solving defect detection problems in complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of an embodiment of the present invention;

[0020] Figure 2 An improved feature fusion module according to an embodiment of the present invention;

[0021] Figure 3 It is a graph convolution enhanced channel attention module of an embodiment of the present invention;

[0022] Figure 4 The spatial-channel self-attention squeeze excitation module of an embodiment of the present invention;

[0023] Figure 5 This is a multi-scale LSTM attention module of an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The technical approach of the present invention is further described in detail below through specific embodiments and drawings.

[0025] like Figures 1 to 5 As shown, the present invention is based on an embodiment of a metal defect recognition method based on YOLOv8, comprising the following steps:

[0026] 1. Create a dataset for metal defect recognition

[0027] A multi-angle industrial camera array was used, with a polarizing filter to suppress specular reflections. An adaptive exposure algorithm was used to ensure clear visibility of defect features. OpenCV was used for image normalization to a uniform 640×640 resolution, and CLAHE contrast enhancement and non-local means denoising were performed. A total of 1,630 images were collected, including 1,430 for training and 200 for validation.

[0028] 2. Perform necessary preprocessing operations on the image

[0029] The input image is resized to 640×640 using the OpenCV library. The collected images are semi-automatically annotated using the LabelMe semi-automatic annotation tool, and the existing YOLOv8 model is used to detect and feedback. For labeled samples with poor detection results, the LabelImage annotation tool is used with manual annotation correction.

[0030] 3. Add an improved feature fusion module to replace the Concat module in the YOLOv8 network

[0031] In order to effectively overcome the problems of missed detection and false detection caused by complex background and small target in metal defect detection, this patent proposes to add an improved feature fusion module to replace the Concat module in the YOLOv8 network. Specifically, Figure 2As shown in the figure, the Softplus function is applied. The Softplus function is a smooth activation function, which ensures that the weight update is more stable and there will be no problems of negative weight or gradient explosion. Compared with the traditional ReLU or Sigmoid, this function avoids the "dead neuron" problem and the weight update is smoother; a feature enhancement layer is added to perform deeper processing of the input features. The convolution operation can improve the expressiveness of the features, which helps to enhance the representativeness and robustness of the features; a gating mechanism layer is added to dynamically adjust the fusion ratio of the enhanced features and the original features, thereby enhancing the adaptability of the model and allowing the model to dynamically select which features to use as needed.

[0032] Specifically, the randomly initialized weight parameters are transformed nonlinearly through the Softplus function. The specific formula is as follows:

[0033] (Softplus function)

[0034] Among them, ω0, ω1 are randomly initialized weight parameters, ω0,

[0035] The advantage of this function over the traditional ReLU is that it is omnidirectional, which can effectively alleviate the gradient vanishing problem while ensuring that the weight is always positive (mathematical constraints: f(ω i )≥0).

[0036] Then the normalization operation is performed on the transformed weights. The specific formula is as follows:

[0037]

[0038] where ε = 10 -4 It is a numerical stabilization term to prevent numerical instability caused by the denominator being zero.

[0039] This step solves the problem of difficulty in quantifying the importance of features at different levels when fusing multiple features by establishing a competition mechanism between weights.

[0040] The input feature maps x1(B, C1, H, W) and x2(B, C2, H, W) are respectively multiplied by the normalized weights. This process essentially weights the spatial attention of features at different levels, strengthening the response strength of important feature areas. The weighted features are concatenated (concat) along the channel dimension to form a fused feature. The specific formula is as follows:

[0041]

[0042]

[0043] in is the input feature map, B is the batch size (BatchSize), C1, C2: the number of input channels, H×W: feature map resolution, ⊙ is element-by-element multiplication.

[0044] The feature enhancement stage uses a 3x3 convolution kernel (parameters: kernel_size = 3, padding = 1) to perform a two-dimensional convolution operation. Its core function is to establish spatial context association by expanding the local receptive field, enhancing the feature expression ability of small-scale targets. Then, the feature distribution is standardized through the batch normalization layer (BatchNorm). This operation improves the training convergence speed by stabilizing the statistical distribution of the activation values ​​of the intermediate layers. Finally, the ReLU activation function is applied to achieve feature sparsity while preserving feature nonlinearity and suppressing invalid feature interference. The specific formula is as follows:

[0045] x conv =W conv ·x cat

[0046]

[0047] x enhance =max(0,x bn )

[0048] is the 3×3 convolution kernel parameter (C out is the number of output channels), × is a two-dimensional convolution operation (Stride = 1, Padding = 1), μ, σ are the mean and standard deviation of the feature map of the current batch. γ, are learnable scaling and translation parameters.

[0049] Entering the gating mechanism layer, this layer adopts a dual-path design, the original splicing feature x cat After the channel is compressed through a 1×1 convolution, a sigmoid function is used to generate gating coefficients, and feature reconstruction is achieved through a dynamic fusion formula. This layer solves the problem of rigid superposition of enhanced features and original features in traditional feature fusion methods, enabling the model to adaptively adjust the enhancement strength based on the current input characteristics. The specific formula is as follows:

[0050] gate=σ(W gate ·x cat )

[0051] output=gate⊙x enhanced +(1-gate)⊙x cat

[0052] in The 1×1 convolution kernel channel is compressed to 1, σ(x) is the Sigmoid function, gate⊙x enhanced is the enhanced defect feature (such as crack edge). (1-gate)⊙x cat To preserve the original texture features (such as metal base).

[0053] 4. Insert the GCNECA module into the last layer of the YOLOv8 network backbone network

[0054] To address the problems of insufficient multi-scale feature extraction, lack of long-range dependency modeling between channels, and weakened feature response under noise interference in complex industrial scenarios for tiny defects on metal surfaces (such as micron-level cracks and submillimeter-level pores), this patent proposes inserting a GCNECA module at the end of the YOLOv8 backbone network. Specifically, Figure 3 As shown in the figure, multi-scale convolution kernels are used instead of fixed-size convolution kernels to capture features of different scales and enhance the model's sensitivity to different feature scales; a graph convolutional network (GCN) layer is added to more effectively capture the dependencies between channels through the graph structure, thereby enhancing the interaction between channels; a feature weighting module is added to the model through a combination of a fully connected layer and a Sigmoid activation function to perform weighted processing on the input features to adjust the contribution of each feature.

[0055] Specifically, the input feature map X∈R B×C×H×W , the spatial dimensions are compressed to (B, C, 1, 1) through the Adaptive Average Pooling layer. This operation eliminates spatial redundancy in the feature map through global averaging, retaining the global statistical characteristics of each channel and resolving the global information loss problem caused by local receptive fields in traditional attention mechanisms. A reshape operation is then performed to convert it into a two-dimensional vector of (b, c), providing an adaptive data structure for subsequent channel-level modeling. The specific formula is as follows:

[0056]

[0057] Where X∈R B×C×H×W F is the input feature map, the dimension is batch b × number of channels c × height h × width w. pool ∈R b×c×h×w After pooling, each channel is compressed to a global mean of 1×1.

[0058] Entering the multi-scale convolution branch, this module performs multi-granular feature extraction on the reshaped features by deploying three convolution kernels of different sizes in parallel (when the base kernel size is set to k, kernel sizes of (k+4, k+2, and k) are used respectively). The large kernel focuses on long-range channel correlations across regions, such as large areas of corrosion on metal sheets. The medium kernel captures medium-range structures, such as the extension direction of scratches. The k×k small kernel focuses on local detailed features, such as the edge texture of microcracks. The output of the multi-scale convolution is fused into an enhanced representation through feature concatenation. This design breaks through the limitations of the single-scale convolution of the traditional ECA module and significantly improves the model's perception of tiny objects (which rely on the details of the small kernel) and complex textures (which rely on the global relationships of the large kernel). The specific formula is as follows:

[0059] Kernel size = (k+4)×(k+4)

[0060] Kernel size = (k+2)×(k+2)

[0061] Kernel size = k × k

[0062]

[0063] Among them F resh ape ∈R B×C Conv is the two-dimensional matrix of the feature after pooling and reshaping. k+Δ is a convolution operation with a kernel size of (k+Δ)×(k+Δ) (Δ=4, 2, 0, is the output feature of convolution kernels of different sizes, F concat ∈R B×C It is the multi-scale feature after splicing along the channel dimension.

[0064] Subsequently, the graph convolutional layer (GCN) performs channel topology modeling on the concatenated features: by introducing the channel reduction ratio r, the feature dimension is reduced to This layer abstracts channels as nodes in a graph structure and uses an adjacency matrix to model non-local dependencies between channels. This addresses the drawback of traditional ECA modules that only capture the association between adjacent channels through one-dimensional convolution. It particularly enhances the ability to capture nonlinear interactions between distant channels (e.g., semantically complementary but spatially separated channels). The specific formula is as follows:

[0065] F gen =σ(W g ·F concat +b g )

[0066] Where W g ∈R 3c×3c / r is the graph convolution weight matrix, r is the channel reduction ratio, bg ∈R 3c / r is the bias term, σ is the nonlinear activation function ReLU, F gcn ∈R b×3c / r is the graph convolution feature after dimension reduction.

[0067] The reduced-dimensional features are restored to the original number of channels (b, c) through the fully connected layer (FC), reconstructing the complete topological structure of the channel and learning the complex mapping relationship across channels through the weight matrix. The fully connected layer and graph convolution here form a lightweight "dimensionality reduction-dimensionality increase" structure, achieving high-order feature transformation while controlling the number of parameters. Then, the Sigmoid function compresses the eigenvalues ​​to the range [0,1] to generate the channel attention weight matrix A∈R b×c×1×1 , the weight matrix quantifies the importance of each channel, where weights close to 1 represent key feature channels that need to be strengthened, and weights close to 0 correspond to redundant or noisy channels. The specific formula is as follows:

[0068] F fc =W f ·F gcn +b f

[0069] A=Sigmoid(F fc )

[0070] Where W f ∈R c×3c / r is the weight matrix of the fully connected layer, b f ∈R c is the bias term, F fc ∈R b×c , A∈R b×c×1×1 The attention weight matrix generated for Sigmoid.

[0071] Finally, the attention weights are restored to the dimension of (b, c, 1, 1) through a reshaping operation, and are multiplied channel-by-channel (Hadamard Product) with the original input feature map X to achieve dynamic calibration of features: the feature response of high-weight channels is enhanced, and the features of low-weight channels are suppressed. This gating mechanism realizes the dynamic fusion of enhanced features and original features, which not only retains the spatial detail integrity of the original features, but also improves the model's utilization of key information through attention-guided feature selection, especially improving the saliency expression of target features in occluded scenes. The output dimension of the entire processing flow remains (b, c, h, w), which is seamlessly connected to the YOLOv8 backbone network. Under the premise of almost no increase in computational cost, the detection model's robustness to scale changes, complex backgrounds and partially occluded targets is significantly improved through the collaborative mechanism of multi-scale perception, global graph structure modeling and dynamic feature calibration. The specific formula is as follows:

[0072]

[0073] X∈R b×c×h×w is the original input feature map, X∈R b×c×1×1 Attention weight matrix, Y∈R b×c×h×w The calibrated output feature map.

[0074] 5. Insert the TransformerSE module after the P3 layer in the neck layer of the YOLOv8 network

[0075] The local receptive field of traditional models is difficult to capture the cross-regional correlation of metal surface defects (such as the direction of crack extension and the spatial distribution of multi-point corrosion) and the local features of background noise (such as metal scratches and oxidation spots) and real defects (such as cracks and holes) are easily confused. This patent proposes to insert the TransformerSE module after the P3 layer in the neck layer of the YOLOv8 network. Specifically, Figure 4 As shown in the figure, spatial attention is combined with channel attention to enhance multi-dimensional features; a multi-head self-attention layer is used to improve the modeling ability of irregular defects on the metal surface (such as crack bifurcation) by allowing the model to learn the association of different subspace features in parallel, while suppressing gradient explosion through the temperature coefficient and stabilizing the training process; a feedforward network is added to compress the number of parameters, balancing computational efficiency and nonlinear expression capabilities while improving model generalization and preventing overfitting in industrial data.

[0076] Specifically, the input feature map X∈R B×C×H×W , dynamic feature enhancement is achieved through the following step-by-step processing. The spatial attention path compresses the global spatial information of each channel through adaptive average pooling. This operation eliminates redundant position noise through spatial dimension compression and focuses on the significant features of the channel dimension. At the same time, the channel attention path adopts adaptive average pooling to quantify the global dependencies between channels and solve the problem that the local receptive field of traditional convolution is insufficient to model long-range correlations. Then, the outputs of the two paths are reshaped separately through view operations (view), and dual attention fusion is achieved through tensor addition. The attention weights of the spatial and channel dimensions are complementarily superimposed to form a joint feature importance evaluation matrix. The specific formula is as follows:

[0077] S=AdaptiveAvgPool2d(1)(X),S∈R B×C×1×1

[0078]

[0079] C=AdaptiveAvgPool2d((h,w))(X),C∈R B×C×1×1

[0080]

[0081] S′=View(S), C′=View(C)

[0082] A=C′+S',A∈R B×C

[0083] A′=Unsqueeze(A),A′∈R 1×B×C

[0084] Where X is the input feature map, X∈R B×C×H×W , AdaptiveAvgPool2d(1) compresses the spatial dimension of each channel to 1×1, retaining the channel dimension, S is the spatial attention weight matrix, which represents the global spatial importance of each channel, AdaptiveAvgPool2d((h,w)) maintains the spatial dimension of the input, C is the channel attention weight matrix, which represents the global channel importance of each spatial position, View is the tensor dimension reshaping, compressing the four-dimensional tensor (b,c,1,1) into a two-dimensional matrix (b,c), and A is the joint attention matrix after fusion.

[0085] The fused feature A is dimensionally expanded to obtain A′ and input to the multi-head attention mechanism. By dividing the feature into four subspaces, the query matrix Q, key matrix K and value matrix V (all dimensions are R 1×B×C / 4 ), using scaled dot-product attention to implement cross-head interaction, and finally concatenating the multi-head outputs and aggregating them through a linear layer. This mechanism overcomes the locality limitations of convolution operations, explicitly modeling long-range spatial-channel dependencies between features, and particularly improving the ability to capture contextual associations for small and occluded objects. The specific formula is as follows:

[0086] A′ split =Split(A',dim=2,num_heads=4),A' Split ∈RB 1×B×4×C / 4

[0087] Q i =A′ split,i W Q ,K i =A′ split,i W K ,Q i =A′ split,i W Q ,V i =A′ split,i W V ,(i=1,2,3,4)

[0088] Attentioni ∈R 1×C / 44 /

[0089] M concat =Concat(Attention1,...,Attention4),Moncat∈R 1×B×C

[0090] M=M concat W o ,M∈R B×C ,W o ∈R C×C

[0091] Split is to split the channel dimension by the number of heads (when there are 4 heads, the number of channels becomes ), Q i , K i 、V i are the query, key, and value matrices of the i-th head, W o is the output projection matrix, W Q 、W K 、W V To be a learnable parameter, the input is mapped to the query, key, and value spaces respectively.

[0092] Next, the feedforward network (FFN) performs a nonlinear transformation on M: first, the dimensions are compressed using a linear layer, then a ReLU activation function is used to introduce nonlinear representation capabilities, and the original dimensions are restored using a linear layer with a dimension increase. Dropout (p = 0.1) is then used to randomly block some neurons to prevent overfitting. This process, through residual connections and layer normalization, stabilizes gradient propagation while preserving the original feature distribution, alleviating the vanishing gradient problem in deep network training. The specific formula is as follows:

[0093] F1=ReLU(MW1+b1),W1∈R C×C / 4 , b1∈R C / 4

[0094] F2=F1W1+b2,W2∈R C / 4×C ,b2∈R c

[0095] F res =Laryer Norm(M+Dropout(F2))

[0096] Among them, W1 and W2 are linear transformation matrices for dimensionality reduction and dimensionality increase, ReLU is a nonlinear activation function, Dropout is used to randomly block some neurons to prevent overfitting, and LayerNorm is layer normalization to stabilize feature distribution.

[0097] In the dynamic weight generation phase, F′ is input into the Sigmoid function to obtain W. This weight matrix is ​​multiplied element-by-element with the original input feature map X by expanding the dimension to form a gating mechanism. This allows the model to adaptively adjust the importance weights of each channel feature according to task requirements. At the same time, the original feature information is retained through residual connections to avoid the loss of key features caused by excessive screening of the attention mechanism. The specific formula is as follows:

[0098]

[0099]

[0100] W=σ(F res ),W∈[0,1] b×c

[0101]

[0102] The residual term To preserve the original feature information.

[0103] The entire process is enhanced through three levels: spatial-channel attention coordination, multi-head interaction modeling, and dynamic gate calibration.

[0104] 5. Insert the MultiScaleLSTMSimAM module after the P5 layer in the neck layer of the YOLOv8 network

[0105] Aiming at the insufficient fusion of multi-scale features in metal surface defect detection (such as the coexistence of cross-scale defects such as microcracks, scratches and pits), this patent proposes to insert the MultiScaleLSTMSimAM module after the P5 layer in the neck layer of the YOLOv8 network. Specifically, Figure 5 As shown in the figure, multi-scale convolutional layers are used to extract multi-scale features and rich features of the metal surface while combining information from different receptive fields to provide a more comprehensive feature representation; small residual links are used to ensure that the morphological details of micron-level defects are fully transmitted in the deep network, while alleviating the gradient vanishing problem; long-short-term memory networks are added to model the dynamic evolution laws of crack propagation trajectories, oxidation spot drift, etc. during the continuous rolling of galvanized sheets through the memory gate mechanism; the SimAM attention mechanism is added to automatically learn and emphasize important spatial features, suppress unimportant information, and drive the adaptive enhancement of feature channels through energy functions, effectively suppressing the reflection of aluminum alloy oxide layers and the interference of cold-rolled steel strip textures.

[0106] ] Specifically, the input feature map is X∈R B×C×H×W, feature extraction is performed through a multi-scale convolution layer. Through three parallel convolution paths, 1×1 convolution is used to reduce redundant interference from the metal texture background and focus on the core features of the defect. 3×3 convolution is used to extract edge details of defects such as local microcracks and scratches. 5×5 convolution is used to capture large-scale texture anomalies on the metal surface. The specific formula is as follows:

[0107]

[0108]

[0109]

[0110] Where W 1×1 ∈R C′×C×1×1 is the 1x1 convolution kernel weight, W 3×3 ∈R C′×C×3×3 is the 3x3 convolution kernel weight, W 5×5 ∈R C ′×C×5×5 is the 5x5 convolution kernel weight, C′ is the number of output channels r is the compression ratio, b i is the bias term (i=1,3,5), and x(i+u,j+v,c) is the value of the cth channel of the input feature map at position (i+u,j+v).

[0111] Cross-scale feature fusion combines the three convolutions and implements multi-scale feature interaction through 1×1 convolution with channel attention weighting, dynamically suppressing metal background noise (such as oxidation spots) and enhancing the response of defective areas. The specific formula is as follows:

[0112]

[0113]

[0114] a=Softmax(MLP(GAP(F cat )))

[0115] Among them, F cat is the concatenated feature map, the number of channels is C', MLP is a multi-layer perceptron, which generates attention weights, GAP is a global average pooling, and the feature map of each channel is compressed into a scalar.

[0116] The Long Short-Term Memory Network (LSTM) achieves time series modeling optimization through deep coupling of the gating mechanism and the motion characteristics of metal defects. The input gate dynamically screens the importance of the current feature, enhancing the morphological characteristics of oxidation plaques on the surface of stainless steel cold-rolled plates under strong reflective interference. The forget gate controls the decay rate of historical memory to avoid trajectory breakage caused by temporary occlusion of targets between frames in traditional methods. The output gate adjusts the output intensity of features to overcome the loss of edge information caused by motion blur. Its cell state update formula establishes an information transmission channel across time steps, constructing a continuous state space model of pit defects during the stamping process of aluminum alloy plates. The reshape operation restores the output of the LSTM to the spatial dimension and reconstructs the spatial topology of the feature map. The specific formula is as follows:

[0117] F seq =Flatten(F fusion )+P

[0118]

[0119] i t =σ(W xi F seq +W hi h t-1 +W ci ⊙c t-1 +b i )

[0120] f t =σ(W xf F seq +W hf h t-1 +W cf ⊙c t-1 +b f )

[0121] o t =σ(W xo F seq +W ho h t-1 +W co ⊙c t +b o )

[0122]

[0123] Where W xi , W xf , W xo is the weight of the input state. hi , W hf , W ho The weight of the hidden state W ci , W cf , W cois the connection weight, σ is the Sigmoid function, F seq Input feature sequence, i t is the output of the input gate. t is the output of the forget gate, c t is the current cell state, Candidate memory.

[0124] The SimAM attention mechanism first calculates the feature energy function, quantifies the importance of spatial position through statistical characteristics, and solves the deviation problem of manually designed attention parameters; the division operation normalizes the energy value to eliminate the influence of feature scale differences on weight distribution, and the Sigmoid activation function maps the energy to the attention weight, overcoming the morphological deviation problem caused by traditional manually set attention parameters. It is particularly suitable for defect localization in non-uniform backgrounds such as surface reflections of metal parts; it realizes the probabilistic expression of the [0,1] interval; the multiplication operation strengthens the feature response of high-energy areas and suppresses background noise. Its adaptive characteristics solve the feature confusion problem of traditional attention modules in complex scenarios. The residual 1×1 convolution layer preserves the integrity of the original X-ray or optical imaging features while learning spatial feature transformation through lightweight convolution kernels. It dynamically balances the strength of defect feature enhancement and background suppression with adjustable residual weights. Furthermore, the learnable residual weights are used to adjust the information fusion ratio, solving the problems of gradient vanishing and feature degradation in deep network training. Finally, the optimized feature map is output, forming a complete closed loop of collaborative optimization from the three dimensions of space, time, and attention, significantly improving detection accuracy and robustness in complex industrial environments. The specific formula is as follows:

[0125]

[0126]

[0127] y=F(x)+G(x)

[0128] Where μ is the mean of the feature map, N is the total number of feature map elements, and λ is a learnable parameter that adjusts the energy scaling. F(x) is the transformation of the main path, and G(x) is the identity mapping or 1×1 convolution of the residual path (used to adjust the number of channels).

[0129] Compared with the prior art, the present invention has the following beneficial effects:

[0130] (1) The present invention replaces the feature fusion part in the original YOLOv8 network with a combination of weighted splicing operation, feature enhancement layer and gating mechanism based on the idea of ​​bidirectional feature pyramid network. The module adaptively fuses features of different scales through learnable weights, and introduces feature enhancement layer to improve the expressiveness of features. At the same time, a gating mechanism is used to dynamically adjust the ratio of original features and enhanced features to achieve more effective feature fusion. This improvement helps to improve the network's detection ability for metal defects, especially for the extraction and fusion of defect features of different scales and complexities. The module consists of a weight learning unit, a feature splicing operation, a convolution-batch normalization-activation feature enhancement layer and a gating unit based on Sigmoid function, which can improve the model's recognition accuracy for metal defects while maintaining computational efficiency.

[0131] (2) The present invention inserts the last layer of the backbone network of the original YOLOV8 network into the GCNECA module to enhance the basic feature extraction capability. The module integrates multi-scale feature extraction, graph convolutional network and channel attention mechanism, which can greatly improve the model's extraction efficiency and accuracy of metal defect features. Multi-scale convolution captures feature information of different scales through convolution kernels of different sizes, enhancing the model's adaptability to defects of various sizes. The graph convolutional network layer effectively models the non-local relationship between features by constructing the adjacency matrix of the feature graph, thereby improving the ability to understand complex defect morphologies. Among them, the graph convolutional network module uses the graph structure to model features, captures the long-range dependency between features through the message passing mechanism, and enhances the model's perception of the overall structure of the metal surface. The channel attention mechanism adaptively adjusts the weight of each channel feature to highlight key information. These improvements work together to improve the backbone network's ability to extract metal defect features, laying a solid foundation for subsequent detection tasks.

[0132] (3) The present invention inserts the TransformerSE module after the P3 layer in the neck network (Neck) layer of the original YOLOV8 network to enhance the detection effect of small target metal defects. The TransformerSE module contains spatial attention, channel attention and multi-head self-attention mechanisms. It generates spatial and channel attention features respectively through adaptive average pooling, and combines the multi-head self-attention layer for feature enhancement. This module is introduced after the P3 layer and can effectively capture the detailed information of small-scale metal defects. The feedforward network and layer normalization in the module will further process the features and improve the model's recognition ability for complex metal defects. The attention weight is generated by the Sigmoid activation function and applied to the input features to realize dynamic feature weighting. This improvement not only enhances the neck network (Neck) layer's ability to extract and fuse small target features, but also improves the model's recognition accuracy for metal defects of different scales while maintaining computational efficiency.

[0133] (4) The present invention inserts the MultiScaleLSTMSimAM module after the P5 layer in the neck layer of the original YOLOV8 network to enhance the detection effect of large target metal defects. The module integrates the SimAM attention mechanism, long short-term memory network temporal modeling and multi-scale feature extraction, which can significantly improve the model's recognition ability for large-scale metal defects. Among them, the SimAM attention mechanism adaptively adjusts the feature weight by calculating the feature map, and the LSTM layer captures the long-range dependency of the features and effectively models the context information of large targets. At the same time, multi-scale convolution (1x1, 3x3, 5x5) is used to extract features of different receptive fields, which enhances the model's adaptability to defects of different sizes. These features are integrated through the fused convolution layer and combined with the original input through the residual connection, which not only retains the original information but also enhances the feature expression. It not only improves the neck network (Neck) layer's ability to extract and fuse large target features, but also enhances the recognition accuracy of complex morphological metal defects through the LSTM temporal modeling.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific implementation methods of the present invention can still be modified or replaced without departing from the spirit and scope of the present invention. Any modification or replacement should be included within the scope of protection of the claims of the present invention.

Claims

1. A metal defect recognition method based on YOLOv8, characterized in that The method significantly improves the recognition accuracy and adaptability of metal defects by introducing an improved feature fusion module, a GCNECA module, a TransformerSE module, and a MultiScaleLSTMSimAM module. These innovations not only optimize the feature extraction and fusion process, but also enhance the model's ability to recognize complex defects, achieving efficient and accurate real-time detection, and has broad application prospects. The method includes the following steps: (1) preparing a dataset for metal defect recognition; (2) performing necessary preprocessing operations on the image; (3) adding an improved feature fusion module to replace the Concat module in the YOLOv8 network; (4) inserting a GCNECA module into the last layer of the backbone network of the YOLOv8 network; (5) inserting a TransformerSE module after the third layer of the feature extraction network in the neck network layer of the YOLOv8 network; (6) inserting a MultiScaleLSTMSimAM module after the fifth layer of the feature extraction network in the neck network layer of the YOLOv8 network; the said preparation of a dataset for metal defect recognition refers to using an industrial camera array, suppressing mirror reflection interference through a polarizing filter, combining an adaptive exposure algorithm to ensure that the defect features are clearly visible, and using OpenCV for image normalization, unifying the size to 640×640, CLAHE contrast enhancement and non-local mean denoising; the said dataset is preprocessed, which refers to adjusting the size of the input image to 640×640 through the OpenCV library, semi-automatically annotating the collected image through the LabelMe semi-automatic annotation tool, and detecting feedback through the existing YOLOv8 model, and using the LabelImage annotation tool to supplement manual annotation correction for detection results with a confidence level lower than 0.

15.

2. The metal defect recognition method based on the YOLOv8 network according to claim 1 is characterized in that An improved feature fusion module is added to replace the Concat module in the YOLOv8 network. The improved feature fusion module consists of a weighted Concat operation layer based on the BiFPN idea, a feature enhancement layer and a gating mechanism layer. The weighted Concat operation layer based on the BiFPN idea refers to the learnable weight allocation of multi-scale input features, normalizing the weight values ​​through the Softmax normalization function, and then weighted splicing along the channel dimension; the feature enhancement layer refers to a nonlinear transformation module composed of a 1×1 convolution layer, a batch normalization layer and a ReLU activation function connected in sequence, which is used to enhance the expressive power of the fusion feature; the gating mechanism layer refers to generating dynamic gating weights through the Sigmoid function, and weightedly summing the weighted fusion features and the original input features according to the dynamic gating weights to achieve a balance between feature selection and information retention.

3. The metal defect recognition method based on the YOLOV8 network according to claim 1 is characterized in that A GCNECA module is inserted into the last layer of the backbone network of the YOLOV8 network. The GCNECA module consists of a multi-scale feature extraction layer, a graph convolutional network and a channel attention mechanism. The multi-scale extraction layer refers to extracting local features of different scales of the input feature map by deploying three two-dimensional convolution kernels in parallel, and the output features are spliced ​​through the channel dimension to form multi-scale fusion features; the graph convolutional network refers to inputting the spliced ​​multi-scale features into the graph convolution layer, reducing the dimension through the channel dimension through a reduction factor, constructing the graph structure adjacency relationship, and using the fully connected layer to model the non-local dependency between channels; the channel attention mechanism refers to adaptively averaging and compressing the spatial dimension of the graph convolution output features, restoring the number of channels through the fully connected layer, and generating a channel weight matrix through the Sigmoid function. Finally, the weight matrix is ​​multiplied by the original input features channel by channel to achieve key channel feature enhancement.

4. The metal defect recognition method based on the YOLOV8 network according to claim 1 is characterized in that A TransformerSE module is inserted after the third layer of the feature extraction network in the neck network layer of the YOLOv8 network. The TransformerSE module refers to a spatial-channel dual attention layer, a multi-head self-attention layer and a feedforward network connected in sequence. The spatial-channel dual attention layer refers to generating a spatial attention weight matrix and a channel attention weight vector respectively through adaptive average pooling, and adding and fusing the spatial attention with the channel attention after dimension expansion to generate a joint attention feature; the multi-head self-attention layer refers to dividing the input feature map into 4 parallel sub-features, each sub-feature generates a query, key, and value triple through an independent 1×1 convolution, and after calculating the spatial position similarity score matrix, generates an attention weight through scaled dot product and Softmax normalization. Finally, the weight and value matrix are weighted summed and spliced, and the multi-head output features are fused through a fully connected layer; The feedforward network is composed of two levels of linear transformation layers. The first layer compresses the number of channels and applies the ReLU function. The second layer restores the original number of channels. Random dropout regularization is used to prevent overfitting. Finally, the weight is generated by the Sigmoid function and multiplied with the input feature for output.

5. The metal defect recognition method based on the YOLOv8 network according to claim 1 is characterized in that A MultiScaleLSTMSimAM module is inserted after the fifth layer of the feature extraction network in the neck network layer of the YOLOv8 network. The MultiScaleLSTMSimAM module is composed of a multi-scale convolutional layer, a long short-term memory network and a SimAM attention mechanism connected in parallel. The multi-scale convolutional layer extracts local features by parallel 1×1, 3×3, and 5×5 convolution kernels; the long short-term memory network processes feature sequences in a bidirectional long short-term memory network to capture long-range spatial dependencies; the SimAM attention mechanism constructs a parameter-free attention weight calculation module based on neuroscience energy theory. For each spatial position of the input feature map, the energy function is used to calculate the feature similarity difference with the adjacent area to generate a three-dimensional weight matrix reflecting the feature significance.

Citation Information

Cited By

  • DSA-based left atrial appendage automatic segmentation and measurement method

    CN120852457A

  • Surface defect detection method and equipment based on DINOv3 model and medium

    CN121190466A

  • Water turbine coating abrasion detection method and system

    CN121353295A

  • Power transmission line engineering defect identification management system fused with artificial intelligence

    CN121392512A

  • Power transmission line engineering defect identification management system integrated with artificial intelligence

    CN121392512B