Small metal defect detection method and system

Through data enhancement, Multi Scale Fusion module and SE attention mechanism, the YOLOv8 model is improved, which solves the problem of insufficient edge information extraction in small metal defect detection, improves detection accuracy and robustness, and is suitable for complex defects and challenging scenarios.

CN120259188APending Publication Date: 2025-07-04ANHUI POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256073.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing YOLOv8 model has problems such as insufficient edge information extraction, loss of spatial information and insufficient feature fusion in small metal defect detection, resulting in insufficient detection accuracy and robustness.

Method used

The data set is expanded by data augmentation method, the Multi Scale Fusion module and SE attention mechanism are introduced to improve the YOLOv8 model, and pre-trained with ImageNet data set to optimize the model's edge information extraction and feature fusion.

Benefits of technology

It significantly improves the accuracy and robustness of small metal defect detection, enhances the model's detection ability in complex scenarios, reduces redundant information interference, and improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259188A_ABST
    Figure CN120259188A_ABST
Patent Text Reader

Abstract

The invention discloses a small metal defect detection method and system, and belongs to the field of target detection. The method comprises the following steps: S1, acquiring a small metal defect data set and preprocessing the small metal defect data set; step S2, deploying a YOLOv8 model and carrying out pre-training; step S3, for the backbone network of the YOLOv8 model after the pre-training is completed, introducing a Multi Scale Fusion module to carry out improvement; s4, dividing the preprocessed small metal defect data set into a training set, a verification set and a test set, training the improved YOLOv8 model by adopting the training set, and finely adjusting model parameters by adopting the verification set; and S5, after training and verification are completed, evaluating the improved YOLOv8 model by adopting a test set. According to the invention, the precision of small metal defect detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target detection. Specifically, the present invention relates to a small metal defect detection method and system. Background Art

[0002] In industrial production, especially for the task of small metal defect detection, the difficulty of collecting data sets is a prominent problem. The difficulty in collecting small metal data sets mainly stems from the visual difficulty of small-sized objects, the minuteness of defects, and the high-precision requirements for annotation. Since small metal parts are not easily recognizable in images and are easily affected by noise and illumination changes, it is difficult to capture the defect details in the images. In addition, metal surface defects are often very subtle and are easily overlooked or misjudged in low-quality images or complex backgrounds, which further increases the difficulty of data collection. These problems lead to the incompleteness and low quality of the data set, thus affecting the effect of subsequent model training.

[0003] As an advanced target detection model, YOLOv8 performs well in many scenarios. However, in small metal defect detection, there are often deficiencies in the extraction of edge information. First, the YOLOv8 model has limitations in edge information extraction, resulting in insufficient accuracy in the detection of defects with complex shapes and rich details. The edge information learning ability of existing models is weak and difficult to meet the high-precision requirements for fine defects. Second, although the downsampling operation in the existing YOLOv8 model can reduce the computational amount, it often leads to the loss of spatial information, which causes problems in the perception of local structures by the model. Especially in the detection of complex backgrounds and local defects, important spatial information is easily lost, thus affecting the detection accuracy and robustness. Finally, there are deficiencies in feature fusion in the YOLOv8 model, especially in fusing edge features and global features. It is unable to effectively combine edge information with global features, resulting in the model being unable to fully utilize the advantages of various features when processing complex images, thus affecting the detection performance.

[0004] Therefore, the present invention proposes a small metal defect detection method and system. Summary of the Invention

[0005] The present invention aims to overcome the deficiencies of the prior art and proposes a small metal defect detection method and system to achieve the following purpose: improving the accuracy of small metal defect detection.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: a small metal defect detection method, the method comprising the following steps:

[0007] Step S1, obtaining a small metal defect data set and performing preprocessing;

[0008] Step S2: Deploy the YOLOv8 model and perform pre-training;

[0009] Step S3: For the backbone network of the pre-trained YOLOv8 model, introduce the Multi Scale Fusion module for improvement;

[0010] Step S4: Divide the pre-processed small metal defect dataset into a training set, a validation set, and a test set, and use the training set to train the improved YOLOv8 model, and use the validation set to fine-tune the model parameters;

[0011] Step S5: After training and validation are completed, use the test set to evaluate the improved YOLOv8 model.

[0012] Preferably, in the step S1, the pre-processing includes augmenting the small metal defect dataset through data augmentation, and the data augmentation method includes geometric transformation and pixel transformation.

[0013] Preferably, in the step S2, the ImageNet dataset is used for pre-training the YOLOv8 model.

[0014] Preferably, in the step S3, all the original C2f modules in the YOLOv8 model backbone network are replaced with the Multi Scale Fusion module.

[0015] Preferably, the Multi Scale Fusion module is represented by the following formula:

[0016] For the input feature X ∈ R C×H×W , where C is the number of input channels, and H and W are the height and width respectively;

[0017] First, the input feature X is processed by the convolutional layer Conv to generate a new feature map X1, that is:

[0018]

[0019] Then, for the convolutional feature map X1, the Sobel operator is used to extract the image edge feature X sobel :

[0020] X sobel = SobelConv(X1);

[0021] Among them, SobelConv represents the Sobel convolution operation, which is divided into horizontal and vertical directions; the horizontal and vertical Sobel convolution kernels are initialized with fixed weights and are respectively used to calculate the horizontal gradient G x and the vertical gradient G y ;

[0022] Horizontal gradient G x and vertical gradient G y are calculated as follows:

[0023] G x = X1 * K x , G y = X1 * K y ;

[0024] where * represents the convolution operation; the horizontal Sobel kernel K x and the vertical Sobel kernel K y are defined as follows:

[0025]

[0026] Finally, the edge feature map X sobel is:

[0027] X sobel = G x + G y ;

[0028] In addition, the feature map X1 is also subjected to a max pooling operation to obtain the downsampled feature map X pool :

[0029] X pool = MaxPool(X1);

[0030] Then, the edge feature X sobel is concatenated with the pooled operation feature X pool in the channel dimension to obtain the concatenated feature map X concat :

[0031] X concat = concat(x sobel , x pool , dim = 1);

[0032] where dim = 1 indicates concatenation in the channel dimension;

[0033] The concatenated feature map X concat is subjected to a convolution operation to obtain a new feature map X2:

[0034] X2 = Conv2(X concat )

[0035] Finally, the intermediate feature map X2 is subjected to a convolution operation: X out = Conv3(X2); to obtain the final output feature map X out .

[0036] Preferably, in the step S3, a SE attention mechanism module is added after the SPPF layer of the backbone network of the YOLOv8 model.

[0037] Preferably, the SE attention mechanism module performs global average pooling on the input feature map, then models the inter-channel relationship through two fully connected layers, and finally generates channel weights through Sigmoid activation and performs weighted fusion with the input feature map, which is expressed by the formula as follows:

[0038] For the input feature map E ∈ R C×H×W perform global average pooling operation to obtain the global feature z of each channel:

[0039] z = avgpool(E);

[0040] where the result of the pooling operation is to obtain a C-dimensional vector z ∈ R C , and each element represents the average value of the corresponding channel:

[0041] Then, first compress the number of channels through a fully connected layer, then apply the ReLU activation function, then restore the number of channels through another fully connected layer, and finally apply the Sigmoid activation function to output the weight of each channel:

[0042] y = σ(FC2(ReLU(FC1(z))));

[0043] where FC1 and FC2 are the first and second fully connected layers respectively; ReLU represents the ReLU activation function; σ represents the Sigmoid activation function; y ∈ R C is the weight vector of the channel;

[0044] Finally, weight the vector y and the original input feature map E and splice them to obtain the weighted output:

[0045] E out = E × y;

[0046] where × represents element-wise multiplication.

[0047] Preferably, in the step S4, a cross-validation method is used to select the training set and the validation set from the dataset and then train and validate the model.

[0048] Preferably, in the step S5, the evaluation method includes calculating the mean average precision and the F1 score.

[0049] Meanwhile, the present application also proposes a small metal defect detection system, which includes a processor for executing a program constructed according to the above method.

[0050] The technical effects of the present invention are as follows:

[0051] (1) The present invention adopts data augmentation methods, including geometric transformation and pixel transformation, to expand the diversity of the dataset and ensure that the model can adapt to various changes.

[0052] (2) The present invention adopts a transfer learning strategy to pre-train the YOLOv8 model based on the ImageNet dataset, thereby reducing the dependence on a large number of metal datasets and improving the robustness of the model.

[0053] (3) The present invention introduces a Multi Scale Fusion module to improve the YOLOv8 model, significantly enhancing the edge information extraction effect and the detection ability for tiny and complex shape defects. At the same time, combined with the SE attention mechanism, which can adaptively adjust the weights of feature channels, enabling the model to focus on more important features, improving the detection accuracy in complex scenarios, reducing the interference of redundant information, and further enhancing the edge information extraction effect. Description of the Drawings

[0054] Figure 1 It is a flowchart of a small metal defect detection method provided by an embodiment of the present invention. Detailed Embodiments

[0055] The following further details the specific embodiments of the present invention by describing the embodiments with reference to the drawings, aiming to help those skilled in the art have a more complete, accurate, and in-depth understanding of the inventive concept and technical solutions of the present invention and facilitate its implementation. It should be noted that the terms "first", "second", etc. used in this application are only for conveniently describing the technical solutions to distinguish different components and do not limit this application. To make the technical solutions of the present invention clearer, the present invention is explained and illustrated through the following embodiments.

[0056] This embodiment provides a small metal defect detection method, as Figure 1 shown, the method includes the following steps:

[0057] Step S1, obtain a small metal defect dataset and perform preprocessing;

[0058] Step S2, deploy the YOLOv8 model and perform pre-training;

[0059] Step S3, for the backbone network of the pre-trained YOLOv8 model, introduce a Multi Scale Fusion module for improvement;

[0060] Step S4: Divide the preprocessed small metal defect dataset into a training set, a validation set, and a test set, and use the training set to train the improved YOLOv8 model, and use the validation set to fine-tune the model parameters;

[0061] Step S5: After training and validation are completed, use the test set to evaluate the improved YOLOv8 model.

[0062] Specifically, in step S1 of this embodiment, to solve the problems of insufficient data and poor quality caused by the difficulty in collecting the small metal defect dataset in the prior art, this embodiment preprocesses the obtained small metal defect dataset. The preprocessing not only includes data cleaning operations on the dataset to eliminate noise, but also includes augmenting the small metal defect dataset through data augmentation to expand the diversity of the dataset and ensure that the model can adapt to various changes. The data augmentation methods include (such as rotation, cropping, translation, scaling) and pixel transformation (such as color perturbation, brightness adjustment, etc.). These augmentation techniques can not only increase the diversity of the dataset, reduce the risk of overfitting, but also enhance the model's adaptability to different lighting, angle changes, and metal surface defects, thereby improving the detection accuracy.

[0063] The YOLOv8 model is an object detection model that supports image classification, object detection, and instance segmentation tasks. One of its feature points is that it replaces the C3 module in YOLOv5 with a C2f module and introduces an SPPF layer at the end of the backbone network, making the model further lightweight while having strong object detection capabilities and certain detection accuracy. Therefore, in step S2 of this embodiment, the YOLOv8 model is used as the basis for the entire small metal defect detection.

[0064] At the same time, even through data augmentation, the data volume of the small metal defect dataset is still insufficient for model training and cannot guarantee the detection accuracy of the trained model. Therefore, in step S2 of this embodiment, the ImageNet dataset is used for pre-training the YOLOv8 model, thereby reducing the dependence on a large number of metal datasets and improving the robustness of the model. As a standard dataset widely used in the field of computer vision, ImageNet has more than 14 million images. Such a large-scale dataset allows the model to learn extremely rich image features and patterns, thereby improving the generalization ability of the model and enabling it to perform well in the face of various different image tasks.

[0065] The pre-trained YOLOv8 model already has the preliminary ability to detect small metal defects. However, in the backbone network of the YOLOv8 model, although the C2f module can perform feature extraction, its performance in learning edge information and retaining spatial information is limited, especially in the detection accuracy of the subtle defects of small metals. To solve this problem, in step S3 of this embodiment, the original C2f modules in the backbone network of the YOLOv8 model are all replaced with Multi ScaleFusion modules to effectively improve the performance of the model, optimize the extraction of edge information and the effect of feature fusion, thereby significantly improving the accuracy of small metal defect detection.

[0066] Compared with the C2f module, the main goal of the Multi Scale Fusion module is to enhance the model's ability to capture image details by fusing edge features (Sobel convolution) and pooling features, and extract high-level features through multi-layer convolution operations. Specifically, the Multi Scale Fusion module is expressed by the following formula:

[0067] For the input feature X ∈ R C×H×W , where C is the number of input channels, and H and W are the height and width respectively;

[0068] First, the input feature X is processed by the convolutional layer Conv to generate a new feature map X1, that is:

[0069] X1 = Conv1(X) (X1 ∈ R C1×H1×W1 );

[0070] Then, for the convolutional feature map X1, the Sobel operator is used to extract the image edge feature X sobel :

[0071] X sobel = SobelConv(X1);

[0072] Among them, SobelConv represents the Sobel convolution operation, which is divided into horizontal and vertical directions; the horizontal and vertical Sobel convolution kernels are initialized with fixed weights and are used to calculate the horizontal gradient G x and the vertical gradient G y ;

[0073] The calculation formulas for the horizontal gradient G x and the vertical gradient G y are:

[0074] G x = X1 * K x , G y = X1 * K y ;

[0075] Among them, * represents the convolution operation; the horizontal Sobel kernel K x and the vertical Sobel kernel K y are respectively defined as:

[0076]

[0077] Finally, the edge feature map X sobel is:

[0078] X sobel = G x + G y ;

[0079] In addition, a max pooling operation is also performed on the feature map X1 to obtain the downsampled feature map X pool :

[0080] X pool = MaxPool(X1);

[0081] Then, the edge feature X sobel is concatenated with the pooled operation feature X pool in the channel dimension to obtain the concatenated feature map X concat :

[0082] X concat = concat(x sobel , x pool , dim = 1);

[0083] where dim = 1 indicates concatenation in the channel dimension;

[0084] A convolution operation is performed on the concatenated feature map X concat to obtain a new feature map X2:

[0085] X2 = Conv2(X concat )

[0086] Finally, a convolution operation is performed on the intermediate feature map X2: X out = Conv3(X2); to obtain the final output feature map X out .

[0087] The Multi Scale Fusion module enhances the network's ability to capture image details (especially edge information) by combining edge extraction using the Sobel operator and pooling feature extraction. Particularly in multi-scale feature extraction and detail capture, it improves the detection ability for tiny and complex shape defects. Additionally, considering that downsampling operations in existing technologies are prone to losing spatial information, which affects the perception of details and local structures, through multi-layer convolutional operations, the edge information and global features are integrated, effectively retaining the spatial information in the image, enhancing the model's sensitivity and detection accuracy for complex backgrounds and local defects, and finally generating more expressive features that can be used for subsequent detection tasks, improving the accuracy of target localization and boundary recognition.

[0088] Meanwhile, to further improve the accuracy of small metal defect detection, in step S3, this embodiment also adds a SE attention mechanism module after the SPPF layer of the YOLOv8 model backbone network.

[0089] The SE attention mechanism module adjusts the importance of different channels by introducing adaptive channel weights, thereby enhancing the feature representation ability. Specifically, the SE attention mechanism module performs global average pooling on the input feature map, then models the inter-channel relationship through two fully connected layers, and finally generates channel weights through Sigmoid activation and performs weighted fusion with the input feature map, which is expressed by the formula as follows:

[0090] Perform global average pooling operation on the input feature map E∈R C×H×W to obtain the global feature z of each channel:

[0091] z = avgpool(E);

[0092] where the result of the pooling operation is to obtain a C-dimensional vector z∈R C , and each element represents the average value of the corresponding channel;

[0093] Then, first compress the number of channels through a fully connected layer, then apply the ReLU activation function, then restore the number of channels through another fully connected layer, and finally apply the Sigmoid activation function to output the weight of each channel:

[0094] y = σ(FC2(ReLU(FC1(z))));

[0095] where FC1 and FC2 are the first and second fully connected layers respectively; ReLU represents the ReLU activation function; σ represents the Sigmoid activation function; y∈R C is the weight vector of the channel;

[0096] Finally, the weighted vector y is concatenated with the original input feature map E to obtain the weighted output:

[0097] E out = E × y;

[0098] where × represents element-wise multiplication. In this way, the feature maps of each channel are scaled according to their importance, thus highlighting the features of important channels.

[0099] The SE attention mechanism module can adaptively adjust the weights of feature channels, enabling the model to focus on more important features, improving the detection accuracy in complex scenarios, reducing the interference of redundant information, further enhancing the extraction effect of edge information, promoting the effective fusion of edge information and global features, and further improving the detection accuracy.

[0100] In step S4 of this embodiment, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. Among them, the training set is used for model training, the validation set is used for adjusting model parameters, and the test set is used for the final model evaluation to ensure the generalization ability and detection accuracy of the model. In particular, in this embodiment, the cross-validation method is used to select the training set and the validation set from the dataset and then perform model training and validation. On the one hand, the problem of insufficient training data is overcome through the cross-utilization of data, and on the other hand, the generalization ability of the model is improved, helping us find the best model hyperparameter configuration to improve the model detection accuracy.

[0101] After the model is trained and validated, it can be applied to small metal defect detection and output the detection results. However, to further optimize the detection results, it is also necessary to evaluate the model performance in the test set. In step S5 of this embodiment, evaluation methods such as calculating the mean average precision (mAP) and F1 score are used to obtain the standard evaluation indicators of the model to evaluate the model performance. If the evaluation result does not meet the expectation, it is possible to return to the model training and validation stage to further optimize the model to ensure that the final model achieves the best detection accuracy and efficiency.

[0102] Meanwhile, according to the above-mentioned method for detecting small metal defects, this embodiment also proposes a small metal defect detection system, which includes a processor for executing a program constructed according to the above method.

[0103] In summary, through innovations in data processing, transfer learning, model optimization, etc., the present invention effectively solves the problems of difficult collection of small metal datasets and insufficient extraction of edge information, improves the accuracy and efficiency of defect detection, and is particularly suitable for efficient detection tasks in complex defects and challenging scenarios.

[0104] The present invention has been described by way of example in conjunction with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention; or without improvement, the above concept and technical solution of the present invention are directly applied to other occasions, they are all within the protection scope of the present invention.

Claims

1. A small metal defect detection method, characterized in that: The method includes the following steps: Step S1: Obtain a small metal defect dataset and perform preprocessing; Step S2: Deploy the YOLOv8 model and perform pre-training; Step S3: For the backbone network of the pre-trained YOLOv8 model, introduce the Multi Scale Fusion module for improvement; Step S4: Divide the preprocessed small metal defect dataset into a training set, a validation set, and a test set, and use the training set to train the improved YOLOv8 model, and use the validation set to fine-tune the model parameters; Step S5: After training and validation are completed, use the test set to evaluate the improved YOLOv8 model.

2. The small metal defect detection method according to claim 1, wherein: In the step S1, the preprocessing includes augmenting the small metal defect dataset through data augmentation, and the data augmentation method includes geometric transformation and pixel transformation.

3. A small metal defect detection method according to claim 1, characterized in that: In the step S2, the ImageNet dataset is used for pre-training the YOLOv8 model.

4. A small metal defect detection method according to claim 1, characterized in that: In the step S3, all the original C2f modules in the YOLOv8 model backbone network are replaced with the Multi Scale Fusion module.

5. A small metal defect detection method according to claim 1 or 4, characterized in that: The MultiScale Fusion module is expressed by the formula as follows: For the input feature X ∈ R C×H×W , where C is the number of channels of the input, and H and W are the height and width respectively; First, process the input feature X through the convolutional layer Conv to generate a new feature map X1, that is: Next, for the feature map X1 after convolution, the Sobel operator is used to extract the image edge feature X sobel : X sobel = SobelConv(X1); Among them, SobelConv represents the Sobel convolution operation, which is divided into the horizontal direction and the vertical direction; the horizontal and vertical Sobel convolution kernels are initialized with fixed weights and are used to calculate the horizontal gradient G x and the vertical gradient G y ; Horizontal gradient G x and vertical gradient G y are calculated by the following formula: G x = X1 * K x , G y = X1 * K y ; where * represents the convolution operation; the horizontal Sobel kernel K x and the vertical Sobel kernel K y are respectively defined as: Finally, the edge feature map X sobel is as follows: X sobel = G x + G y ; In addition, a max pooling operation is also performed on the feature map X1 to obtain the downsampled feature map X pool : X pool = MaxPool(X1); Next, the edge feature X sobel and the pooling operation feature X pool are concatenated in the channel dimension to obtain the concatenated feature map X concat : X concat = concat(x sobel , x pool , dim = 1); where dim = 1 indicates concatenation in the channel dimension; For the stitched feature map X concat perform a convolution operation to obtain a new feature map X2: X2 = Conv2(X concat ) Finally, perform a convolution operation on the intermediate feature map X2: X out = Conv3(X2); to obtain the final output feature map X out .

6. A small metal defect detection method according to claim 5, characterized in that: In the step S3, add a layer of SE attention mechanism module after the SPPF layer in the YOLOv8 model backbone network.

7. A small metal defect detection method according to claim 6, characterized in that: The SE attention mechanism module performs global average pooling on the input feature map, then models the inter-channel relationship through two fully connected layers, and finally generates channel weights through Sigmoid activation and performs weighted fusion with the input feature map, and is expressed by the formula as follows: For the input feature map E ∈ R C×H×W perform global average pooling operation to obtain the global feature z for each channel: z = avgpool(E); Among them, the result of the pooling operation is to obtain a C-dimensional vector z ∈ R C , and each element represents the average value of the corresponding channel: Then, first compress the number of channels through a fully connected layer, then apply the ReLU activation function, then restore the number of channels through another fully connected layer, and finally apply the Sigmoid activation function to output the weight of each channel: y = σ(FC2(ReLU(FC1(z)))); Among them, FC1 and FC2 are the first and second fully connected layers respectively; ReLU represents the ReLU activation function; σ represents the Sigmoid activation function; y ∈ R C is the weight vector of the channel; Finally, perform weighted concatenation of the weight vector y and the original input feature map E to obtain the weighted output: E out = E × y; where × represents element-wise multiplication.

8. A small metal defect detection method according to claim 1, characterized in that: In the step S4, the cross-validation method is used to select the training set and the validation set from the dataset and then perform model training and validation.

9. A small metal defect detection method according to claim 1, characterized in that: In the step S5, the evaluation method includes calculating the mean average precision and the F1 score.

10. A small metal defect detection system according to the method described in any one of claims 1-9, characterized in that: The system includes a processor, and the processor is used to execute the program constructed according to any one of claims 1-9.