Metal surface defect detection method and system

Through the improved RT-DETR model, the adoption of additive attention encoder and multi-scale residual feature extraction module, the problem of slow detection speed of existing metal surface defect detection models is solved, and more efficient detection speed and lower computational complexity are achieved, suitable for real-time detection and edge device deployment.

CN120125515APending Publication Date: 2025-06-10HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178132.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing metal surface defect detection model has a large number of parameters and slow detection speed, which is not conducive to real-time detection and the deployment of edge equipment.

Method used

The improved RT-DETR model is adopted as the defect detection model. The improved RT-DETR model adopts an encoder with additive attention and a multi-scale residual feature extraction module. It replaces quadratic matrix multiplication by linear element-by-element multiplication to reduce the computational complexity and performs feature fusion through the context information feature fusion module.

Benefits of technology

It improves the speed of metal surface defect detection, reduces the calculation complexity and parameter quantity, and enhances the real-time and adaptability of detection, which is especially suitable for metal surface defect detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125515A_ABST
    Figure CN120125515A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of defect detection, and discloses a metal surface defect detection method and system, and the method comprises the steps: firstly obtaining to-be-detected metal surface image data; inputting the acquired metal surface image data into the trained defect detection model so as to realize defect detection on the metal surface; wherein the defect detection model adopts an improved RT-DETR model, an encoder adopted by the improved RT-DETR model is an additive attention encoder, the encoder generates a query vector and a key vector, and a calculation result is obtained through calculation by utilizing the key vector, the query vector and a global weight matrix. And processing the calculation result and the input of the encoder through addition normalization and a feedforward neural network to obtain the output of the encoder. And an additive attention encoder is adopted to improve the RT-DETR model, so that the defect detection capability is effectively improved, and the detection speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of defect detection, and in particular to a method and system for detecting metal surface defects. Background Art

[0002] Defect detection refers to the process of detecting possible defects, such as cracks, dents, bubbles, stains, etc. in products or materials through surface or internal inspection during industrial production. Defect detection can ensure product quality, protect consumers' rights and interests, and maintain the reputation of enterprises. Therefore, in existing industrial production, defect detection is an indispensable part. In the past, people often used traditional methods such as manual inspection, offline sampling, or a combination of both to detect metal surface defects. However, these traditional methods are not only inefficient but also prone to errors, resulting in a large waste of labor costs. In recent years, the emergence of machine learning has improved the defects of traditional methods and provided a more efficient and accurate alternative method for traditional methods.

[0003] With the continuous improvement of hardware performance and the rapid development of deep learning, the scope of application of neural networks is becoming more and more extensive. Currently, it has been successfully applied to the defect detection of industrial metal surfaces. For example, researchers have proposed MSAF-YOLOv8n, which combines MSAF and AP-RFB modules, to improve the accuracy of defect detection at the cost of speed.

[0004] Currently, the RT-DETR algorithm has been successfully applied to object detection, which can solve the post-processing problem in object detection models and improve the detection speed. On this basis, researchers have improved the RT-DETR algorithm and developed an enhanced EMSC-DETR model, which achieves high-precision detection by enhancing local and global feature interactions. However, the parameters and computational complexity of this model are higher. Researchers have also proposed the MCG-RTDETR network model, which uses DualConv and DeformableConvolution modules to capture complex feature information. However, the computational burden brought by multiple convolutions within the module seriously affects the real-time performance of the algorithm, and the module calculation speed is relatively low.

[0005] In summary, the currently common metal surface defect detection models have a large number of parameters and slow detection speeds, which are not conducive to real-time detection and the deployment of edge devices. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for detecting metal surface defects, aiming at the technical problem of the slow detection speed of existing metal surface defect detection models.

[0007] The present invention provides a method for detecting metal surface defects to solve the above technical problems. The method includes:

[0008] 1) Obtain the metal surface image data to be detected;

[0009] 2) Input the obtained metal surface image data into the trained defect detection model to achieve defect detection on the metal surface; the defect detection model adopts an improved RT-DETR model, and the encoder adopted by the improved RT-DETR model is an encoder based on additive attention. This encoder generates query vectors and key vectors through linear layers, calculates the normalized attention weights through matrix multiplication using the query vectors and the global weight matrix, sums the normalized attention weights with the query vectors after weighting, multiplies the obtained value element by element with the key vectors and then adds it to the original query vectors, adds the added result to the encoder input and normalizes it, and then inputs it into a feed-forward neural network for processing. The processed data is added and normalized with the data before the feed-forward neural network processing, and the result is the output of the encoder.

[0010] Furthermore, the improved RT-DETR model adopts a multi-scale residual feature extraction module as the backbone network. This multi-scale residual feature extraction module uses convolutional kernels of different sizes to extract features from the input image to obtain image features with different receptive fields.

[0011] Furthermore, the multi-scale residual feature extraction module includes a first convolutional layer, a feature extraction layer, and a second convolutional layer. The data input into the multi-scale residual feature extraction module is divided into two groups and processed separately through the first convolutional layer. One group of the convolved data is input into the feature extraction layer, and the feature extraction layer processes the convolved data through convolutional kernels of different sizes to obtain features containing different scale sizes. The output result of the feature extraction layer is concatenated with the data processed by the other group of the first convolutional layer and then processed through the second convolutional layer, and the obtained result is the output of the multi-scale residual feature extraction module.

[0012] Furthermore, the feature extraction layer includes N convolutional kernels of different sizes and a channel adjustment convolutional layer. The data input into the feature extraction layer passes through the N convolutional kernels of different sizes respectively to obtain features corresponding to different scales. The N convolutional results obtained are concatenated with the data input into the feature extraction layer without convolution, and the concatenated result is processed through the channel adjustment convolutional layer to obtain the output of the feature extraction layer, where N≥2.

[0013] Furthermore, the improved RT-DETR model adopts a context information feature fusion module as the feature fusion module to perform information fusion processing on the deep feature map and the shallow feature map.

[0014] Furthermore, the context information feature fusion module includes a fusion layer, a convolutional layer, and a partial convolutional layer group. The partial convolutional layer group contains at least one partial convolutional layer. The fusion layer is used to fuse the depth feature map and the shallow feature map. The output of the fusion layer is processed by the convolutional layer and the partial convolutional layer group respectively, and the output of the context information feature fusion module is obtained through splicing.

[0015] Furthermore, the fusion layer includes an average pooling module, a max pooling module, and a convolutional module. After the depth feature map is input into the fusion module, it is processed by the average pooling module and the max pooling module respectively. The average pooling result and the max pooling result are concatenated and then the convolutional module extracts feature information and generates a weight matrix. The weight matrix is combined with the shallow feature map using the Hadamard product operation to obtain the output of the fusion layer.

[0016] A metal surface defect detection system includes a processor, which is used to run a computer program to implement the metal surface defect detection method.

[0017] The beneficial effects of the present invention are as follows: As an improved invention, the present invention uses an improved RT-DETR model as the defect detection model. Among them, the encoder of the improved RT-DETR model uses an encoder with additive attention. The encoder with additive attention calculates the calculation result using the global weight matrix and the key vector and query vector generated by the linear layer, and processes the calculation result and the encoder input through addition normalization and a feed-forward neural network to obtain the output of the encoder. The encoder with additive attention is used to improve the RT-DETR model. By introducing linear element-wise multiplication to replace the quadratic matrix multiplication operation, the computational complexity is reduced and the detection speed is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the present invention;

[0019] Figure 2 is a schematic structural diagram of an efficient encoder with additive attention (EEAA);

[0020] Figure 3 is a schematic structural diagram of a multi-scale residual feature extraction (MSRFE) module;

[0021] Figure 4 is a schematic structural diagram of a context information fusion module (CFIF);

[0022] Figure 5 is a comparison diagram of detection effects compared with other models;

[0023] Figure 6 is a schematic structural diagram of the overall structure of the improved RT-DETR model. DETAILED DESCRIPTION OF THE INVENTION

[0024] The following further describes the specific implementation manners of the present invention with reference to the accompanying drawings.

[0025] The present invention uses an improved RT-DETR model for defect detection on the metal surface. Among them, the encoder adopted by the improved RT-DETR model is an encoder with additive attention. This encoder generates query vectors and key vectors through linear layers, performs matrix multiplication calculation using the query vectors and the global weight matrix to obtain normalized attention weights, weights and sums the normalized attention weights with the query vectors, multiplies the obtained values element by element with the key vectors and then adds them to the original query vectors, adds and normalizes the addition result with the encoder input and then inputs it into a feed-forward neural network for processing, and adds and normalizes the processed data with the data before being processed by the feed-forward neural network to obtain the output of the encoder. By using the encoder with additive attention, the operation of quadratic matrix multiplication is replaced by introducing linear element-wise multiplication, reducing the computational complexity and effectively improving the defect detection speed on the metal surface.

[0026] An embodiment of a method for detecting defects on a metal surface

[0027] As Figure 1 shown, the present invention proposes a method for detecting defects on a metal surface. First, image data of the metal surface to be detected is obtained. Subsequently, the obtained image data of the metal surface is input into a trained defect detection model to achieve defect detection on the metal surface. The overall structure of the improved RT-DETR model is as Figure 6 shown.

[0028] Among them, the backbone of the improved RT-DETR model is a multi-scale residual feature extraction module MSRFE. This feature extraction module uses different convolutional kernels to obtain different receptive fields for feature extraction. In the improved RT-DETR model, the feature fusion module CCFM of the original RT-DETR model is replaced by a context information feature fusion module CFIF. Through this feature fusion module, information fusion processing is performed on the deep feature map and the shallow feature map to obtain a feature fusion result. The AIFI encoder of the original RT-DETR model is replaced by an encoder with additive attention EEAA.

[0029] Specifically, as Figure 3 shown, the improved RT-DETR model replaces the backbone network part with a lightweight multi-scale residual feature extraction module MSRFE. This module uses convolutional kernels of different sizes to perform feature extraction on the input image to obtain image features with different receptive fields.

[0030] The multi-scale residual feature extraction module includes a first convolutional layer, a feature extraction layer MSFE, and a second convolutional layer. Among them, the first convolutional layer can be a 3×3 convolutional layer, and the second convolutional layer can be a 1×1 convolutional layer. After the data is input into the multi-scale residual feature extraction module and divided into two groups and processed by the first convolutional layer respectively, one group of the convolved data is input into the feature extraction layer MSFE. The feature extraction layer MSFE includes N convolutional kernels and a channel adjustment convolutional layer, where N≥2. The convolved data is grouped according to the number of convolutional kernels in the feature extraction layer MSFE, and the number of groups is the number of convolutional kernels plus one. Since the shapes and sizes of metal surface defects are irregular, convolutional kernels of different sizes are used to perform convolutional processing on the data except the first group respectively, and the convolutional results and the non-convolved features are Concat spliced, and the spliced result is processed by the channel adjustment convolutional layer to obtain the output of the feature extraction layer MSFE. The output of the feature extraction layer is spliced with the data processed by the other group of the first convolutional layer and then processed by the second convolutional layer to obtain the output of the multi-scale residual feature extraction module.

[0031] Taking the example of dividing the convolved data into four groups, the convolved data is divided into four parts xi∈{x1, x2, x3, x4}, and the number of channels of each xi is the same, and the number of channels of xi is reduced to one-fourth of the original. The three convolutional kernels in the multi-convolutional kernel module are 3×3 convolution, 5×5 convolution, and 7×7 convolution respectively. Among them, the data in the x1 part is not convolved and retains the original features to obtain the result y1. x2, x3, and x4 are respectively convolved by 3×3 convolution, 5×5 convolution, and 7×7 convolution to obtain the processed results y2, y3, and y4. Then, the processed results yi∈{y1, y2, y3, y4} are Concat spliced, and the number of channels of the spliced result is adjusted by 1×1 convolution to obtain the feature extraction result.

[0032] The multi-scale residual feature extraction layer uses convolutional kernels of different sizes for feature extraction, which helps the model obtain different receptive fields. It can not only help the model capture local feature information and understand global feature information, but also reduce the number of model parameters and computational complexity, and improve the model detection speed.

[0033] As Figure 4 shown, the improved RT-DETR model uses the context information feature fusion module as the feature fusion module to perform information fusion processing on the deep feature map and the shallow feature map. The context information feature fusion module includes a fusion layer, a convolutional layer, and a partial convolutional layer group. The partial convolutional layer group contains at least one partial convolutional layer, and the convolutional layer can be a 3×3 convolutional layer. The deep feature map and the shallow feature map are processed and combined by the fusion layer to obtain the output of the fusion layer. The output of the fusion layer is processed by the convolutional layer and the partial convolutional layer group respectively, and the output of the context information feature fusion module is obtained through splicing.

[0034] Among them, the fusion layer includes an average pooling module, a max pooling module, and a convolution module. After the depth feature map is input into the fusion layer, it is processed by the average pooling module and the max pooling module respectively. After connecting the average pooling result and the max pooling result, the convolution module extracts feature information and generates a weight matrix. The weight matrix is combined with the shallow feature map using the Hadamard product operation to obtain the output of the fusion layer. The convolution module can select a 7×7 convolution.

[0035] The context information feature fusion module can help the model perform efficient feature fusion, improve the detection performance of the model, and reduce the number of model parameters.

[0036] Such as Figure 2 As shown, the encoder based on additive attention generates query vectors and key vectors through a linear layer, calculates the normalized attention weights using the query vectors and the trainable global weight matrix obtained by initializing the RT-DETR model model parameters, and weights and sums the query vectors using the normalized attention weights. The obtained value is multiplied element-wise with the key vectors and then added to the original query vectors. The added result is added and normalized with the encoder input and then input into a feed-forward neural network for processing. The processed data is added and normalized with the data before the feed-forward neural network processing to obtain the output of the encoder.

[0037] The encoder module based on additive attention can help the model distinguish the features of different types of defects, improve the detection ability, and increase the detection speed of the model.

[0038] As another implementation, the improved RT-DETR model can also replace only the encoder with an encoder based on additive attention, or on the basis of using an encoder based on additive attention, replace the backbone with a multi-scale residual feature extraction module or replace the feature fusion module with a context information feature fusion module.

[0039] In order to verify that the improved RT-DETR model proposed by the present invention has a fast detection speed and accurate detection results, the improved RT-DETR model proposed in this embodiment is used to process the metal defect dataset.

[0040] The present invention uses a total of three datasets. The first is the steel surface defect dataset (NEU-DET) provided by Northeastern University, which contains 6 typical defects, namely rolled-in scale (Rs), patches (Pa), crazing (Cr), pitted surface (Ps), inclusion (In), and scratches (Sc). Each type of defect has 300 samples, for a total of 1800 grayscale images, and the resolution of each image is 200×200. The second is a public dataset of steel surface defects: GC10-DET. This dataset contains 2294 steel plate surface defect images with a resolution of 2048×1000 pixels, covering ten surface defect types: punching (Pu), weld (Wl), crescent pattern (Cg), water stain (Ws), oil stain (Os), silk stain (Ss), inclusion (In), roll mark (Rp), crease mark (Cr), and edge wave (Wf). The last one is a public dataset of aluminum plate surface defects, ASSDD. It contains a total of 1400 aluminum plate surface defect images, each image with a size of 640×480 pixels, and contains four defect types:

[0041] Pinhole (Pi), scratch (Sc), stain (Di), and wrinkle (Wr).

[0042] To verify the effectiveness of the improved RT-DETR model, a comparative experiment was first conducted on the NEU-DET dataset, as shown in Table 1.

[0043] Table 1. Comparative Experiment

[0044]

[0045] Table 1 shows that in the original RT-DETR model, its mAP is 74.0%, the FPS speed is 119.04, the Params parameter quantity is 38.31M, and the GFLOPs calculation amount is 57.0. When using the improved RT-DETR model, the mAP reaches 76.4%, the FPS reaches 158.73, which is a 33.3% increase compared to the original model; the Params parameter quantity is only 16.94M, which is a 55.7% decrease compared to the original model; the GFLOPs calculation amount is reduced to 28.4, a reduction of 50.1%.

[0046] The proposed improved RT-DETR model was compared with other advanced models using the NEU-DET dataset, including the original RT-DETR, RT-DETR-L, Faster RCNN, YOLOv7, YOLOv8l, and YOLOv10m, and the comparison results are presented in Table 2.

[0047] Compared with other advanced algorithms, the method proposed in the present invention achieves the highest detection accuracy for defects In, Rs, and Sc, and is also the most outstanding in terms of average precision index, detection speed, and lightweight index. These findings highlight the excellent performance of the improved RT-DETR model, making it particularly suitable for metal surface defect detection tasks.

[0048] In addition, Figure 5 The comparison of the detection effects of multiple models on the NEU-DET dataset is shown. It can be found that the improved RT-DETR model has the best detection ability, making up for the missed detection and misdetection of other models.

[0049] Table 2. Comparison with other algorithms on the NEU-DET dataset

[0050]

[0051] As can be seen from Table 2, the RT-DETR model proposed in the present invention also achieves an excellent detection speed of 158.73 FPS, far exceeding other models. At the same time, the RT-DETR model proposed in the present invention also achieves excellent results in terms of lightweight index, with the number of parameters Params and GFLOPs values being 16.94 and 28.4 respectively, much lower than other algorithms. By comparing the detection result data of the improved RT-DETR model with other algorithms, the excellent performance of the improved RT-DETR model in terms of accuracy, speed, and lightweight design is highlighted, making it particularly suitable for metal surface defect detection tasks.

[0052] The performance of different algorithms is compared using two other datasets, GC10-DET and ASSDD, and the results are shown in Tables 3 and 4. The comparison of performance parameters in the tables shows that the algorithm proposed in the present invention has excellent generalization and excellent performance on different datasets.

[0053] Table 3. Comparison with other algorithms on the GC10-DET dataset

[0054]

[0055] Table 4. Comparison with other algorithms on the ASSDD dataset

[0056]

[0057] An embodiment of a metal surface defect detection system

[0058] The present invention proposes a metal surface defect detection system, including a processor for running a computer program to implement the metal surface defect detection method.

[0059] The specific implementation process has been described in detail in the embodiments of the defect detection method, and will not be elaborated here.

Claims

1. A method for detecting metal surface defects, characterized in that: The method includes: 1) Obtaining image data of the metal surface to be detected; 2) The acquired metal surface image data is input into the trained defect detection model to realize defect detection on the metal surface; the defect detection model adopts an improved RT-DETR model, and the encoder adopted by the improved RT-DETR model is an encoder based on additive attention. The encoder generates a query vector and a key vector through a linear layer, and uses the query vector and the global weight matrix to perform matrix multiplication to obtain a normalized attention weight. The normalized attention weight is weighted and summed with the query vector, and the obtained value is multiplied element by element with the key vector and then added to the original query vector. The addition result is added to the encoder input and normalized, and then input into the feedforward neural network for processing. The processed data is added and normalized with the data before being processed by the feedforward neural network, and the result is the output of the encoder.

2. The metal surface defect detection method according to claim 1, characterized in that: The improved RT-DETR model adopts a multi-scale residual feature extraction module as a backbone network. The multi-scale residual feature extraction module uses convolution kernels of different sizes to extract features from the input image to obtain image features of different receptive fields.

3. The metal surface defect detection method according to claim 2, characterized in that: The multi-scale residual feature extraction module includes a first convolution layer, a feature extraction layer and a second convolution layer. The data input to the multi-scale residual feature extraction module is divided into two groups and processed by the first convolution layer respectively. One group of convolved data is input to the feature extraction layer, and the feature extraction layer processes the convolved data through convolution kernels of different sizes to obtain features containing different scales. The output result of the feature extraction layer is spliced ​​with another group of data processed by the first convolution layer and then processed by the second convolution layer, and the result is the output of the multi-scale residual feature extraction module.

4. The metal surface defect detection method according to claim 3, characterized in that: The feature extraction layer includes N convolution kernels of different sizes and a channel adjustment convolution layer. The data input to the feature extraction layer are respectively passed through N convolution kernels of different sizes to obtain features of corresponding scales. The N convolution results obtained are spliced ​​with the data input to the feature extraction layer and not convolved, and the splicing results are processed by the channel adjustment convolution layer to obtain the output of the feature extraction layer, where N≥2.

5. The metal surface defect detection method according to claim 1, characterized in that: The improved RT-DETR model adopts a context information feature fusion module as a feature fusion module to perform information fusion processing on the deep feature map and the shallow feature map.

6. The metal surface defect detection method according to claim 5, characterized in that: The context information feature fusion module includes a fusion layer, a convolution layer and a partial convolution layer group. The partial convolution layer group includes at least one partial convolution layer. The fusion layer is used to fuse the deep feature map and the shallow feature map. The output of the fusion layer is processed by the convolution layer and the partial convolution layer group respectively, and the output of the context information feature fusion module is obtained by splicing.

7. The metal surface defect detection method according to claim 6, characterized in that: The fusion layer includes an average pooling module, a maximum pooling module and a convolution module. After the deep feature map is input into the fusion module, it is processed by the average pooling module and the maximum pooling module respectively. After the average pooling result and the maximum pooling result are connected, the feature information is extracted through the convolution module and a weight matrix is ​​generated. The weight matrix is ​​combined with the shallow feature map using the Hadamard product operation to obtain the output of the fusion layer.

8. A metal surface defect detection system, characterized in that: A processor is included, and the processor is used to run a computer program to implement the metal surface defect detection method as described in any one of claims 1 to 7.