Steel surface defect detection method and system based on improved RT-DETR model

By improving the multi-scale residual connection backbone network and hybrid query propagation mechanism of the RT-DETR model, the problems of low computational efficiency and insufficient detection accuracy in low-contrast scenarios in existing methods are solved, and efficient and accurate steel surface defect detection is achieved.

CN120747602APending Publication Date: 2025-10-03NANJING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510840571.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing Transformer-based steel surface defect detection methods have low computational efficiency and are unable to meet the needs of efficient and accurate detection. In addition, it is difficult to distinguish defects from background in low-contrast scenes, resulting in insufficient detection accuracy.

Method used

An improved RT-DETR model is adopted to enhance the feature fusion capability and dynamic adjustment of the decoder structure by constructing a multi-scale residual connection backbone network, a hybrid query propagation mechanism and a dual hybrid strategy, thereby achieving multi-scale feature detection and supervision signal balance.

Benefits of technology

It improves the accuracy and efficiency of steel surface defect detection, enhances adaptability to lighting and background interference, reduces computing overhead, and achieves high-precision real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747602A_ABST
    Figure CN120747602A_ABST
Patent Text Reader

Abstract

The invention discloses a steel surface defect detection method based on an improved RT-DETR model, and the method specifically comprises the steps: firstly obtaining a steel defect surface detection data set, carrying out the input preprocessing of an image, and dividing the image into a training set, a test set and a verification set according to a preset proportion; then, an improved RT-DETR model is constructed; then, inputting the preprocessed training set image into the improved RT-DETR model for training, and obtaining a trained improved RT-DETR model; and finally, inputting a to-be-detected test set image into the trained improved RT-DETR model to obtain a steel surface defect detection result. According to the method, the phenomena of missing detection and false detection in the steel surface defect detection process are effectively reduced, the steel surface defect detection precision and detection efficiency in a complex scene are improved, and the method has high adaptability to variable changes such as illumination and background interference in an actual production environment and can be directly deployed in an industrial production line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steel surface defect detection, and in particular to a steel surface defect detection method and system based on an improved RT-DETR model. Background Art

[0002] Surface defects such as cracks, patches, inclusions, and scratches are common during the steel production process, posing a serious threat to structural safety and reliability. Therefore, surface defect detection has become a critical component of quality control in the steel industry. In traditional industrial settings, surface defects have long relied on manual detection. However, this method has significant limitations, such as high labor intensity, high subjective variability, and low efficiency in detecting subtle or multi-scale defects, making it difficult to meet the efficient and accurate detection needs of modern industry.

[0003] As DETR introduced the transformer structure into the field of target detection, the transformer-based detection network has shown potential in steel defect detection due to its powerful long-distance dependency modeling capabilities. However, existing methods still have two major bottlenecks. First, the computational efficiency is low. The high computational complexity of the self-attention mechanism limits the application of the model to multi-scale features. The mainstream solution is forced to use a single-level feature map for prediction, resulting in slow convergence and insufficient detection accuracy. In addition, it has poor adaptability to low-contrast scenes. When the defect is similar to the background, the standard self-attention layer is difficult to effectively distinguish the target features, resulting in a large amount of redundant calculations. Therefore, it is of urgent industrial value to develop a steel surface defect detection method with both high precision and low computational overhead. Summary of the Invention

[0004] The object of the present invention is to provide a steel surface defect detection method and system based on an improved RT-DETR model with high detection accuracy and high detection efficiency.

[0005] The technical solution for achieving the purpose of the present invention is: a steel surface defect detection method based on an improved RT-DETR model, comprising the following steps:

[0006] Step 1: Obtain a steel defect surface detection dataset, pre-process the image input, and divide it into a training set, a test set, and a validation set according to a preset ratio;

[0007] Step 2: Build an improved RT-DETR model, including:

[0008] Step 2.1: Construct a multi-scale feature enhancement module and stack it with the Bottleneck module to form a multi-scale residual connection backbone network.

[0009] Step 2.2: Construct a hybrid query propagation mechanism to balance the supervision signals between different decoding layers by collecting and propagating query vectors from the middle layer.

[0010] Step 2.3: Construct a dual hybrid strategy and dynamically adjust the decoder structure;

[0011] Step 3: Input the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model;

[0012] Step 4: Input the test set images to be detected into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

[0013] Furthermore, the steel defect surface detection dataset described in step 1 is obtained, the image is input preprocessed, and divided into a training set, a test set, and a validation set according to a preset ratio, as follows:

[0014] Step 1.1: Obtain a dataset of steel defect surface detection, normalize the size of the original image, scale the image to 800 × 800 with any resolution, maintain the aspect ratio, and fill the blank area with 114 grayscale values;

[0015] Step 1.2: Normalize the image pixels and perform z-score standardization with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225].

[0016] Step 1.3: Perform Mosaic data enhancement on the image;

[0017] Step 1.4: Divide the preprocessed detection data set into training set, test set and validation set according to the preset ratio.

[0018] Furthermore, the multi-scale feature enhancement module described in step 2.1 is constructed as follows:

[0019] (1) Use depth-wise separable convolution with a convolution kernel of 7×7 to process the original input image and obtain the initial feature map;

[0020] (2) Split the initial feature map into two branches, directly input the feature map of one branch into a convolution layer with a convolution kernel of 1×1 for processing; input the feature map of the other branch into a convolution layer with a convolution kernel of 1×1 and a ReLU activation function for processing;

[0021] (3) Perform vector dot product operations on the feature maps processed by the two branches to achieve interaction and fusion between features;

[0022] (4) The fused feature maps are sequentially input into the fully connected network, and refined by the depthwise separable convolution with a convolution kernel of 7×7 and the DropOut layer to obtain the final output features;

[0023] (5) Finally, the Dropout layer outputs the refined feature map.

[0024] Furthermore, the multi-scale residual connection backbone network described in step 2.1 consists of six stages connected in sequence:

[0025] The first stage consists of an initial module, including a convolutional layer with a 3×3 kernel, 64 output channels, a stride of 2, and a ReLu activation function.

[0026] The second and third stages each consist of a multi-scale feature enhancement module. The first module includes a depthwise separable convolution layer with a 7×7 convolution kernel, 256 output channels, and a stride of 1, a feature transformation layer with a dual-way convolution kernel of 1×1 and 768 output channels, a feature fusion layer based on element-wise multiplication, a feature compression layer with a 1×1 convolution kernel and 256 output channels, and a residual connection structure. The second module has 512 input channels and 1024 output channels, and its module structure is the same as that of the second stage.

[0027] The fourth stage is composed of a stack of Bottleneck modules, which include a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 2, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure.

[0028] The fifth stage consists of two stacked Bottleneck modules, both of which contain a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 1, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure;

[0029] The sixth stage is composed of three stacked Bottleneck modules. The first module contains a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512, and a stride of 2, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure; the last two modules contain a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512, and a stride of 1, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure;

[0030] An 800×800×3 image is input to the multi-scale residual connection backbone network. In the third stage, a 100×100×512 feature map is output. In the fifth stage, a 50×50×1024 feature map is output. In the sixth stage, a 25×25×2048 feature map is output to form a multi-scale defect feature pyramid. The 100×100×512, 50×50×1024, and 25×25×2048 feature maps are aligned to 256 channels by a convolution with a convolution kernel of 1×1 and then spliced. The spliced ​​vector is flattened into a vector of size 13125×256 as the input of the Transformer encoder.

[0031] Furthermore, the hybrid query propagation mechanism described in step 2.2 is constructed as follows:

[0032] (1) Input the initial query vector q0 into the first layer of the decoder D1 to generate the intermediate feature set D1(q0);

[0033] (2) Merge q0 and D1(q0) to form the first-layer output set O1 = q0 ∪ {D1(q0)};

[0034] (3) For the decoder layer i D i , where i = 2, 3, ..., n, output O from the previous layer i-1 Extract the query vector q i-1 , q i-1 Input D i , generating set D i (q i-1 ), merge q i-1 With D i (q i-1 ) Get the current layer output O i =O i-1 ∪{D i (q)};

[0035] (4) The sixth layer output O6 is used as the result of the entire decoder;

[0036] The formula of the hybrid query propagation mechanism is as follows:

[0037] O0={q0}

[0038] O i =O i-1 ∪{D i (q)|q∈O i-1}

[0039] Among them, D i represents the i-th layer of the decoder, q i represents the query after being processed by the i-th layer of the decoder, O i represents the output of the i-th layer of the decoder.

[0040] Furthermore, the dual hybrid strategy described in step 2.3 includes a hybrid strategy in the training round dimension and a hybrid strategy in the decoder layer dimension. That is, in the training process with a total of 200 rounds, the conventional query propagation mechanism is applied to all layers of the six-layer decoder in the first 150 rounds of training; and the hybrid query propagation mechanism is switched to the last four layers of the six-layer decoder in the last 50 rounds of training.

[0041] Furthermore, in the conventional query propagation mechanism, the initial query vector passes through the six layers of decoders in sequence, and the i-th layer decoder is based on the input query q i-1 Generate output query q i , i=2,3,…,n, and finally the output q6 of the sixth layer is used as the result of the entire decoder.

[0042] A steel surface defect detection system based on an improved RT-DETR model is provided. The system is used to implement the steel surface defect detection method based on the improved RT-DETR model. The system includes a first module to a fourth module, wherein:

[0043] The first module obtains a steel defect surface detection dataset, pre-processes the images, and divides them into training, test, and validation sets according to a pre-set ratio.

[0044] The second module is to build an improved RT-DETR model, including:

[0045] The first sub-unit builds a multi-scale feature enhancement module, which is stacked and connected with the Bottleneck module to form a multi-scale residual connection backbone network;

[0046] The second sub-unit builds a hybrid query propagation mechanism, which balances the supervision signals between different decoding layers by collecting and propagating query vectors in the middle layer;

[0047] The third subunit builds a dual hybrid strategy and dynamically adjusts the decoder structure;

[0048] The third module inputs the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model;

[0049] In the fourth module, the test set images to be detected are input into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

[0050] A mobile terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for detecting steel surface defects based on the improved RT-DETR model is implemented.

[0051] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the steel surface defect detection method based on the improved RT-DETR model.

[0052] Compared with the prior art, the present invention has the following significant advantages:

[0053] (1) Using a multi-scale residual connection backbone network, it achieves real-time inference performance while ensuring high robustness. It is highly adaptable to changes in variables such as lighting and background interference in actual production environments and can be directly deployed on industrial production lines.

[0054] (2) A multi-scale feature fusion module is designed and stacked with the Bottleneck module to form an improved multi-scale residual connection backbone network. This effectively improves the detection accuracy of small target defects while expanding the channel dimension of the feature map.

[0055] (3) Through the hybrid query propagation mechanism, the supervision intensity gradients of different decoder layers are dynamically allocated, which improves the accuracy of steel surface defect detection in complex scenarios;

[0056] (4) The dual hybrid strategy synergistically optimizes the training convergence speed and computational overhead, accelerates the late convergence of the model and reduces the computational overhead, solves the accuracy bottleneck caused by computational redundancy of the traditional Transformer detector, and improves the accuracy and efficiency of steel surface defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 The present invention is a flow chart of a steel surface defect detection method based on an improved RT-DETR model.

[0058] Figure 2 Schematic diagram of the structure of the improved RT-DETR model in the present invention.

[0059] Figure 3 Schematic diagram of the structure of the multi-scale feature enhancement module in the present invention.

[0060] Figure 4 Schematic diagram of the hybrid query propagation mechanism in the present invention.

[0061] Figure 5 Schematic diagram of the structure of the mixed round strategy in the present invention.

[0062] Figure 6 This is a comparison chart of the results of steel surface defect detection using different methods in an embodiment of the present invention. DETAILED DESCRIPTION

[0063] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0064] Combine Figure 1 The present invention provides a steel surface defect detection method based on an improved RT-DETR model, comprising the following steps:

[0065] Step 1: Obtain a steel defect surface detection dataset, pre-process the image input, and divide it into a training set, a test set, and a validation set according to a preset ratio;

[0066] Step 2: Build an improved RT-DETR model, including:

[0067] Step 2.1: Construct a multi-scale feature enhancement module and stack it with the Bottleneck module to form a multi-scale residual connection backbone network.

[0068] Step 2.2: Construct a hybrid query propagation mechanism to balance the supervision signals between different decoding layers by collecting and propagating query vectors from the middle layer.

[0069] Step 2.3: Construct a dual hybrid strategy and dynamically adjust the decoder structure;

[0070] Step 3: Input the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model;

[0071] Step 4: Input the test set images to be detected into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

[0072] As a specific example, in step 1, a steel defect surface detection dataset is obtained, the images are input preprocessed, and divided into a training set, a test set, and a validation set according to a preset ratio, as follows:

[0073] Step 1.1: Obtain a steel defect surface detection dataset, such as the NEU-DET dataset or the GC10-DET dataset. Normalize the original image size and scale the image to 800 × 800 with any resolution, maintaining the aspect ratio. Fill the blank areas with 114 grayscale values.

[0074] Step 1.2: Normalize the image pixels and perform z-score standardization with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225].

[0075] Step 1.3: Perform Mosaic data enhancement on the image;

[0076] Step 1.4: Divide the preprocessed detection data set into a training set, a test set, and a validation set in a ratio of 8:1:1.

[0077] As a specific example, in step 2, an improved RT-DETR model is constructed, such as Figure 2 As shown, the details are as follows:

[0078] Step 2.1: Construct a multi-scale feature enhancement module and stack it with the Bottleneck module to form a multi-scale residual connection backbone network.

[0079] The multi-scale feature enhancement module is constructed as follows:

[0080] (1) Use depth-wise separable convolution with a convolution kernel of 7×7 to process the original input image and obtain the initial feature map;

[0081] (2) Splitting the initial feature map into two branches, directly inputting the feature map of one branch into a convolution layer with a convolution kernel of 1×1 for processing; inputting the feature map of the other branch into a convolution layer with a convolution kernel of 1×1 and a ReLU activation function for processing;

[0082] (3) Perform vector dot product operations on the feature maps processed by the two branches to achieve interaction and fusion between features;

[0083] (4) The fused feature maps are sequentially input into the fully connected network, and refined by the depthwise separable convolution with a convolution kernel of 7×7 and the DropOut layer to obtain the final output features;

[0084] (5) Finally, the Dropout layer outputs the refined feature map.

[0085] like Figure 3 As shown, the multi-scale residual connection backbone network includes six stages connected in sequence:

[0086] The first stage consists of an initial module, including a convolutional layer with a 3×3 kernel, 64 output channels, a stride of 2, and a ReLu activation function.

[0087] The second and third stages each consist of a multi-scale feature enhancement module. The first module includes a depthwise separable convolution layer with a 7×7 convolution kernel, 256 output channels, and a stride of 1, a feature transformation layer with a dual-way convolution kernel of 1×1 and 768 output channels, a feature fusion layer based on element-wise multiplication, a feature compression layer with a 1×1 convolution kernel and 256 output channels, and a residual connection structure. The second module has 512 input channels and 1024 output channels, and its module structure is the same as that of the second stage.

[0088] The fourth stage is composed of a stack of Bottleneck modules, which include a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 2, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure.

[0089] The fifth stage consists of two stacked Bottleneck modules, both of which contain a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 1, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure;

[0090] The sixth stage is composed of three stacked Bottleneck modules. The first module contains a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512 and a stride of 2, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure; the last two modules contain a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512 and a stride of 1, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure;

[0091] An 800×800×3 image is input to the multi-scale residual connection backbone network. In the third stage, a 100×100×512 feature map is output. In the fifth stage, a 50×50×1024 feature map is output. In the sixth stage, a 25×25×2048 feature map is output to form a multi-scale defect feature pyramid. The 100×100×512, 50×50×1024, and 25×25×2048 feature maps are aligned to 256 channels by a convolution with a convolution kernel of 1×1 and then spliced. The spliced ​​vector is flattened into a vector of size 13125×256 as the input of the Transformer encoder.

[0092] Furthermore, the concatenated feature maps are input into the Transformer encoder, which adopts a 6-layer standard Transformer structure. Its implementation refers to the RT-DETR open source code, and its output feature sequence dimension is (256, 13125) and contains learnable position encoding information.

[0093] Furthermore, the feature sequence of dimension (256, 13125) output by the Transformer encoder is input into the improved decoder as the key and value. The decoder takes 300 learnable query vectors as input. The query propagation mechanism of the decoder is executed according to the hybrid query propagation strategy; the training strategy is executed according to the dual hybrid strategy.

[0094] Step 2.2: Construct a hybrid query propagation mechanism to balance the supervision signals between different decoding layers by collecting and propagating query vectors from the middle layer.

[0095] like Figure 4 As shown, the hybrid query propagation mechanism is constructed as follows:

[0096] (1) Input the initial query vector q0 into the first layer of the decoder D1 to generate the intermediate feature set D1(q0);

[0097] (2) Merge q0 and D1(q0) to form the first-layer output set O1 = q0 ∪ {D1(q0)};

[0098] (3) For the decoder layer i D i , where i = 2, 3, ..., n, output O from the previous layer i-1 Extract the query vector q i-1 , change q i-1 Input D i , generating set D i (q i-1 ), merge q i-1 With D i (q i-1 ) Get the current layer output O i =O i-1 ∪{D i (q)};

[0099] (4) The sixth layer output O6 is used as the result of the entire decoder.

[0100] The formula of the hybrid query propagation mechanism is as follows:

[0101] O0={q0}

[0102] O i =O i-1 ∪{D i (q)|q∈O i-1}

[0103] Among them, D i represents the i-th layer of the decoder, q i represents the query after being processed by the i-th layer of the decoder, O i represents the output of the i-th layer of the decoder.

[0104] Step 2.3: Construct a dual hybrid strategy and dynamically adjust the decoder structure;

[0105] like Figure 5 As shown, the dual hybrid strategy includes a hybrid strategy in the training round dimension and a hybrid strategy in the decoder layer dimension. That is, in a training process with a total of 200 rounds, the conventional query propagation mechanism is applied to all layers of the six-layer decoder in the first 150 rounds of training; and the hybrid query propagation mechanism is switched to the last four layers of the six-layer decoder in the last 50 rounds of training.

[0106] In the conventional query propagation mechanism, the initial query vector passes through the six layers of decoders in sequence. The i-th layer decoder is based on the input query q i-1 Generate output query q i , i=2,3,…,n, and finally the output q6 of the sixth layer is used as the result of the entire decoder.

[0107] As a specific example, in step 4, the test set images to be detected are input into the trained improved RT-DETR model to obtain the detection results of steel surface defects, as follows:

[0108] The 300×256 feature vector output by the decoder is passed through the detection head to generate a prediction result. The detection head adopts the DETR standard detection head and contains two branches. Among them, the bounding box regression branch contains a 3-layer fully connected network with dimensions of 256, 256, and 4, and outputs normalized bounding box coordinates (x, y, w, h). The defect classification branch contains a 3-layer fully connected network with dimensions of 256, 256, and 6, and outputs the probabilities of six types of defects. Finally, the model outputs the defect location and category information.

[0109] The present invention also provides a steel surface defect detection system based on an improved RT-DETR model, which is used to implement the steel surface defect detection method based on the improved RT-DETR model. The system includes a first module to a fourth module, wherein:

[0110] The first module obtains a steel defect surface detection dataset, pre-processes the images, and divides them into training, test, and validation sets according to a pre-set ratio.

[0111] The second module is to build an improved RT-DETR model, including:

[0112] The first sub-unit builds a multi-scale feature enhancement module, which is stacked and connected with the Bottleneck module to form a multi-scale residual connection backbone network;

[0113] The second sub-unit builds a hybrid query propagation mechanism, which balances the supervision signals between different decoding layers by collecting and propagating query vectors in the middle layer;

[0114] The third subunit builds a dual hybrid strategy and dynamically adjusts the decoder structure;

[0115] The third module inputs the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model;

[0116] In the fourth module, the test set images to be detected are input into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

[0117] The present invention also provides a mobile terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for detecting steel surface defects based on the improved RT-DETR model is implemented.

[0118] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the steel surface defect detection method based on the improved RT-DETR model.

[0119] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0120] Example

[0121] This embodiment uses the publicly available steel surface defect dataset NEU-DET, which contains 1,800 images of six defect types, with 300 images of each defect. The dataset is divided into training set, validation set, and test set in a ratio of 8:1:1. The training platform includes a Xeon(R) Gold 6430 CPU with a main frequency of 2.1GHz and an RTX4090 GPU. The computer version is Ubuntu 22.04, Python version is 3.10, and CUDA version is 11.8. In the PyTorch deep learning framework, the training set is input into the improved RT-DETR model for training, and the trained model is tested using the test set and compared with Faster-RCNN, RetinaNet, CenterNet, YOLOv5, YOLOv8, YOLOv11, Deformable-DETR, DINO, DAB-DETR, and RT-DETR. All methods are trained using the Adam optimization algorithm with a batch value of 4.

[0122] Visualization of detection results Figure 6As shown in the figure, (a) is the original image, (b) is the true value, (c) is the detection result of the improved RT-DETR on the test set, (d) is the detection result of the YOLOV11 on the test set, (e) is the detection result of the RetinaNet on the test set, (f) is the detection result of the YOLOV8 on the test set, and (g) is the detection result of the Faster-RCNN on the test set. The comparative test results are shown in Table 1.

[0123] Table 1 Comparative test results of steel surface defect detection using different methods

[0124]

[0125] Depend on Figure 6 As can be seen from Table 1, the proposed method can effectively reduce the missed detection rate and the repeated detection rate, and significantly improve the detection accuracy compared to traditional methods. Moreover, due to the ability to adaptively control the decoder structure, the detection efficiency of this method is comparable to that of traditional methods.

[0126] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A steel surface defect detection method based on an improved RT-DETR model, characterized in that: The following steps are involved: Step 1: Obtain a steel defect surface detection dataset, pre-process the image input, and divide it into a training set, a test set, and a validation set according to a preset ratio; Step 2: Build an improved RT-DETR model, including: Step 2.1: Construct a multi-scale feature enhancement module and stack it with the Bottleneck module to form a multi-scale residual connection backbone network. Step 2.2: Construct a hybrid query propagation mechanism to balance the supervision signals between different decoding layers by collecting and propagating query vectors from the middle layer. Step 2.3: Construct a dual hybrid strategy and dynamically adjust the decoder structure; Step 3: Input the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model; Step 4: Input the test set images to be detected into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

2. The steel surface defect detection method based on the improved RT-DETR model according to claim 1, characterized in that: As described in step 1, obtain the steel defect surface detection dataset, perform input preprocessing on the image, and divide it into training set, test set and validation set according to the preset ratio, as follows: Step 1.1: Obtain a dataset of steel defect surface detection, normalize the size of the original image, scale the image to 800 × 800 with any resolution, maintain the aspect ratio, and fill the blank area with 114 grayscale values; Step 1.2: Normalize the image pixels and perform z-score standardization with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225]. Step 1.3: Perform Mosaic data enhancement on the image; Step 1.4: Divide the preprocessed detection data set into training set, test set and validation set according to the preset ratio.

3. The steel surface defect detection method based on the improved RT-DETR model according to claim 2, characterized in that: The construction of the multi-scale feature enhancement module described in step 2.1 is as follows: (1) Use a depth-wise separable convolution with a convolution kernel of 7×7 to process the original input image and obtain the initial feature map; (2) Split the initial feature map into two branches, and directly input the feature map of one branch into a convolution layer with a convolution kernel of 1×1 for processing; The feature map of the other branch is sequentially input into a convolution layer with a convolution kernel of 1×1 and a ReLU activation function for processing; (3) Perform vector dot product operations on the feature maps processed by the two branches to achieve interaction and fusion between features; (4) The fused feature maps are sequentially input into the fully connected network, and refined by the depthwise separable convolution with a convolution kernel of 7×7 and the DropOut layer to obtain the final output features; (5) Finally, the Dropout layer outputs the refined feature map.

4. The steel surface defect detection method based on the improved RT-DETR model according to claim 3, characterized in that: The multi-scale residual connection backbone network described in step 2.1 consists of six stages connected in sequence: The first stage consists of an initial module, including a convolutional layer with a 3×3 kernel, 64 output channels, a stride of 2, and a ReLu activation function. The second and third stages each consist of a multi-scale feature enhancement module. The first module includes a depthwise separable convolution layer with a 7×7 convolution kernel, 256 output channels, and a stride of 1, a feature transformation layer with a dual-way convolution kernel of 1×1 and 768 output channels, a feature fusion layer based on element-wise multiplication, a feature compression layer with a 1×1 convolution kernel and 256 output channels, and a residual connection structure. The second module has 512 input channels and 1024 output channels, and its module structure is the same as that of the second stage. The fourth stage is composed of a stack of Bottleneck modules, which include a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 2, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure. The fifth stage consists of two stacked Bottleneck modules, both of which contain a dimensionality reduction convolution with a 1×1 convolution kernel and 256 output channels, a spatial convolution with a 3×3 convolution kernel, 256 output channels and a stride of 1, a channel expansion convolution with a 1×1 convolution kernel and 1024 output channels, and a residual connection structure; The sixth stage is composed of three stacked Bottleneck modules. The first module contains a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512, and a stride of 2, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure; the last two modules contain a dimensionality reduction convolution with a convolution kernel of 1×1 and an output channel number of 512, a spatial convolution with a convolution kernel of 3×3, an output channel number of 512, and a stride of 1, a channel expansion convolution with a convolution kernel of 1×1 and an output channel number of 2048, and a residual connection structure; An 800×800×3 image is input to the multi-scale residual connection backbone network. In the third stage, a 100×100×512 feature map is output. In the fifth stage, a 50×50×1024 feature map is output. In the sixth stage, a 25×25×2048 feature map is output to form a multi-scale defect feature pyramid. The 100×100×512, 50×50×1024, and 25×25×2048 feature maps are aligned to 256 channels by a convolution with a convolution kernel of 1×1 and then spliced. The spliced ​​vector is flattened into a vector of size 13125×256 as the input of the Transformer encoder.

5. The steel surface defect detection method based on the improved RT-DETR model according to claim 4, characterized in that: The hybrid query propagation mechanism described in step 2.2 is constructed as follows: (1) Input the initial query vector q0 into the first layer of the decoder D1 to generate the intermediate feature set D1(q0); (2) Merge q0 and D1(q0) to form the first-layer output set O1 = q0 ∪ {D1(q0)}; (3) For the decoder layer i D i , where i = 2, 3, ..., n, output O from the previous layer i-1 Extract the query vector q i-1 , change q i-1 Input D i , generating set D i (q i-1 ), merge q i-1 With D i (q i-1 ) Get the current layer output O i =O i-1 ∪{D i (q)}; (4) The sixth layer output O6 is used as the result of the entire decoder; The formula of the hybrid query propagation mechanism is as follows: O0={q0} O i =O i-1 ∪{D i (q)|q∈O i-1 } Among them, D i represents the i-th layer of the decoder, q i represents the query after being processed by the i-th layer of the decoder, O i represents the output of the i-th layer of the decoder.

6. The steel surface defect detection method based on the improved RT-DETR model according to claim 5, characterized in that: The dual hybrid strategy described in step 2.3 includes a hybrid strategy in the training round dimension and a hybrid strategy in the decoder layer dimension. That is, in the training process with a total of 200 rounds, the conventional query propagation mechanism is applied to all layers of the six-layer decoder for the first 150 rounds of training; and the hybrid query propagation mechanism is switched to the last four layers of the six-layer decoder for the next 50 rounds of training.

7. The steel surface defect detection method based on the improved RT-DETR model according to claim 6, characterized in that: In the conventional query propagation mechanism, the initial query vector passes through the six layers of decoders in sequence. The i-th layer decoder is based on the input query q i-1 Generate output query q i , i=2,3,…,n, and finally the output q6 of the sixth layer is used as the result of the entire decoder.

8. A steel surface defect detection system based on an improved RT-DETR model, characterized in that: The system is used to implement the steel surface defect detection method based on the improved RT-DETR model according to any one of claims 1 to 7, and the system includes a first module to a fourth module, wherein: The first module obtains a steel defect surface detection dataset, pre-processes the images, and divides them into training, test, and validation sets according to a pre-set ratio. The second module is to build an improved RT-DETR model, including: The first sub-unit builds a multi-scale feature enhancement module, which is stacked and connected with the Bottleneck module to form a multi-scale residual connection backbone network; The second sub-unit builds a hybrid query propagation mechanism, which balances the supervision signals between different decoding layers by collecting and propagating query vectors in the middle layer; The third subunit builds a dual hybrid strategy and dynamically adjusts the decoder structure; The third module inputs the preprocessed training set images into the improved RT-DETR model for training to obtain the trained improved RT-DETR model; In the fourth module, the test set images to be detected are input into the trained improved RT-DETR model to obtain the detection results of steel surface defects.

9. A mobile terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steel surface defect detection method based on the improved RT-DETR model as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the steel surface defect detection method based on the improved RT-DETR model are implemented.

Citation Information

Cited By

  • Station building facility disease detection method and system based on hybrid model

    CN121505356A

  • A self-evolving visual inspection method and system for copper-based stripes based on improved DETR

    CN122415619A