Steel Surface Defect Detection Method, Device and Equipment Based on Deep Learning
The SH-DETR model addresses gradient vanishing and long-range dependency issues in deep learning for steel defect detection, achieving efficient and accurate defect identification through enhanced feature extraction and fusion.
Patent Information
- Application Number
- CN202510643083.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The prior art has insufficient detection accuracy, is sensitive to environmental factors and has a large amount of calculation in the detection of steel surface defects, making it difficult to effectively identify complex steel surface defects.
Using a deep learning-based method, image features are extracted through the backbone network, channel shuffling and self-attention coding are combined with SH-encoder, weighted feature fusion module and IoU-perceived query mechanism, and finally defect detection is used using Transformer decoder.
The calculation amount has been greatly reduced, and the detection speed and accuracy have been improved, especially the detection ability of steel surface defects in complex backgrounds has been significantly improved.
Smart Images

Figure CN120164081B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of steel surface defect detection, and particularly to a steel surface defect detection method, device, and equipment based on deep learning. Background Art
[0002] Steel, as a basic raw material in industrial production, plays a crucial role. However, in production practice, defects may occur on the steel surface due to corrosion or deformation, making the detection of steel surface defects particularly critical. Common defect types include scratches and inclusions, etc., and these defects may seriously affect the performance of steel. Steel is widely used not only in fields such as machining, but also the close relationship between its quality and surface defects is directly related to the safety and reliability of products. With the development of society, the requirements for the quality of industrial materials are constantly increasing, and material quality has become the focus of attention. Traditional steel surface defect detection methods are mainly divided into two categories: manual detection and automatic detection. Manual detection methods have many limitations. Especially in large-scale production, their efficiency and reliability are difficult to meet the needs of modern industry. Automatic detection methods rely on efficient algorithms. Although early technologies such as threshold segmentation have improved the detection efficiency to a certain extent, they are sensitive to environmental factors and it is difficult to achieve the required detection accuracy. Therefore, how to accurately identify steel surface defects and how to achieve efficient detection of steel have become urgent problems to be solved in the industrial field.
[0003] Driven by deep learning, many innovative methods have emerged in the field of industrial defect detection. In the existing technical achievements, a multi-level feature fusion network (MFN) is proposed to achieve the fusion of multi-layer features. This operation not only integrates rich semantic information, but also highlights the region of interest through the region proposal network (RPN), thus promoting the research on steel defect detection. In the field of steel surface defect detection, due to the complexity of defects, it is necessary to use multi-layer neural networks for feature extraction. However, as the number of network layers increases, the problems of gradient disappearance and overfitting become the main obstacles restricting the performance of deep learning models. Traditional convolutional neural networks (CNNs) lack the ability to capture long-range dependencies in visual tasks, while the long-range dependency ability of Transformer provides the key to solving this problem. However, directly applying Transformer to the steel defect detection task is not appropriate. The existing research results provide new ideas and technical means for the field of industrial defect detection. However, there are certain limitations in the existing results in terms of detection accuracy and position information. Summary of the Invention
[0004] Based on this, it is necessary to provide a steel surface defect detection method, device, and equipment based on deep learning for the above technical problems.
[0005] A steel surface defect detection method based on deep learning, the method comprising:
[0006] Obtain a steel surface image.
[0007] Use a backbone network to extract features from the steel surface image to obtain image features.
[0008] After performing two downsamplings on the image features, use an SH-encoder for encoding to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffling, address encoding, and self-attention mechanism.
[0009] Fuse the image features, the first downsampled features of the image features, and the encoded features using a weighted feature fusion module to obtain fused features; the weighted feature fusion module is used to fuse the input features using SimAM, concatenation operation, and convolution operation.
[0010] Process the fused features using an IoU-aware query mechanism to obtain initial target query features.
[0011] Decode the initial target query features using a Transformer decoder with an auxiliary prediction head to obtain the detection result of the steel surface defect.
[0012] A steel surface defect detection device based on deep learning, the device comprising:
[0013] A steel surface image acquisition module for obtaining a steel surface image.
[0014] A feature extraction module for extracting features from the steel surface image using a backbone network to obtain image features.
[0015] A feature encoding module for encoding the image features after two downsamplings using an SH-encoder to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffling, address encoding, and self-attention mechanism.
[0016] A feature fusion module for fusing the image features, the first downsampled features of the image features, and the encoded features using a weighted feature fusion module to obtain fused features; the weighted feature fusion module is used to fuse the input features using SimAM, concatenation operation, and convolution operation.
[0017] A decoding module for processing the fused features using an IoU-aware query mechanism to obtain initial target query features; decoding the initial target query features using a Transformer decoder with an auxiliary prediction head to obtain the detection result of the steel surface defect.
[0018] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0019] The above method, device, and equipment for detecting steel surface defects based on deep learning. The method includes: extracting features from a steel surface image using a backbone network, performing two downsamplings on the extracted image features, and then encoding them using an SH-encoder to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffle, address encoding, and self-attention mechanism; fusing the image features, the first downsampling features of the image features, and the encoded features using a weighted feature fusion module to obtain fused features; the weighted feature fusion module is used to fuse the input features using SimAM, concatenation operation, and convolution operation; processing the fused features using an IoU-aware query mechanism to obtain initial target query features; decoding the initial target query features using a Transformer decoder with an auxiliary prediction head to obtain the detection results of steel surface defects. This method greatly reduces the computational amount, improves the computational speed, and enhances the accuracy of steel surface defect detection. Brief Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of a method for detecting steel surface defects based on deep learning in an embodiment;
[0021] Figure 2 It is a structural diagram of a method for detecting steel surface defects based on deep learning in another embodiment;
[0022] Figure 3 It is a schematic diagram of the channel shuffle process in another embodiment;
[0023] Figure 4 It is a data sample diagram in an embodiment, where (a) is a sample diagram of a pit, (b) is a sample diagram of an inclusion, (c) is a sample diagram of a patch, (d) is a sample diagram of a pitted surface, (e) is a sample diagram of rolling scale, and (f) is a sample diagram of a scratch;
[0024] Figure 5 It is a diagram of the number of samples of each category in another embodiment;
[0025] Figure 6 It is an experimental result diagram on the NEU dataset in another embodiment, where (a) is a schematic diagram of the training GIoU loss, (b) is a schematic diagram of the training L1 loss, (c) is a schematic diagram of the accuracy metric, (d) is a schematic diagram of the recall metric, (e) is a schematic diagram of the validation GIoU loss, (f) is a schematic diagram of the validation L1 loss, (g) is a schematic diagram of the mAP50 metric, and (h) is a schematic diagram of the mAP@0.5:0.95 metric;
[0026] Figure 7 It is a confusion matrix diagram in another embodiment;
[0027] Figure 8 It is an experimental result diagram on the GC10-DE dataset in one embodiment, where (a) is a schematic diagram of the training GIoU loss, (b) is a schematic diagram of the training L1 loss, (c) is a schematic diagram of the accuracy metric, (d) is a schematic diagram of the recall metric, (e) is a schematic diagram of the validation GIoU loss, (f) is a schematic diagram of the validation L1 loss, (g) is a schematic diagram of the mAP50 metric, and (h) is a schematic diagram of the mAP@0.5:0.95 metric;
[0028] Figure 9 It is an internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0030] In one embodiment, as Figure 1 shown, a steel surface defect detection method based on deep learning is provided, and the method includes the following steps:
[0031] Step 100: Obtain a steel surface image.
[0032] Specifically, the steel surface image is a steel surface defect image. The types of steel surface defects include: cracks, patches, inclusions, pitted surface, crazing, and scratches.
[0033] Step 102: Extract features from the steel surface image by using a backbone network to obtain image features.
[0034] Specifically, a backbone network based on a convolutional neural network is used to extract features from the steel surface image to obtain image features.
[0035] Step 104: Downsample the image features twice and then encode them by using an SH-encoder to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffling, address encoding, and self-attention mechanism.
[0036] Specifically, the SH-encoder uses channel shuffle followed by a layer of positional encoding and self-attention mechanism to specifically process the last layer of CNN features (i.e., image features) output by the backbone network. Traditional convolutional operations are usually limited within the channel groups they are assigned to, restricting the information interaction between different channels. Through channel shuffle, the channel order is rearranged, promoting the information exchange between different channel groups, thereby enhancing the feature extraction ability. This innovation enables the model to effectively learn and represent defect features while reducing the number of parameters and computational volume. The feature form of the feature map is pulled into a high-dimensional vector and then handed over to the encoder for processing, that is, the multi-scale features are flattened into a sequence (RB×C×H×W→RB×N×C), spliced into a vector with a long sequence, and then the self-attention technology is used for multi-scale feature interaction. The self-attention mechanism focuses more on the semantics of features. Compared with the shallower features in the CNN network, the deeper features have richer and more advanced semantic information. Therefore, the SH-encoder part only processes high-order features, which not only greatly reduces the computational volume and improves the computational speed, but also does not damage the performance.
[0037] Aiming at the problems of large computational volume and gradient descent in traditional convolution, the SH-encoder proposes a method of connecting shuffled channels in grouped convolution to better capture and retain local feature details. While ensuring the improvement of accuracy and model robustness, it reduces the complexity and the number of parameters of the model. The decoder usually has the characteristics of high computational complexity and low efficiency. The multi-head segmentation module is adopted to divide the feature map into different modules, and then the shuffled feature vectors are input into the encoder and the self-attention module for connection of mutual relationships and embedding of position information, which not only enhances the global attention, but also reduces the original computational volume to a linear level. In addition, a channel shuffle module is introduced to solve the problem that the positions of image encodings in the channels are fixed and the connections are not close enough, which greatly improves the efficiency and accuracy of the visual encoder when processing complex visual tasks.
[0038] By combining ResNet18 with the global self-attention mechanism, the feature extraction ability of the model is enhanced. Secondly, an innovative improvement is made to the encoder, and a multi-channel shuffle coding module is introduced, which reduces the number of parameters and effectively encodes different features through the interaction of different convolutional groups.
[0039] Step 106: The image features, the first downsampled features of the image features, and the encoded features are fused by a weighted feature fusion module to obtain fused features; the weighted feature fusion module is used to fuse the features input into it by using SimAM, splicing operation, and convolutional operation.
[0040] Specifically, aiming at the complexity of steel surface defects and the density of small target defects, a weighted feature fusion module is proposed. In the weighted feature fusion module, the UPC-SimAM cross-scale feature fusion module, convolution operation, and splicing operation are used to optimize the feature processing in the backbone network. Although the deep network can extract rich global features, the redundancy of its deep channel information leads to an increase in the model size and reduces the attention to the detection of dense small targets. In addition, traditional feature fusion methods, such as simple addition or superposition, do not yield ideal results. Therefore, the UPC-SimAM cross-scale feature fusion module uses a weighted fusion strategy to make the model pay more attention to the target area, improve the feature fusion efficiency, and reduce the redundancy of noise features.
[0041] In existing models, a common practice is to first adjust the scale of the feature map through upsampling (Upsample), and then fuse the features of different branches or different scales through concatenation (Concat), which often leads to a large amount of useless feature redundancy and may mask some features that are not significant but important. The UPC-SimAM cross-scale feature fusion module adopts a weighted fusion mechanism to integrate feature maps from different scales. At each fusion node, multiple input feature maps are fused by weighted summation, and the weighting coefficients are automatically learned by the network and used as learnable parameters. These parameters control the contribution degrees of features at different scales, ensuring that more important or more discriminative features can obtain larger weights during the fusion process, so as to more effectively capture the multi-scale information of the target.
[0042] Step 108: Process the fused features using an IoU-aware query mechanism to obtain initial target query features.
[0043] Specifically, through the IoU-aware query mechanism, a certain number of image features are selected from the output sequence of feature fusion as the initial target query of the decoder.
[0044] Step 110: Decode the initial target query features using a Transformer decoder with an auxiliary prediction head to obtain the detection results of steel surface defects.
[0045] Specifically, the decoder is equipped with an auxiliary prediction head to iteratively optimize the target query, thereby generating accurate bounding boxes and confidence scores, effectively monitoring the defects, and thus improving the comprehensive performance of the detector. This efficient model significantly improves the detection accuracy of small defects.
[0046] The core architecture of the steel surface defect detection model based on deep learning (abbreviation: SH-DETR) consists of four main parts, including: backbone network, SH-encoder, weighted feature fusion module, IoU-aware query mechanism, and Transformer decoder with auxiliary prediction heads. These components together constitute the overall framework of the SH-DETR model, ensuring its efficient performance and accurate object detection ability. The specific architecture of the steel surface defect detection model based on deep learning is as Figure 2 shown.
[0047] The SH-DETR model is an efficient and powerful single-stage object detection framework that integrates advantages such as real-time performance, accuracy, and stability. Compared with early models, the SH-DETR model achieves real-time end-to-end object detection without post-processing steps, thus maintaining a consistent speed during the inference process without introducing additional latency. In addition, the model also introduces an IoU-aware query mechanism, significantly improving the model performance and providing a more effective method for the initialization of object queries.
[0048] This method aims to solve the complexity of steel surface defects and the challenge of multi-scale feature extraction. To address this problem, a multi-scale feature extraction module is designed, which combines Transformer and CNN and uses different-sized convolutional kernels to extract features at different scales. In addition, an effective channel shuffle encoder module is proposed to solve the problems of feature disappearance and insufficient interaction between features. The introduction of the UPC-SimAM cross-scale feature fusion module further enhances the model's feature fusion ability, while SimAM improves the performance of the convolutional neural network through the method of ability weighting. This application also deepens the backbone network to further enhance the model's feature extraction ability. Ablation experiments on the publicly available NEU-DET and GC10-DE databases verify the effectiveness of this method, and comparisons with multiple state-of-the-art (SOTA) object detection models also prove the advantages of this method. Experimental results show that this method has achieved remarkable performance in terms of mAP@0.5 and mAP@0.5:0.95 metrics.
[0049] In the above-mentioned steel surface defect detection method based on deep learning, the method includes: using a backbone network to extract features from the steel surface image, performing two downsamplings on the extracted image features, and then encoding them using an SH-encoder to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffle, address encoding, and self-attention mechanism; using a weighted feature fusion module to fuse the image features, the first downsampling features of the image features, and the encoded features to obtain fused features; the weighted feature fusion module is used to fuse the input features using SimAM, concatenation operation, and convolution operation; processing the fused features using an IoU-aware query mechanism to obtain initial target query features; decoding the initial target query features using a Transformer decoder with an auxiliary prediction head to obtain the detection result of the steel surface defects. This method greatly reduces the computational amount, improves the computational speed, and enhances the accuracy of steel surface defect detection.
[0050] In one embodiment, the backbone network in step 102 is a convolutional neural network with a Resnet architecture.
[0051] Specifically, the backbone network is responsible for feature extraction and is mainly composed of a ConvBN module and a basic block module; among them, the ConvBN module includes a convolutional layer and a batch normalization layer, effectively expanding the receptive field of the network; the basic block module is based on the ResNet architecture and adopts a design of two-layer convolution and residual connection, effectively alleviating the problem of gradient disappearance and enhancing the expression ability and performance of the model.
[0052] In one embodiment, the SH-encoder includes: a channel shuffle module, a dynamic position encoding module, and a Transformer self-attention mechanism; step 104 includes: performing two downsamplings on the image features and then inputting them into the channel shuffle module to rearrange the channel order to obtain a feature vector; inputting the feature vector into the dynamic position encoding module to obtain position-encoded features; inputting the position-encoded features into the Transformer self-attention mechanism to obtain encoded features.
[0053] Specifically, the image features are downsampled twice and then input into the channel shuffle module to rearrange the channel order, obtaining an embedding vector. The embedding vector is passed through a linear layer to generate three vectors: Query, Key, and Value. For each input, the similarity between its query and all keys (through dot product) is calculated to generate attention weights, enabling each element in the input sequence to focus on elements in other positions, thereby capturing the dependencies within the sequence. After these weights are normalized (using Softmax), they are dot-producted with the corresponding Value vectors to generate the output vector. Since there is no convolutional or recurrent structure in the SH-encoder to capture the sequence order, explicit positional information needs to be added. Through Positional Encoding, the position of each element in the input sequence is added to its corresponding embedding vector. At each position, a learned position embedding vector is appended to the sequence. The multi-head attention captures different dependencies through different heads, and the outputs of the multi-heads are concatenated and then processed through a linear layer to obtain the final attention output. Finally, the output is adjusted back to two dimensions, denoted as F5, for subsequent cross-scale feature fusion. Through this mechanism, the self-attention in the SH-encoder can capture the global dependencies in the sequence without relying on the sequence order, thereby improving the robustness and generalization ability of the model.
[0054] The channel shuffle process is as Figure 3 shown in Figure 3 where the gated convolution is: Gated Convolution, abbreviated as: GConv.
[0055] In one embodiment, the weighted feature fusion module includes: a UPC-SimAM cross-scale feature fusion module, a first convolutional module, and a second convolutional module; step 106 includes: inputting the encoded features into the first first convolutional module to obtain first convolutional features; inputting the first convolutional features and the first downsampled features of the image features into the first UPC-SimAM cross-scale feature fusion module to obtain first fusion features; inputting the first fusion features into the second first convolutional module to obtain second convolutional features; inputting the second convolutional features and the image features into the second UPC-SimAM cross-scale feature fusion module to obtain second fusion features; inputting the second fusion features into the first second convolutional module to obtain third convolutional features; inputting the third convolutional features and the second convolutional features into the third UPC-SimAM cross-scale feature fusion module to obtain third fusion features; inputting the third fusion features into the second second convolutional module to obtain fourth convolutional features; inputting the fourth convolutional features and the first convolutional features into the fourth UPC-SimAM cross-scale feature fusion module to obtain fourth fusion features; splicing the second fusion features, the third fusion features, and the fourth fusion features to obtain fusion features.
[0056] In one embodiment, the UPC-SimAM cross-scale feature fusion module includes an upsampling layer and a SimAM module; inputting the first convolutional features and the first downsampled features of the image features into the first UPC-SimAM cross-scale feature fusion module to obtain first fusion features includes: upsampling the first convolutional features through the upsampling layer and then splicing them with the first downsampled features of the image features to obtain spliced features; inputting the spliced features into the SimAM module to obtain first fusion features.
[0057] Specifically, the UPC-SimAM cross-scale feature fusion module utilizes the characteristic that the similarity between adjacent pixels in an image is strong and the similarity between distant pixels is weak, and generates attention weights by calculating the similarity between each pixel in the feature map and its adjacent pixels. Different from existing channel / space domain attention modules, SimAM is a lightweight attention module that improves the performance of convolutional neural networks in a simple way. It does not need to introduce additional parameters or expensive computational overheads, but can still effectively capture important feature information. It enhances the performance of CNN by calculating the local self-similarity of the feature map. The operation of this module is mainly based on the selection of the defined energy function, avoiding excessive structural adjustments and deriving 3D attention weights for the feature map without additional parameters. It uses binary labels and adds a regular term, and uses the energy of each pixel in the feature map to judge its contribution to the model task. Specifically, the minimum energy can be calculated by the following formula:
[0058] ;
[0059] Among them, is the minimum energy, is the regularization term, is the k th neuron on a single channel of the input feature map. is the mean of all neurons on a single channel, is the variance of all neurons on a single channel. The smaller the value, the lower the energy, and the greater the difference between the neuron k and its surrounding neurons, which is also more important for visual processing.
[0060] The output feature map is:
[0061] ;
[0062] Among them, E is the set of all neurons, X is the input feature map, Sigmoid is the activation function, whose purpose is to limit the E value and enhance the features. Through this mechanism, the UPC-SimAM cross-scale feature fusion module can capture the global dependencies in the sequence without relying on the sequence order, improving the robustness and generalization ability of the model.
[0063] The UPC-SimAM cross-scale feature fusion module then uses features at different levels for feature fusion. The fusion part consists of two convolutional layers with a kernel size of 1×1 and multiple convolutional layers with a kernel size of 3×3, giving full play to the integration advantages of features at different scales. The encoded feature vectors are integrated into the internal scale feature interaction module, enhancing the network's ability to capture global dependencies and complex local details in the image. The UPC-SimAM cross-scale feature fusion module fuses the feature maps output by other modules while upsampling. These feature maps carry specific position information, which is beneficial for more effective feature extraction and multi-frequency fusion. In addition, the UPC-SimAM cross-scale feature fusion module introduces a dynamically adjusted offset to highlight the key features, thereby achieving precise localization of the defect area and greatly improving the efficiency and accuracy of subsequent decoding for complex visual tasks. The multi-feature fusion strategy fuses the initial feature maps from shallow to deep, enhancing the different attentions of the feature maps at each stage while retaining the detailed features to improve the recognition of tiny features.
[0064] In one embodiment, the first convolutional module includes: a convolutional layer with a kernel size of , a batch normalization layer, and a SiLU activation function; in one embodiment, the second convolutional module includes: a convolutional layer with a kernel size of , a batch normalization layer, and a SiLU activation function.
[0065] In a specific embodiment, the loss function of the steel surface defect detection model used in the steel surface defect detection method proposed in this application is the L1 loss or the GIoU loss.
[0066] The L1 loss, also known as the Mean Absolute Error (MAE for short), is widely used in regression tasks. This loss function measures the error by calculating the sum of the absolute differences between the model prediction values and the actual values. Compared with the L2 loss (i.e., the Mean Squared Error MSE), the L1 loss shows higher robustness to outliers because it is less sensitive to extreme values in the data. This characteristic makes the L1 loss often provide more robust performance when dealing with datasets containing outliers. The L1 loss is:
[0067] ;
[0068] where, are the actual value and the corresponding predicted value respectively, is the number of training samples.
[0069] Since the L1 loss calculates the absolute difference rather than the squared difference between the predicted value and the true value, it has relatively less impact on outliers. This characteristic makes the L1 loss more robust in the face of extreme values in the data. The gradient of the L1 loss is constantly ±1, which may lead to less frequent gradient updates in some optimization algorithms, resulting in sparse updates. However, the L1 loss directly reflects the average magnitude of the prediction error, and its unit is consistent with the original data, which provides good interpretability for it. In tasks such as image denoising, the L1 loss helps to maintain the accuracy of image pixel values and avoids over-punishing large errors.
[0070] The Generalized Intersection over Union (GIoU) loss is an advanced loss function in object detection tasks. Compared with the standard Intersection over Union (IoU) loss, the GIoU loss performs better in optimizing the bounding box localization accuracy. Especially when there is no overlap between bounding boxes, it can still provide effective optimization signals. The application of this loss function in the object detection model significantly improves the performance of the model in bounding box regression.
[0071] The IoU calculation formula is:
[0072] ;
[0073] where, IoU is the standard intersection over union, is the area of the intersection between the predicted bounding box and the ground truth bounding box. is the area of the union between the predicted bounding box and the ground truth bounding box.
[0074] Minimum enclosing box C: Calculate the area of the smallest bounding rectangle region that contains both the predicted bounding box and the ground truth bounding box (referred to as the enclosing box area). The GIoU calculation formula is:
[0075] ;
[0076] where, is the Generalized Intersection over Union, is the area of the minimum enclosing box C.
[0077] GIoU Loss:
[0078] ;
[0079] where, is loss.
[0080] The value range of the GIoU loss is [-1, 1]. When the two bounding boxes completely overlap, the value of the GIoU loss is 1; while when the two bounding boxes do not intersect at all, the value of the GIoU loss is less than 0. Therefore, the value range of the GIoU loss is between [0, 2], and the smaller the value, the higher the matching degree between the predicted bounding box and the ground truth bounding box. Compared with the standard IoU loss, the latter cannot generate effective gradients when the bounding boxes do not intersect, which limits the learning ability of the model. The GIoU loss can provide an optimization signal even when the bounding boxes do not intersect, thus promoting the model to converge more quickly. Compared with the IoU loss, the GIoU loss provides more stable gradient updates, which enables the model to more accurately regress the target position. The GIoU loss significantly improves the accuracy of the model in bounding box prediction by reducing the localization error.
[0081] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0082] In a validation example, the efficiency of the SH-DETR model proposed in this application was comprehensively evaluated using two different benchmark datasets. First, the used benchmark datasets and experimental conditions were described in detail to ensure the transparency and reproducibility of the evaluation process. To verify the effectiveness of the proposed SH-DETR model on the test dataset, its performance was compared with several existing methods. In terms of performance evaluation metrics, key metrics including recognition accuracy, recall rate, mAP@0.5, and mAP@0.5:0.95 were adopted, which can comprehensively reflect the detection performance of the model. Finally, the transfer learning method was used to further optimize the model, and the results showed that the model optimized by transfer learning had a significant improvement in accuracy. These results not only demonstrated the effectiveness of the SH-DETR model but also showed its potential and flexibility in object detection tasks.
[0083] (1) Training equipment and test process
[0084] In this example, the SH-DETR model was first trained on the training set and then evaluated on the test set to verify its reliability. In the case where the dataset size is small and cannot be fully divided into training set, validation set, and test set, an 80%-20% data splitting strategy is usually adopted, that is, 80% of the data is used for model training and the remaining 20% of the data is used for model validation.
[0085] This example will show the experimental results on the NEU-DET dataset. To comprehensively evaluate the performance of the model, key metrics such as recognition accuracy, recall rate, mAP@0.5, and mAP@0.5:0.95 were adopted to verify the reliability of the model. These metrics can comprehensively reflect the performance of the model in object detection tasks. In addition, the SH-DETR model was implemented on a computer configured with an Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz, 2.39 GHz (dual processors), and 128 GB of memory, using a Python environment and the PyTorch framework. This experimental environment ensured the efficiency and stability of the model training and testing processes.
[0086] (2) Dataset details
[0087] The NEU-DET dataset is designed specifically for steel surface defect detection and is widely used in the fields of machine vision and deep learning for the development and evaluation of various defect detection algorithms. The NEU-DET dataset used in this study contains 1,800 steel surface defect images, each with a resolution of 200 pixels × 200 pixels. This dataset covers 6 different types of steel surface defects, including Cracks, Patches, Inclusions, Pitted Surface, Crazing, and Scratches.
[0088] Examples of steel surface defect types are as Figure 4 shown, where Figure 4 (a) in it is an example diagram of crazing, Figure 4 (b) in it is an example diagram of inclusions, Figure 4 (c) in it is an example diagram of patches, Figure 4 (d) in it is an example diagram of pitted surface, Figure 4 (e) in it is an example diagram of rolling scale, Figure 4 (f) in it is an example diagram of scratches.
[0089] The GC10-DET dataset is a publicly available industrial defect detection dataset that specifically collects and annotates typical defect samples on the steel surface. This dataset contains 10 different types of defects, including Punching (Pu), Weld (Wl), Crescent Gap (Cg), Water Spot, Oil Spot (Os), Silk Spot (Ss), Inclusion (In), Rolling Pit (Rp), Crease (Cr), and Waist Crease (Wf). Each defect presents different shapes, sizes, and complexities, comprehensively covering common steel defect situations in industry. The GC10-DET dataset is renowned for its high image quality and fine annotations, so it is often used as a standard dataset when studying steel defect detection and related deep learning and computer vision methods. GC10-DET provides defect annotation information with bounding boxes, facilitating the training and evaluation of models. Given that the dataset contains various types and shapes of defects, this dataset has high practical value in designing and testing defect detection algorithms, especially suitable for verifying object detection models based on deep learning. The number of samples in each category is as Figure 5 shown.
[0090] (3)Experimental Results
[0091] 1) Model Performance (NEU-DET Dataset)
[0092] Due to the influence of light changes and material differences, the defect images in the NEU-DET dataset show variations in grayscale, resulting in significant differences in the appearance of defects within the same category (intra-class), while similar characteristics are exhibited among defects of different categories (inter-class). These characteristics are both challenges and opportunities for the model. They contribute to more precise validation on the NEU-DET dataset and endow the model with important value in real-world applications. During the training process, the trends of loss and accuracy change significantly: in the initial dozens of iterations, the loss value drops significantly and the accuracy increases substantially. Subsequently, as the iterations continue, the loss gradually stabilizes and the accuracy increases slightly until the model converges. Finally, the precision of the model reaches 0.9172 and the recall rate reaches 0.7844.
[0093] As shown in Table 1, the model has achieved satisfactory accuracy in six different categories of defects. The accuracy of the Inclusions category is the lowest, at 0.703. In terms of the recall rate, the Patches category ranks first with a score of 0.93, and the other categories also achieve good results, but the recall rate of the Crazing category is the lowest, only 0.239. This may be because the color of the Crazing is relatively light and it is easy to blend in with the background, resulting in a low recall rate. In addition, the impact of datasets of different scales on the model performance cannot be ignored. From the mAP metric, the mAP@0.5 of Patches, Pitted Surface, and Scratches exceeds 0.9, and the overall mAP@0.5 reaches 0.8303, while mAP@0.5:0.95 reaches 0.455. Due to the high contrast between cracks and the background, their precision is relatively better. The confusion matrix further shows that although the model performs well in category classification, it is still affected by subtle differences in the background, which precisely proves the superiority of the model and its sensitivity to details. The experimental results are as Figure 6 shown, where Figure 6 (a) in it is the schematic diagram of the training GIoU loss, Figure 6 (b) in it is the schematic diagram of the training L1 loss, Figure 6 (c) in it is the schematic diagram of the accuracy metric, Figure 6 (d) in it is the schematic diagram of the recall rate metric, Figure 6 (e) in it is the schematic diagram of the validation GIoU loss, Figure 6 (f) in it is the schematic diagram of the validation L1 loss, Figure 6 (g) in it is the schematic diagram of the mAP50 metric, Figure 6 (h) in it is the schematic diagram of the mAP@0.5:0.95 metric. The confusion matrix is as Figure 7 shown.
[0094] Table 1 Recognition Results of Six Different Types of Defects
[0095]
[0096] 2) Model Performance (GC10-DE Dataset)
[0097] To further verify the performance of the model, experiments were conducted on the GC10-DET dataset. GC10-DET is a surface defect dataset collected from an actual industrial environment, containing 3570 images with a resolution of 2048 pixels × 1000 pixels, and divided into a training set and a validation set according to an 8:2 ratio. The experimental results are shown in Table 2. This method shows good effectiveness on the GC10-DET dataset, with an accuracy of 76.73% and a recall rate of 63.84%. The values of the two loss functions were reduced to 0.5795 and 0.3526 respectively. Since the GC10-DET dataset was collected from a real industrial environment, its accuracy may be slightly lower than that of other datasets, but it has higher reference value for actual steel surface defect detection. It can be seen from the table that good recognition accuracies were achieved for different types of defects in the model, especially for the waist crease category, with an accuracy as high as 0.922. In the mAP@0.5 metric, the accuracies of punching, weld seam, and crescent-shaped gap exceeded 90%, while the mAP@0.5 of inclusions, rolling pits, and creases were the lowest, at 28.3%, 24.4%, and 24.7% respectively. The overall mAP@0.5 of the categories reached 65.03%. The number of datasets has a significant impact on model performance. For the rolling pit and crease categories, due to the small number of samples, the accuracy is low. Punching and weld seam achieved the highest accuracies in mAP@0.5:0.95, exceeding 53%, and the average accuracy was 32%. It can be seen from the mAP@0.5:0.95 data that the accuracy has a profound impact on the background relationship. In real-life scenarios, steel surface defect detection is easily interfered by the background. Therefore, appropriately increasing the number of samples of certain categories and balancing the samples of each category can improve the generalization ability of the model. In the ten categories, this method can accurately identify the crack category and its location and distinguish it from the background, showing that the model has strong robustness and generalization ability. The experimental results on the GC10-DE dataset are as Figure 8 shown, where Figure 8 (a) in it is the schematic diagram of the training GIoU loss, Figure 8 (b) in it is the schematic diagram of the training L1 loss, Figure 8 (c) in it is the schematic diagram of the accuracy metric, Figure 8 (d) in it is the schematic diagram of the recall rate metric, Figure 8 (e) in it is the schematic diagram of the validation GIoU loss, Figure 8 (f) in it is the schematic diagram of the validation L1 lossFigure 8 In (g) is a schematic diagram of the mAP50 metric, Figure 8 In (h) is a schematic diagram of the mAP@0.5:0.95 metric.
[0098] Table 2 Defect recognition results of the GC10-DE dataset
[0099]
[0100] The comparison of the detection results of this method verified on the NEU-DET dataset and the GC10-DE dataset is shown in Table 3.
[0101] Table 3 Comparison of defect detection results on the NEU-DET dataset and the GC10-DE dataset
[0102]
[0103] 3) Comparison with sota
[0104] On the public NEU-DET dataset, a series of comparative experiments were carried out, and the performance comparison of multiple models (I) is shown in Table 4. In this embodiment, this method was compared with multiple state-of-the-art models, including ScaledYOLOv4-csp, YOLO-MSFE-EFF, YOLOv7-tiny, Mask-R-CNN, and YOLOv7. The experimental results show that this method exhibits excellent performance under both the mAP@0.5 and mAP@0.5:0.95 metrics. It can be seen from the data in Table 4 that the number of parameters of this method is only slightly higher than that of YOLOv7-tiny and YOLO-MSFE-EFF. Specifically, this method reached 83.30% and 45.55% on mAP@0.5 and mAP@0.5:0.95 respectively, which are 11.06% and 8.35% higher than the YOLOv7 model on the NEU-DET dataset. This method only contains 15.78M parameters and is more lightweight compared to other basic advanced models. On mAP@0.5 and mAP@0.5:0.95, this method is 1.83% and 10.25% higher than the Mask-R-CNN model respectively, and the number of parameters is 56.94M less than that of Mask-R-CNN. The number of parameters of YOLOv7-tiny and YOLO-MSFE-EFF increases in turn, and their mAP performance also improves accordingly. This method is not only superior to traditional YOLO models but also more economical in terms of the number of parameters.
[0105] Following the comparison of the NEU-DET dataset, additional tests were conducted on the GC10-DET dataset. The performance comparison of multiple models (Part II) is shown in Table 5. The detection effect of this method is better than current mainstream industrial object detection algorithms, including YOLOv5-S, YOLOv5-M, YOLOv5-L, YOLOv5-X, and the Mask-R-CNN model, achieving an accuracy of 76.73% and an mAP@0.5 of 65.03%. The YOLO series variants achieved an accuracy of 60.6% to 70.8% and an mAP@0.5 of 55.6% to 62.5% on the GC10-DET dataset. The Mask-R-CNN model achieved an mAP@0.5 of 65.3% and an mAP@0.5:0.95 of 33.5% on the GC10-DET dataset, with a recall rate of 61.3%. As shown in the data in Table 5, although this method surpasses the YOLO variant models in terms of accuracy and achieves comparable accuracy to the Mask-R-CNN on other datasets. This method achieves the accuracy of advanced models on other datasets with fewer parameters. These results indicate that this method maintains good performance on the GC10-DET dataset, demonstrating the robustness of this model.
[0106] Table 4 Performance Comparison of Multiple Models (Part I)
[0107]
[0108] Table 5 Performance Comparison of Multiple Models (Part II)
[0109]
[0110] 4) Ablation and Visualization Experiments
[0111] The contribution of each module in the model was verified through ablation experiments. The results of the ablation experiments are shown in Table 6. Model A is a basic model based on the ResNet18 backbone network combined with Transformer, with its mAP@0.5 and mAP@0.5:0.95 reaching 79.40% and 43.55% respectively. Model B introduced the SH encoder module proposed in this paper on the basis of Model A. Through single-channel convolution and channel shuffling techniques, not only 4M parameters were reduced, but also the interaction between different channels was enhanced, resulting in an increase of 2.24% and 0.67% in mAP@0.5 and mAP@0.5:0.95 compared to the basic model, while significantly reducing the computational amount of the model.
[0112] Model C is the complete model proposed in this application. It further integrates multi-feature fusion and parameter-free attention modules on the basis of Model B. This method of fusing features of different scales, combined with the parameter-free attention mechanism, not only controls the growth of the model's parameter quantity but also enhances the interaction ability between features, increasing mAP@0.5 and mAP@0.5:0.95 to 83.03% and 45.55% respectively. The data in the table shows that the lack of these key modules will lead to varying degrees of decline in the model's performance, thus confirming the importance of these modules for the overall model performance.
[0113] For the CNN-based Module E, in this embodiment, ResNet50 is selected as the backbone network for convolutional embedding for experiments. The results show that although the parameter quantity and computational amount of Model E increase significantly on the NEU-DET dataset, its accuracy improvement is not obvious. The excessive increase in parameter quantity and computational amount may limit the application of the model in real life. Therefore, in this embodiment, ResNet18 with fewer parameters and higher accuracy is selected as the backbone network.
[0114] Based on Model C, Model D is further added to explore whether increasing the number of encoders will improve the model's accuracy. The experimental results show that the LSwin Transformer block with a deep MLP has the best effect. Without the deep MLP, mAP@0.5 increases by 2.5% and mAP@0.5:0.95 increases by 1.9%. In addition, the attention patch merging and convolutional embedding modules increase mAP@0.5 by 1.2% and 1.3% respectively. These results further confirm the effectiveness of this method and provide valuable references for future improvements.
[0115] Table 6 Ablation Experiment Results
[0116]
[0117] After analyzing the accuracy of each module, a comparative analysis is further carried out by visualizing the detection boxes and categories on the NEU-DET dataset. This method can not only accurately identify defect categories such as cracks, but also has a relatively high recognition accuracy. Taking crazing as an example, when compared with the original image, the basic model can already identify the defect category, but after introducing the SH-encoder module, the recognition range is expanded. In the complete model proposed in this application, after combining the UPC-SimAM cross-scale feature fusion module, the recognition range is further enlarged, indicating that the model's ability to handle interfering similar backgrounds has been enhanced. In the inclusion category, with the addition of different modules, the recognition of cracks is significantly improved, and the model can even identify small cracks that were not detected before, significantly improving the recognition ability of small targets, thus proving the accuracy of the model. For the patches category, the basic model has redundant and repeated detection boxes in target recognition, and the range score of the recognized categories is not high. However, the model proposed in this application can more accurately mark the position and size of the cracks, which obviously reflects the powerful performance of the modules proposed in this application in feature extraction and processing capabilities.
[0118] Generally speaking, these visualization results not only demonstrate the model's recognition ability for different defect categories, but also prove that by introducing the SH-encoder and UPC-SimAM cross-scale feature fusion modules, this method has significant advantages in dealing with complex backgrounds and tiny defects, thus improving the overall performance and accuracy of the model.
[0119] In one embodiment, there is provided a schematic flowchart device of a steel surface defect detection method based on deep learning, including: a steel surface image acquisition module, a feature extraction module, a feature encoding module, a feature fusion module, and a decoding module, where:
[0120] The steel surface image acquisition module is used to acquire a steel surface image;
[0121] The feature extraction module is used to extract features from the steel surface image by using a backbone network to obtain image features;
[0122] The feature encoding module is used to perform two downsamplings on the image features and then encode them by using an SH-encoder to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffling, address encoding, and self-attention mechanism;
[0123] The feature fusion module is used to fuse the image features, the first downsampling features of the image features, and the encoded features by using a weighted feature fusion module to obtain fused features; the weighted feature fusion module is used to fuse the input features by using SimAM, splicing operation, and convolution operation;
[0124] A decoding module, which is used to process the fused features by using an IoU-aware query mechanism to obtain initial target query features; and decode the initial target query features by using a Transformer decoder with an auxiliary prediction head to obtain the detection results of steel surface defects.
[0125] In one embodiment, the backbone network in the feature extraction module is a convolutional neural network with a Resnet architecture.
[0126] In one embodiment, the SH-encoder includes: a channel shuffle module, a dynamic position encoding module, and a Transformer self-attention mechanism; the feature encoding module is further configured to downsample the image features twice and then input them into the channel shuffle module to rearrange the channel order to obtain feature vectors; input the feature vectors into the dynamic position encoding module to obtain position-encoded features; and input the position-encoded features into the Transformer self-attention mechanism to obtain encoded features.
[0127] In one embodiment, the weighted feature fusion module includes: a UPC-SimAM cross-scale feature fusion module, a first convolutional module, and a second convolutional module; the feature fusion module is further configured to input the encoded features into the first first convolutional module to obtain first convolutional features; input the first convolutional features and the first downsampled features of the image features into the first UPC-SimAM cross-scale feature fusion module to obtain first fusion features; input the first fusion features into the second first convolutional module to obtain second convolutional features; input the second convolutional features and the image features into the second UPC-SimAM cross-scale feature fusion module to obtain second fusion features; input the second fusion features into the first second convolutional module to obtain third convolutional features; input the third convolutional features and the second convolutional features into the third UPC-SimAM cross-scale feature fusion module to obtain third fusion features; input the third fusion features into the second second convolutional module to obtain fourth convolutional features; input the fourth convolutional features and the first convolutional features into the fourth UPC-SimAM cross-scale feature fusion module to obtain fourth fusion features; and splice the second fusion features, the third fusion features, and the fourth fusion features to obtain fused features.
[0128] In one embodiment, the UPC-SimAM cross-scale feature fusion module includes an upsampling layer and a SimAM module; the feature fusion module is further configured to splice the first convolutional features after upsampling through the upsampling layer with the first downsampled features of the image features to obtain spliced features; and input the spliced features into the SimAM module to obtain first fusion features.
[0129] In one embodiment, the first convolutional module in the feature fusion module includes: a convolutional layer with a convolutional kernel of , a batch normalization layer, and a SiLU activation function; in one embodiment, the second convolutional module includes: a convolutional layer with a convolutional kernel of , a batch normalization layer, and a SiLU activation function.
[0130] For the specific limitations of the steel surface defect detection device based on deep learning, reference can be made to the limitations of the steel surface defect detection method based on deep learning in the above text, which will not be elaborated here. Each module in the above-mentioned steel surface defect detection device based on deep learning can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0131] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a decoding module method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0132] Those skilled in the art can understand that Figure 9 the structure shown in
[0133] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0137] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A steel surface defect detection method based on deep learning, characterized in that, The method includes: Obtain an image of the steel surface; Use a backbone network to extract features from the image of the steel surface to obtain image features; After performing two downsamplings on the image features, use an SH-encoder for encoding to obtain encoded features; the SH-encoder is used to encode the image features through channel shuffle, address encoding, and self-attention mechanism; Use a weighted feature fusion module to fuse the image features, the encoded features, and the first downsampling features of the image features to obtain fused features; the weighted feature fusion module is used to fuse the input features by using SimAM, concatenation operation, and convolution operation; Process the fused features by using an IoU-aware query mechanism to obtain initial target query features; Decode the initial target query features by using a Transformer decoder with an auxiliary prediction head to obtain the detection result of the steel surface defect.
2. The method according to claim 1, characterized in that, The backbone network is a convolutional neural network with a Resnet architecture.
3. The method according to claim 1, wherein The SH-encoder includes: a channel shuffle module, a dynamic position encoding module, and a Transformer self-attention mechanism; After performing two downsamplings on the image features and using the SH-encoder for encoding to obtain encoded features, it includes: After performing two downsamplings on the image features, input them into the channel shuffle module to rearrange the channel order to obtain a feature vector; Input the feature vector into the dynamic position encoding module to obtain position-encoded features Input the position-encoded features into the Transformer self-attention mechanism to obtain encoded features.
4. The method according to claim 1, characterized in that, The weighted feature fusion module includes: a UPC-SimAM cross-scale feature fusion module, a first convolution module, and a second convolution module; Using the weighted feature fusion module to fuse the image features, the encoded features, and the first downsampling features of the image features to obtain fused features, includes: Input the encoded features into the first first convolution module to obtain first convolution features; Input the first convolution features and the first downsampling features of the image features into the first UPC-SimAM cross-scale feature fusion module to obtain first fused features; Input the first fused features into the second first convolution module to obtain second convolution features; Input the second convolution features and the image features into the second UPC-SimAM cross-scale feature fusion module to obtain second fused features; Input the second fused features into the first second convolution module to obtain third convolution features; Input the third convolution features and the second convolution features into the third UPC-SimAM cross-scale feature fusion module to obtain third fused features; Input the third fused features into the second second convolution module to obtain fourth convolution features; Input the fourth convolution features and the first convolution features into the fourth UPC-SimAM cross-scale feature fusion module to obtain fourth fused features; Concatenate the second fusion feature, the third fusion feature, and the fourth fusion feature to obtain a fusion feature.
5. The method according to claim 4, characterized in that, The UPC-SimAM cross-scale feature fusion module includes an upsampling layer and a SimAM module; Input the first convolutional feature and the first downsampled feature of the image feature into the first UPC-SimAM cross-scale feature fusion module to obtain a first fusion feature, including: Upsample the first convolutional feature through an upsampling layer and concatenate it with the first downsampled feature of the image feature to obtain a concatenated feature; Input the concatenated feature into the SimAM module to obtain a first fusion feature.
6. The method according to claim 4, characterized in that The first convolutional module includes: a convolutional layer with a convolutional kernel of , a batch normalization layer, and a SiLU activation function.
7. The method according to claim 4, wherein The second convolutional module includes: a convolutional layer with a convolutional kernel of , a batch normalization layer, and a SiLU activation function.
8. A steel surface defect detection device based on deep learning, characterized in that, The device includes: A steel surface image acquisition module for acquiring a steel surface image; A feature extraction module for extracting features from the steel surface image using a backbone network to obtain an image feature; A feature encoding module for encoding the image feature after two downsamplings using an SH-encoder to obtain an encoded feature; the SH-encoder is used to encode the image feature through channel shuffle, address encoding, and self-attention mechanism; A feature fusion module for fusing the image feature, the encoded feature, and the first downsampled feature of the image feature using a weighted feature fusion module to obtain a fusion feature; the weighted feature fusion module is used to fuse the features input into it using SimAM, concatenation operation, and convolution operation; A decoding module for processing the fusion feature using an IoU-aware query mechanism to obtain an initial target query feature; decoding the initial target query feature using a Transformer decoder with an auxiliary prediction head to obtain a detection result of steel surface defects.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the deep learning-based steel surface defect detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A speed and precision balanced steel product surface defect detection method
CN113628178A
Steel surface defect detection method and system based on deep learning
CN118134877A