Steel surface defect identification method and device, computer equipment and storage medium

Through the backbone network, adaptive feature interaction encoder and dynamic pooling pyramid and GhostConv's feature fusion module, combined with IoU-aware query selection strategy and decoder with auxiliary prediction head, the shortcomings of traditional detection technology in environmental adaptability and anti-interference capabilities are solved, and the accurate identification and capture of full-scale defects is achieved.

CN120164082AActive Publication Date: 2025-06-17NANCHANG YANNUO TECH CO LTD +1

Patent Information

Application Number
CN202510643474.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Traditional steel surface defect detection technology has weak environmental adaptability and poor anti-interference ability, which makes it difficult to meet high-end manufacturing requirements, especially when dealing with complex defect types of multiple scales and multiple forms.

Method used

Feature extraction is performed using a backbone network, combining the adaptive feature interaction encoder and the feature fusion module of dynamic pooling pyramid and GhostConv, and the identification of steel surface defects is achieved through IoU-aware query selection strategy and a decoder with auxiliary prediction head.

Benefits of technology

While maintaining computing efficiency, the full-scale defect capture capability from macro deformation to micro cracks is achieved, which improves the accuracy and robustness of detection, and meets the needs of high-end manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164082A_ABST
    Figure CN120164082A_ABST
Patent Text Reader

Abstract

The invention relates to a steel surface defect identification method and device, computer equipment and a storage medium. According to the method, a backbone network is adopted to carry out feature extraction on a steel surface image, and down-sampling is carried out on the extracted image features twice; converting the second down-sampling feature into an image feature vector, and processing the image feature vector by using an adaptive feature interaction encoder to obtain an encoding feature; processing the image features, the coding features and the first down-sampling features by using a feature fusion module based on a dynamic pooling pyramid and GhostConv to obtain a fusion feature vector; screening a fixed number of image features from the fusion feature vector by adopting an IoU-perceived query selection strategy to obtain an initial query vector; and processing the initial query vector by adopting a decoder with an auxiliary prediction head to obtain a steel surface defect identification result. According to the method, the full-scale defect capturing capability from macroscopic deformation to microcosmic cracks is realized while the calculation efficiency is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, computer device and storage medium for identifying steel surface defects. Background Art

[0002] In the industrial production system, steel, as an indispensable basic material, its quality directly affects the safety and reliability of end products. It is worth noting that surface damage phenomena are likely to occur during the manufacturing process, mainly including typical defect types such as corrosion deformation, mechanical scratches and impurity embedding. These surface abnormalities will not only significantly weaken the mechanical properties of the material, but also pose potential risks of structural failure. Therefore, building an accurate surface defect detection system plays a key role in ensuring the quality of steel and has become an important quality control link that cannot be ignored in modern manufacturing processes. Driven by the dual upgrades of the industrial manufacturing field and the improvement of material quality standards, material quality control has risen to a key technical indicator. Traditional manual detection methods have gradually revealed the shortcomings of insufficient efficiency and lack of stability in large-scale production scenarios, and are difficult to meet the precision requirements of modern industry. Although intelligent detection technologies based on deep learning have achieved breakthrough progress in detection efficiency, traditional image processing algorithms represented by threshold segmentation still have inherent defects such as weak environmental adaptability and poor anti-interference ability, resulting in the robustness and accuracy of detection results being difficult to meet the requirements of high-end manufacturing.

[0003] Aiming at the problems of multi-scale and multi-form detection commonly existing in the steel surface defect detection scenario, including complex defect types such as slender scratches, irregular oxidation patches and dispersed inclusions, traditional convolutional neural networks (CNNs) have inherent defects in long-range spatial dependence modeling due to being limited by the local receptive field mechanism. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, computer device and storage medium for identifying steel surface defects in view of the above technical problems.

[0005] A method for identifying steel surface defects, the method includes: Using a backbone network to extract features from the steel surface image to obtain image features.

[0006] Performing two downsamplings on the image features to obtain a first downsampled feature and a second downsampled feature.

[0007] Converting the second downsampled feature into an image feature vector and then processing it using an adaptive feature interaction encoder to obtain an encoded feature.

[0008] The image features, encoded features, and first downsampled features are processed using a feature fusion module based on a dynamic pooling pyramid and GhostConv to obtain a fused feature vector. The feature fusion module is used to process the image features, encoded features, and first downsampled features using convolution, a cross-scale feature interaction module, and a concatenation operation to obtain a fused feature vector. The cross-scale feature interaction module is used to perform multi-granularity feature integration on the features input to the module through convolution, a dynamic pooling pyramid structure, and GhostConv.

[0009] An IoU-aware query selection strategy is used to select a fixed number of image features from the fused feature vector to obtain an initial query vector.

[0010] The initial query vector is processed using a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result.

[0011] A steel surface defect recognition device, the device includes: A feature extraction module, used to extract features from the steel surface image using a backbone network to obtain image features.

[0012] An encoding module, used to perform two downsamplings on the image features to obtain a first downsampled feature and a second downsampled feature; after converting the second downsampled feature into an image feature vector, it is processed using an adaptive feature interaction encoder to obtain encoded features; the image features, encoded features, and first downsampled features are processed using a feature fusion module based on a dynamic pooling pyramid and GhostConv to obtain a fused feature vector. The feature fusion module is used to process the image features, encoded features, and first downsampled features using convolution, a cross-scale feature interaction module, and a concatenation operation to obtain a fused feature vector. The cross-scale feature interaction module is used to perform multi-granularity feature integration on the features input to the module through convolution, a dynamic pooling pyramid structure, and GhostConv.

[0013] An initial query vector determination module, used to select a fixed number of image features from the fused feature vector using an IoU-aware query selection strategy to obtain an initial query vector.

[0014] A steel surface defect recognition module, used to process the initial query vector using a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result.

[0015] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0016] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0017] The above steel surface defect recognition method, device, computer equipment and storage medium, the method uses a backbone network to extract features from the steel surface image, performs two downsamplings on the extracted image features to obtain the first downsampled feature and the second downsampled feature; converts the second downsampled feature into an image feature vector and then processes it with an adaptive feature interaction encoder to obtain an encoded feature; processes the image feature, the encoded feature and the first downsampled feature with a feature fusion module based on dynamic pooling pyramid and GhostConv to obtain a fused feature vector; uses an IoU-aware query selection strategy to filter a fixed number of image features from the fused feature vector to obtain an initial query vector; processes the initial query vector with a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result. This method realizes the full-scale defect capture ability from macroscopic deformation to microscopic cracks while maintaining computational efficiency. Brief Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of the steel surface defect recognition method in one embodiment; Figure 2 It is a flowchart of the steel surface defect recognition method in another embodiment; Figure 3 It is a structural diagram of the feature fusion module based on dynamic pooling pyramid and GhostConv in another embodiment; Figure 4 It is a schematic diagram of six types of typical industrial defect examples in another embodiment, where (a) is a pit sample diagram, (b) is an inclusion sample diagram, (c) is a plaque sample diagram, (d) is a depression sample diagram, (e) is a rolling scale sample diagram, and (f) is a scratch sample diagram; Figure 5 It is an experimental result diagram on the NEU-DET dataset in another embodiment, where (a) is a schematic diagram of the training GIoU loss, (b) is a schematic diagram of the training L1 loss, (c) is a schematic diagram of the accuracy index, (d) is a schematic diagram of the recall rate index, (e) is a schematic diagram of the validation GIoU loss, (f) is a schematic diagram of the validation L1 loss, (g) is a schematic diagram of the mAP50 index, and (h) is a schematic diagram of the mAP@0.5:0.95 index; Figure 6 It is a confusion matrix diagram in another embodiment; Figure 7It is a diagram of experimental results on the GC10-DE dataset in another embodiment, where (a) is a schematic diagram of the training GIoU loss, (b) is a schematic diagram of the training L1 loss, (c) is a schematic diagram of the accuracy metric, (d) is a schematic diagram of the recall metric, (e) is a schematic diagram of the validation GIoU loss, (f) is a schematic diagram of the validation L1 loss, (g) is a schematic diagram of the mAP50 metric, and (h) is a schematic diagram of the mAP@0.5:0.95 metric; Figure 8 It is an internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0020] In one embodiment, as Figure 1 shown, a method for identifying steel surface defects is provided, and the method includes the following steps: Step 100: Use a backbone network to extract features from a steel surface image to obtain image features.

[0021] Specifically, the backbone network is mainly used to extract features from the steel surface image.

[0022] The backbone network can select a convolutional neural network, which integrates convolutional layers, batch normalization, and activation functions, and expands the receptive field range while reducing the computational complexity.

[0023] Step 102: Perform two downsamplings on the image features to obtain a first downsampled feature and a second downsampled feature.

[0024] Specifically, the first downsampled feature is the feature obtained by performing one downsampling on the image features; the second downsampled feature is the feature obtained by performing one downsampling on the first downsampled feature.

[0025] Step 104: Convert the second downsampled feature into an image feature vector and then process it using an adaptive feature interaction encoder to obtain an encoded feature.

[0026] Specifically, the adaptive feature interaction encoder (AIFI) is used as the core feature aggregation unit, which innovatively integrates the cross-scale image feature interaction mechanism. This module realizes the efficient fusion of multi-granularity features within a single encoding layer by constructing a bidirectional cross-head attention network, and effectively captures the relationships between concept entities in the image.

[0027] Step 106: Process the image features, encoded features, and the first downsampled features using a feature fusion module based on dynamic pooling pyramid and GhostConv to obtain a fused feature vector; the feature fusion module is used to process the image features, encoded features, and the first downsampled features through convolution, cross-scale feature interaction module, and concatenation operation to obtain a fused feature vector; the cross-scale feature interaction module is used to perform multi-granularity feature integration on the features input to the module through convolution, dynamic pooling pyramid structure, and GhostConv.

[0028] Specifically, a feature fusion module based on dynamic pooling pyramid and GhostConv (abbreviated as SPPF-GhostConv module) is used to integrate the initial feature map from the shallow layer to the deep layer, while maintaining details, thereby improving the detection of subtle features. The core technology of this feature fusion module lies in establishing a joint optimization space for cross-resolution feature mapping after fusing feature maps of different sizes. Through the concatenation strategy guided by convolution and mapping operations, compared with traditional single-scale feature processing, this architecture improves the global feature extraction ability and effectively balances the requirements of macroscopic morphology perception and microscopic feature analysis in steel surface defect detection. The feature integration architecture realizes cross-resolution feature alignment through an upsampling strategy, and then constructs a hierarchical feature interaction space. After combining images of different scales, the response intensity of the defect area is effectively enhanced in the dynamic pooling pyramid, which helps more effective feature extraction and multi-frequency fusion; the combination of GhostConv allows for adjustment of important features, enabling effective extraction of the defect area in the dynamic pooling pyramid and significantly improving the efficiency and accuracy of subsequent feature extraction tasks in complex vision applications.

[0029] The cross-scale feature interaction module realizes multi-granularity feature integration through a collaborative architecture of convolution and pooling. This module innovatively combines the dynamic pooling pyramid structure and GhostConv to construct a lightweight hybrid architecture. Without significantly increasing the number of parameters, this module effectively solves the problem of lack of interactivity between different convolutional groups and improves the ability to extract feature maps, thereby improving the efficiency and effect of feature recognition.

[0030] Step 108: Select a fixed number of image features from the fused feature vector using an IoU-aware query selection strategy to obtain an initial query vector.

[0031] Specifically, this method introduces an IoU-aware query selection strategy to select a fixed number of image features from the output sequence of the front-end model as the initial query vector of the decoder.

[0032] Step 110: Process the initial query vector using a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result.

[0033] Specifically, the decoder architecture integrates an auxiliary prediction module and a prediction head, and dynamically corrects the bounding box coordinates and confidence evaluation through iterative optimization of the target query. This mechanism optimizes the feature representation through defect localization tracking, significantly improving the recognition accuracy of small-scale defects while maintaining the detection efficiency.

[0034] The process structure of the steel surface defect recognition method is as Figure 2 shown.

[0035] In the above steel surface defect recognition method, the method uses a backbone network to extract features from the steel surface image, performs two downsamplings on the extracted image features to obtain the first downsampled feature and the second downsampled feature; converts the second downsampled feature into an image feature vector and then processes it using an adaptive feature interaction encoder to obtain an encoded feature; processes the image feature, the encoded feature, and the first downsampled feature using a feature fusion module based on dynamic pooling pyramid and GhostConv to obtain a fused feature vector; uses an IoU-aware query selection strategy to filter a fixed number of image features from the fused feature vector to obtain an initial query vector; processes the initial query vector using a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result. This method realizes the full-scale defect capture ability from macroscopic deformation to microscopic cracks while maintaining the computational efficiency.

[0036] In one of the embodiments, as Figure 3 shown, the feature fusion module based on dynamic pooling pyramid and GhostConv includes: two first convolution modules, two second convolution modules, and three cross-scale feature interaction modules; step 106 includes: after passing the encoded feature through the first first convolution module, obtaining a first convolution feature; splicing the first convolution feature with the first downsampled feature to obtain a first spliced feature; processing the first spliced feature through the first cross-scale feature interaction module to obtain a cross-scale feature interaction feature; processing the cross-scale feature interaction feature using the second first convolution module to obtain a second convolution feature; splicing the second convolution feature with the image feature to obtain a second spliced feature; processing the second spliced feature through the second cross-scale feature interaction module and the first second convolution module to obtain a third convolution feature; splicing the third convolution feature with the second convolution feature to obtain a third spliced feature; processing the third spliced feature through the third cross-scale feature interaction module and the second second convolution module to obtain a fourth convolution feature; splicing the fourth convolution feature with the first convolution feature to obtain a fourth spliced feature; splicing the fourth spliced feature, the third spliced feature, and the second spliced feature to obtain a fused feature vector.

[0037] Specifically, the innovation of the feature fusion module based on the dynamic pooling pyramid and GhostConv lies in the hierarchical feature fusion mechanism and the cross-level feature reuse strategy. By constructing a cascaded pooling structure, multi-scale information capture is achieved, and target shape, size, and spatial distribution features are extracted using pooling kernels of different granularities. To address the problems of increased model complexity caused by channel redundancy and weakened features of small targets, the GhostConv module with low computational cost and high feature generation efficiency is adopted, and the discriminative representation of complex texture defects (such as oxidation spots) and geometric abnormal defects (such as microcracks) on the steel surface is enhanced through a feature channel dynamic screening mechanism.

[0038] The feature fusion module adopts a hybrid serial-parallel pooling path to reduce the parameter order while maintaining the multi-scale perception ability, and combines the feature reuse strategy to strengthen the response intensity of small-size defects. Traditional solutions usually perform bilinear interpolation for feature map upsampling and then execute channel-level cascade fusion, which results in high redundancy and the fine-grained information in the shallow high-resolution features is easily covered by the deep semantic features. Therefore, the model proposed in this application makes up for the traditional defects and greatly reduces the number of parameters caused by previous convolutions. This module proposes a three-stage feature enhancement process: First, perform upsampling on the feature map to enhance cross-scale feature interaction and detail information fusion. Then, use the Spatial Pyramid Pooling Layer (SPPF) to achieve multi-granularity feature aggregation, and reduce the parameter order through the feature splicing operation of parallel pooling branches. Finally, use the GhostConv module to perform implicit feature expansion on the fused features, and improve the feature expression ability using the redundant feature generation mechanism.

[0039] In one embodiment, the cross-scale feature interaction module includes: two convolutional modules, a dynamic pooling pyramid module, and a GhostConv module; the first concatenated feature is processed by the first cross-scale feature interaction module to obtain cross-scale feature interaction features, including: the first convolutional feature is processed by the first convolutional module to obtain a convolutional feature; the convolutional feature is processed by the dynamic pooling pyramid module to obtain three pooled features of different scales; the convolutional feature and the three pooled features of different scales are concatenated and then processed by the second convolutional module to obtain a pooled convolutional feature; the pooled convolutional feature is processed by the GhostConv module to obtain cross-scale feature interaction features.

[0040] In one embodiment, the dynamic pooling pyramid module includes three max-pooling layers; the convolutional features are processed by the dynamic pooling pyramid module to obtain pooling features of three different scales, including: after the convolutional features are processed by the first max-pooling layer, the first-scale pooling features are obtained; after the first-scale pooling features are processed by the second max-pooling layer, the second-scale pooling features are obtained; after the second-scale pooling features are processed by the third max-pooling layer, the third-scale pooling features are obtained.

[0041] In one embodiment, in the GhostConv module: a mapping operation is performed on each channel feature of the features input to the GhostConv module to obtain corresponding Ghost feature maps; the first mapping operation is an identity mapping, and the remaining mappings are lightweight transformations based on depthwise separable convolutions; the features input to the GhostConv module and the Ghost feature maps are concatenated in the channel dimension to obtain the output features of the GhostConv module.

[0042] Specifically, the specific workflow of the GhostConv module: First, the intrinsic feature map F needs to be generated, which is completed through a conventional convolution operation and serves as the basic feature representation. For each channel feature of the intrinsic feature map F, i mapping operations are performed: one of them is an identity mapping, and the remaining i - 1 times use lightweight transformations such as depthwise separable convolutions to generate Ghost feature maps as shown in Equation (2). Finally, the original intrinsic feature map and the Ghost feature maps are concatenated in the channel dimension to form the output result as shown in Equation (3).

[0043] (1) (2) (3) where F represents the input features of the GhostConv module, represents the mapping of the original feature map, represents the convolution operation for the nth channel alone, represents the nth lightweight transformation based on depthwise separable convolution, represents the output features of the GhostConv module.

[0044] In one embodiment, the first convolutional module includes: a convolutional layer with a convolutional kernel size of , a batch normalization layer, and a SiLU activation function; the second convolutional module includes: a convolutional layer with a convolutional kernel size of , a batch normalization layer, and a SiLU activation function.

[0045] In one embodiment, the backbone network is the backbone network of the ResNet structure.

[0046] Specifically, the core structure of the backbone network of the ResNet structure adopts the residual connection design. The residual structure not only alleviates the problem of gradient disappearance in deep networks but also improves the detection accuracy of small targets through cross-layer feature fusion. As the preferred backbone network, ResNet18 is adopted. The basic residual module of the ResNet architecture series forms a residual connection through cross-layer identity mapping while extracting features by convolution. The residual connection mechanism of ResNet18 allows the network to directly transmit the underlying feature information. This double-convolution layer configuration reduces the number of parameters through the parameter sharing strategy, while the cross-layer connection improves the training convergence speed.

[0047] In one embodiment, the steel surface defect recognition model is composed of a backbone network, an adaptive feature interaction encoder, a feature fusion module based on dynamic pooling pyramid and GhostConv, an IoU-aware query selection strategy, and a decoder with an auxiliary prediction head. The loss function used in the training process of the steel surface defect recognition model is the L1 loss and the GIoU loss.

[0048] The L1 loss (Mean Absolute Error, MAE), as the core evaluation index in the regression task, quantifies the model bias by calculating the mean absolute error between the predicted value and the true value y i . Compared with the L2 loss (MSE) using squared error, the linear penalty mechanism of the L1 loss effectively avoids the amplification effect of errors and shows stronger stability when dealing with noisy data. This robustness stems from the constant characteristic of its loss function gradient, making the model training process less susceptible to extreme values and thus improving the generalization ability in scenarios with complex data distributions.

[0049] (4) where represents the L1 loss, , represent the true value and the corresponding predicted value respectively, and n represents the number of training samples.

[0050] The L1 loss function constructs an error metric system based on the absolute difference between the predicted value and the true value. Its linear penalty mechanism makes the contribution of outliers to the overall loss significantly lower than that of the L2 loss with squared error. Therefore, it has a natural anti-interference ability against outliers. During the optimization process, this function maintains a fixed gradient characteristic of ±1, resulting in a sparse feature in parameter updates, which is suitable for modeling scenarios that require feature screening. Analyzing from the dimension of error interpretation, the mean absolute error output by the L1 loss is consistent with the unit of the original data, enhancing the readability of the model evaluation results. In the image reconstruction task, the balanced optimization of its pixel-level error avoids overfitting extreme noise points and effectively balances the dialectical relationship between denoising effect and detail retention.

[0051] The Generalized Intersection over Union Loss (GIoU loss), as an improved loss function in the field of object detection, demonstrates significant advantages over the traditional IoU loss in the bounding box regression task. Its core mechanism effectively solves the defect of the original IoU loss where the gradient vanishes when there is no overlap between the boxes by introducing the calculation parameter of the smallest enclosing rectangle C, that is, the area of the smallest closed region that covers both the predicted box and the ground truth box.

[0052] The GIoU loss extends its value range to the interval [-1, 1]. When the predicted box and the ground truth box completely overlap, the metric value reaches the upper limit of 1; if the two boxes are completely separated and the distance increases, the metric value approaches -1. Compared with the defect of the traditional IoU loss where the gradient becomes zero when there is no overlap between the boxes, the GIoU loss can still generate an effective gradient signal in the non-overlapping state through the design of a compensation term for the difference in the area of the closure region. This characteristic enables the model to continuously obtain the direction of parameter correction during the bounding box regression process, significantly accelerating the convergence speed. Mathematical derivation shows that the GIoU loss effectively improves the accuracy of bounding box coordinate regression through a stable gradient propagation mechanism, especially showing stronger optimization ability when dealing with small object localization and complex spatial relationships.

[0053] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0054] In one embodiment, an adaptive feature interaction encoder is replaced by a feature encoding module based on a dynamic sparse attention mechanism, which consists of a dynamic sparse attention mechanism, dynamic positional encoding, and a multi-head attention mechanism; the feature encoding module adopts a selective feature processing mechanism and only processes the top-K high-order features screened by the dynamic sparse attention. This not only significantly reduces the computational load and improves the processing speed, but also maintains the performance. In the Transformer encoding architecture, due to the lack of the inherent sequence modeling ability of convolutional or recurrent structures, explicit positional representations become a key design element. This embodiment adopts a dynamic positional encoding strategy to inject spatial coordinate information into the feature vectors, effectively solving the problem that the encoding mechanism is insensitive to spatial topology. To improve the performance of feature extraction, global semantic information is further explored on the basis of the local receptive field. Subsequently, after screening the feature maps with high feature correlation coefficients, cross-scale semantic fusion is performed through the multi-head self-attention mechanism of the Transformer. The dynamic sparse attention mechanism ignores the local features with low correlation coefficients. Therefore, the self-attention mechanism of the Transformer can well enrich the detailed features of more local and global features on the existing feature descriptions, and supplement the key parts that have subtle defects in the steel cracks but are ignored due to the coarse-grained screening of the dynamic sparse attention.

[0055] In the feature encoding module based on the dynamic sparse attention mechanism: the image features are downsampled twice to obtain downsampled features; the downsampled features are captured by the dynamic sparse attention mechanism to obtain long-distance feature correlations across image regions, resulting in dynamic attention features; the dynamic attention features are encoded and then positionally encoded using the dynamic positional encoding strategy to obtain a position encoding result; the position encoding result is subjected to cross-scale semantic fusion using the self-attention mechanism of the Transformer to obtain cross-scale semantic fusion features; the cross-scale semantic fusion features are dimensionally reshaped to obtain encoded features.

[0056] The dynamic sparse attention mechanism is used to divide the downsampled features into discrete semantic units, and then through high-dimensional space projection, obtain the Q matrix, K matrix, and V matrix; according to the Q matrix and K matrix, determine the adjacency matrix of the regional correlation between the Q matrix and the K matrix; dynamically screen the top k relevant regions of each region according to the adjacency matrix and the correlation threshold to obtain the Top-K key regions; aggregate the Top-K key regions with the discrete semantic units of the K matrix and the V matrix respectively to obtain the aggregated key and value vectors; the aggregated key and value vector pairs are processed by the self-attention mechanism and then integrated through a linear transformation layer to obtain the attention output; the attention output is dimensionally reshaped to obtain the dynamic attention features.

[0057] The steel surface defect recognition model after module replacement in this embodiment is verified using a phased verification process: First, parameter optimization is completed based on the training set, and then verification is performed on an independent test set. In scenarios with limited sample sizes, this method implements a retention verification strategy, allocating the original samples in a 4:1 ratio, where 80% is used for model training and 20% is used as the verification set. This embodiment will conduct experimental verification on the effectiveness of the model based on the NEU-DET and GC10-DET datasets. AdamW optimizer is used for parameter optimization during model training, with the base learning rate configured as 1e-4 and the momentum coefficient set to 0.9. To comprehensively evaluate the model's performance, a multi-dimensional evaluation system is constructed, covering core indicators such as classification accuracy, recall rate, and mean average precision (mAP@0.5 and mAP@0.5:0.95).

[0058] (1) Datasets As a benchmark dataset dedicated to steel surface anomaly detection, NEU-DET mainly serves the fields of computer vision and deep learning algorithm research. This dataset contains 1,800 standardized industrial image samples, with each frame image size uniformly standardized to 200 pixels × 200 pixels, covering six typical industrial defects: cracks, patches, inclusions, pitted surfaces, crazing, and scratches, providing a standardized evaluation benchmark for defect classification and localization algorithms. Examples of the six typical industrial defects are shown as Figure 4 shown, where Figure 4 (a) in is an example diagram of crazing, Figure 4 (b) in is an example diagram of inclusions, Figure 4 (c) in is an example diagram of patches, Figure 4 (d) in is an example diagram of pitted surfaces, Figure 4 (e) in is an example diagram of rolling mill scale, Figure 4 (f) in is an example diagram of scratches; GC10-DET, as an open-source industrial steel material quality inspection dataset, integrates ten types of typical surface defect samples: punching (Pu), weld (Wl), crescent gap (Cg), water stain (Ws), oil stain (Os), wire stain (Ss), inclusion (In), rolling pit (Rp), crease (Cr), and waist fold (Wf). This dataset is professionally labeled, and the samples cover diverse morphological features ranging from micron-level punctate defects to centimeter-level planar damages. Each defect category shows significant differences in geometric morphology, size distribution, and texture complexity. It is renowned for its high image quality and fine annotation. This dataset provides precise bounding box defect annotation data, facilitating the training and performance verification of detection models. Due to covering diverse defect morphology categories, it has important application value for developing industrial defect recognition algorithms, especially in the verification link of deep learning-driven detection models, and can effectively evaluate the generalization ability of algorithms to complex defect patterns. The distribution of the number of samples in each category is presented through a visualization chart. The distribution of the number of samples in each category is shown in Table 1.

[0059] Table 1 Distribution of the number of samples in each category

[0060] (2) Experimental results 1) Model performance on the NEU-DET dataset Under the influence of dynamic lighting conditions and differences in material surface properties in the NEU-DET dataset, the defect samples exhibit gray-scale features. This phenomenon leads to significant morphological differences among samples of the same type of defect, while cross-category defects show texture similarity. This dual characteristic poses dual challenges to the defect recognition model: it is necessary to overcome the problem of discrete distribution of intra-class samples and strengthen the ability to capture subtle differences between classes. Model training under such complex data distribution conditions can effectively improve the anti-interference ability of the industrial quality inspection system to lighting and material adaptability, providing a technical verification basis for the engineering deployment of the online steel surface detection system. In this embodiment, the model is trained for 250 iterations. During the 250 training cycles, the model performance shows typical convergence characteristics. In the initial stage of training (the first 50 cycles), there is a rapid optimization stage: the GIoU loss and L1 loss are respectively reduced to 0.3757 and 0.3469. After entering the mid-term training, the loss function enters a stable period, and the accuracy index is continuously optimized through parameter fine-tuning, and finally reaches the convergence state. The experimental results on the NEU-DET dataset are as Figure 5 shown, where Figure 5 (a) in it is a schematic diagram of the training GIoU loss, Figure 5 (b) in it is a schematic diagram of the training L1 loss, Figure 5 (c) in it is a schematic diagram of the accuracy index, Figure 5 (d) in it is a schematic diagram of the recall rate index, Figure 5In (e) is the schematic diagram for verifying the GIoU loss, Figure 5 In (f) is the schematic diagram for verifying the L1 loss, Figure 5 In (g) is the schematic diagram for the mAP50 metric, Figure 5 In (h) is the schematic diagram for the mAP0.5:0.95 metric. The model maintains a detection accuracy of 92.55% and a recall rate of 0.7952. This training trajectory verifies the effectiveness of the gradient optimization strategy, especially demonstrating stable performance in balancing the precision and recall rate metrics in object detection tasks.

[0061] The evaluation results of the NEU-DET dataset are shown in Table 2. The model shows different performance in the detection of six types of defects. In the classification task, the recognition accuracy of the Inclusion category is the lowest (60.2%), while the peak recall rates of the Patches, Inclusion, and Pitted_surface categories exceed 0.9. The recall rate of the Crazing category drops sharply to 0.377 due to the fusion of light-colored features with the background. There is a significant correlation between the dataset size and the model performance. The mAP@0.5 of the Patches, Pitted_surface, and Scratches categories all exceed 90%, driving the overall mAP@0.5 to 83.14%, and the mAP@0.5:0.95 is stable at 47.37%. High-contrast crack samples achieve high detection accuracy due to the advantage of distinguishable features. As Figure 6 shown in the confusion matrix diagram, the confusion matrix analysis reveals that the model has strong discriminative power in fine-grained classification tasks, but is still sensitive to subtle perturbations of background textures. This not only reflects the advantage of the algorithm in capturing defect features but also exposes the optimization space in complex industrial scenarios.

[0062] Table 2 Evaluation Results of the NEU-DET Dataset

[0063] 2) Model Performance on the GC10-DE Dataset To further evaluate the performance of the model of this method, model verification is carried out on the GC10-DET steel defect dataset, which contains 3,570 high-resolution industrial images (2048 pixels × 1000 pixels) and is divided into a training set and a validation set in a 4:1 ratio. After 300 training epochs of optimization, the model achieves a detection accuracy of 79.61% and a recall rate of 0.6459. The GIOU loss and the L1 loss converge to 0.5542 and 0.3359 respectively, verifying the effectiveness of the multi-task optimization mechanism. The experimental results on the GC10-DE dataset are as Figure 7 shown, where Figure 7Among them, (a) is a schematic diagram of the training GIoU loss, Figure 7 among which, (b) is a schematic diagram of the training L1 loss, Figure 7 among which, (c) is a schematic diagram of the accuracy metric, Figure 7 among which, (d) is a schematic diagram of the recall metric, Figure 7 among which, (e) is a schematic diagram of the validation GIoU loss, Figure 7 among which, (f) is a schematic diagram of the validation L1 loss, Figure 7 among which, (g) is a schematic diagram of the mAP50 metric, Figure 7 among which, (h) is a schematic diagram of the mAP@0.5:0.95 metric. The experimental results show that the detection framework has stable feature extraction ability in complex industrial scenarios, and its multi-scale feature fusion mechanism effectively balances the localization accuracy and classification performance.

[0064] Verification on the GC10-DET industrial dataset shows that the model can exhibit differentiated performance characteristics in real-scenario applications. In the detection of specific defect types, the weld (Wl) category leads with an accuracy of 90.6%, while the mAP@0.5 of the inclusion (In) and rolling pit (Rp) categories is lower than 30%, being 29.6% and 27.7% respectively. It is worth noting that although the overall mAP@0.5 of this dataset reaches 67.27%, the detection accuracies of the punching (Pu), weld (Wl), and crescent gap (Cg) defect types exceed 90%, confirming the model's advantage in feature extraction for high-contrast defects. This performance difference reveals that in industrial inspection scenarios, the visual separability between the target and the background has a decisive impact on the model's performance.

[0065] The sample size of the GC10-DET dataset shows a significant correlation with the model's detection performance. For the rolled pit (Rp) and crease (Cr) classes, the detection accuracy is limited due to the scarcity of samples. However, for the punched hole (Pu) and weld seam (Wl) classes, the mAP@0.5:0.95 metrics reach 52.8% and 52.9% respectively, significantly better than the overall average of 34.09%. The distribution characteristics of mAP@0.5:0.95 indicate that the visual separability between the target and the background becomes a key limiting factor. In industrial scenarios, complex texture backgrounds are prone to causing feature confusion, especially having a significant impact on low-contrast defects. This performance difference reveals that the balance of inter-class sample distribution and background complexity in the dataset jointly affect the model's generalization ability. Therefore, the number of datasets has a great impact on the parameter optimization of the model. By increasing the data size of low-sample categories such as cracks and optimizing the distribution structure, the model's feature separation ability in complex backgrounds is enhanced. In the detection of ten types of defects, the framework of this method achieves an accuracy of 79.61% in the crack recognition task. From the accuracies of the inclusion (In) and rolled pit (Rp) categories, it can be observed that due to the small number of datasets, the accuracy cannot be correspondingly improved, confirming the positive effect of data distribution optimization on the model's robustness. The accuracies of the 10 categories are shown in Table 3; the total accuracies of the two datasets are shown in Table 4.

[0066] Table 3 Accuracies of 10 categories

[0067] Table 4 Total accuracies of two datasets

[0068] (3) Comparative experiments The steel surface defect recognition module with the feature encoding module based on the dynamic sparse attention mechanism replacing the adaptive feature interaction encoder is compared with the current advanced models on the NEU-DET dataset. The experiments show that this model exhibits significant advantages in terms of accuracy metrics. The experimental results of the current advanced models on the NEU-DET dataset are shown in Table 5.

[0069] Table 5 Experimental results of current advanced models on the NEU-DET dataset

[0070] After comparing the NEU-DET dataset, supplementary tests are conducted using the GC10-DET dataset. The experimental results of the YOLO series models and this method on the GC10-DET dataset are shown in Table 6. The experimental results of the YOLO5 model and this method on the GC10-DET dataset are shown in Table 7. This method has the advantage of stability, further verifying its comprehensive reliability.

[0071] Table 6 Experimental results of YOLO series models and this method on the GC10-DET dataset

[0072] Table 7 Experimental results of YOLO5 model and this method on the GC10-DET dataset

[0073] (3) Ablation and visualization experiments In this embodiment, ablation experiments are carried out on NEU-DET to verify the performance contributions of each module. The results of the ablation experiments on NEU-DET are shown in Table 8.

[0074] The baseline model A integrates the ResNet18 backbone network and the DETR architecture, achieving 79.40% mAP@0.5 and 43.55% mAP@0.5:0.95. On the basis of model A, model B introduces a feature encoding module based on the dynamic sparse attention mechanism, integrates the dynamic sparse attention mechanism, and further optimizes the detection effect. While increasing the parameter scale by 0.21M, this scheme improves mAP@0.5 and mAP@0.5:0.95 by 3.27% and 1.45% respectively compared with the baseline model by strengthening the image information of associated tokens. This structural improvement significantly optimizes the performance without excessive increase in the overall computational load while maintaining the detection accuracy.

[0075] Model C introduces a feature fusion module (SPPF-GhostConv) based on dynamic pooling pyramid and GhostConv into the architecture of model B, integrating the dynamic pooling pyramid mechanism and GhostConv. Through the synergistic effect of the dynamic pooling pyramid mechanism and GhostConv, this design effectively suppresses the parameter growth while strengthening the feature interaction, realizing the synchronous optimization of detection accuracy and model efficiency. This scheme improves mAP@0.5 and mAP@0.5:0.95 to 83.14% and 47.37% respectively. The ablation experiment data shows that removing any core component causes a step-by-step decay in accuracy, confirming the decisive influence of the synergistic effect of the dynamic pooling pyramid mechanism and GhostConv on the model performance.

[0076] Model D constructed based on the architecture of model C aims to verify the impact of the stacking of encoding layers on performance. The experiment finds that excessive stacking of Transformer encoding layers will cause performance degradation, specifically manifested as a significant decline of 0.77% in the mAP@0.5 index and a synchronous decrease of 2.57 percentage points in mAP@0.5:0.95. This reverse optimization phenomenon not only confirms the rationality of the original architecture but also provides crucial experimental basis for the subsequent optimization path.

[0077] In the verification of the CNN-based module E, ResNet50 is adopted as the infrastructure of the convolutional embedding framework. Experiments show that although the parameter quantity and operation overhead of this scheme increase significantly on the NEU-DET dataset, the improvement in detection accuracy is limited. Quantitative analysis shows that the detection accuracy index mAP@0.5:0.95 only increases by 1.23% compared with the baseline model A, and the synchronous optimization of performance is not achieved. It can also be judged that excessive parameter quantity and computational complexity will weaken the engineering deployment value of the model. Finally, ResNet18 is selected as the feature extraction backbone in this method, and this architecture achieves a better balance configuration between computational efficiency and detection accuracy.

[0078] Table 8 Module ablation experiment

[0079] After completing the module-level accuracy evaluation, the following conducts a visual comparison study of the detection results for the NEU-DET dataset. In this embodiment, the model after module replacement not only realizes the accurate classification and localization of crack-like defects, but also shows superior performance in typical samples. Taking the detection of crazing as an example, although the baseline model can complete defect classification, this scheme has obvious advantages in terms of localization accuracy and feature distinguishability. Introducing the dynamic sparse attention mechanism significantly expands the detection coverage area. In the integrated model of GhostConv, this combined architecture shows stronger environmental adaptability. Especially in the inclusion-like scenarios, the crack recognition accuracy is systematically improved. Experimental data shows that the newly added module effectively improves the suppression efficiency of complex background interference, enabling the algorithm to successfully capture the micro-scale crack features that were previously missed. For the patches-like detection scenario, the baseline model has the dual defects of detection redundancy and insufficient area coverage. In the pitted_surface detection task, the optimized module shows a gain effect: reducing the number of redundant candidate boxes while maintaining the detection accuracy. Visualization confirms that this architecture has the ability to accurately represent the geometric features of defects, especially reducing the recognition error of the crack distribution range and morphological features, highlighting the breakthrough progress of the module in the aspect of feature space optimization.

[0080] To more significantly verify the recognition ability of the model proposed in this application, the recognition ability on the GC10-DET dataset is also visually presented. Through the visualization of water stains (Ws), creases (Cr), and the mixture of different defects, it can be seen that as the modules are stacked, the recognition of different types of defects becomes more accurate. It can be seen from the following aspects: as the modules are stacked, the redundant bounding boxes for the recognition of several defects decrease, and the recognition scores increase significantly, finally achieving an excellent final effect. This progressive optimization mechanism not only enhances the localization accuracy of small defects but also effectively suppresses the background noise interference in industrial scenarios through cross-layer feature fusion, ultimately verifying the robust generalization ability of the model in multi-source data scenarios.

[0081] The comprehensive experimental results confirm that the model after module replacement in this embodiment shows strong advantages in complex industrial scenarios: it not only enhances the robustness to background noise but also improves the recognition sensitivity to micro-defects. Visual analysis verifies the effectiveness of the algorithm in multi-type defect detection, where the localization error of weak-texture defects is reduced.

[0082] In one embodiment, a steel surface defect recognition device is provided, including: a feature extraction module, an encoding module, an initial query vector determination module, and a steel surface defect recognition module, where: The feature extraction module is used to extract features from the steel surface image using a backbone network to obtain image features.

[0083] The encoding module is used to perform two downsamplings on the image features to obtain the first downsampled feature and the second downsampled feature; after converting the second downsampled feature into an image feature vector, it is processed using an adaptive feature interaction encoder to obtain encoded features; the image features, encoded features, and the first downsampled feature are processed using a feature fusion module based on dynamic pooling pyramid and GhostConv to obtain a fused feature vector; the feature fusion module is used to process the image features, encoded features, and the first downsampled feature using convolution, cross-scale feature interaction module, and concatenation operation to obtain a fused feature vector; the cross-scale feature interaction module is used to perform multi-granularity feature integration on the features input to the module through convolution, dynamic pooling pyramid structure, and GhostConv.

[0084] The initial query vector determination module is used to screen a fixed number of image features from the fused feature vector using an IoU-aware query selection strategy to obtain an initial query vector.

[0085] The steel surface defect recognition module is used to process the initial query vector using a decoder with an auxiliary prediction head to obtain the steel surface defect recognition result.

[0086] In one embodiment, the feature fusion module based on the dynamic pooling pyramid and GhostConv includes: two first convolution modules, two second convolution modules, and three cross-scale feature interaction modules; the encoding module is further configured to obtain first convolution features after passing the encoded features through the first first convolution module; splice the first convolution features with the first downsampled features to obtain first spliced features; process the first spliced features through the first cross-scale feature interaction module to obtain cross-scale feature interaction features; process the cross-scale feature interaction features with the second first convolution module to obtain second convolution features; splice the second convolution features with the image features to obtain second spliced features; process the second spliced features through the second cross-scale feature interaction module and the first second convolution module to obtain third convolution features; splice the third convolution features with the second convolution features to obtain third spliced features; process the third spliced features through the third cross-scale feature interaction module and the second second convolution module to obtain fourth convolution features; splice the fourth convolution features with the first convolution features to obtain fourth spliced features; splice the fourth spliced features, the third spliced features, and the second spliced features to obtain a fused feature vector.

[0087] In one embodiment, the cross-scale feature interaction module includes: two convolution modules, a dynamic pooling pyramid module, and a GhostConv module; the encoding module is further configured to process the first convolution features with the first convolution module to obtain convolution features; process the convolution features with the dynamic pooling pyramid module to obtain three pooled features at different scales; splice the convolution features and the three pooled features at different scales and then process them with the second convolution module to obtain pooled convolution features; process the pooled convolution features with the GhostConv module to obtain cross-scale feature interaction features.

[0088] In one embodiment, the dynamic pooling pyramid module includes three maximum pooling layers; the encoding module is further configured to process the convolution features through the first maximum pooling layer to obtain first-scale pooled features; process the first-scale pooled features through the second maximum pooling layer to obtain second-scale pooled features; process the second-scale pooled features through the third maximum pooling layer to obtain third-scale pooled features.

[0089] In one embodiment, the encoding module is further configured to, in the GhostConv module: perform a mapping operation on each channel feature of the feature input to the GhostConv module to obtain a corresponding Ghost feature map; where the first mapping operation is an identity mapping, and the remaining mappings are lightweight transformations based on depthwise separable convolutions; and splice the feature input to the GhostConv module and the Ghost feature map in the channel dimension to obtain the output feature of the GhostConv module.

[0090] In one embodiment, the first convolutional module includes: a convolutional layer with a convolutional kernel size of a batch normalization layer, and a SiLU activation function; the second convolutional module includes: a convolutional layer with a convolutional kernel size of a batch normalization layer, and a SiLU activation function.

[0091] In one embodiment, the backbone network in the feature extraction module is the backbone network of the ResNet structure.

[0092] For the specific limitations of the steel surface defect recognition device, reference can be made to the limitations on the steel surface defect recognition method in the foregoing text, which will not be elaborated here. Each module in the above steel surface defect recognition device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0093] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a steel surface defect recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0094] Those skilled in the art can understand that Figure 8The structure shown is only a block diagram of some of the structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component arrangement.

[0095] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiment are implemented.

[0096] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0097] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0098] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0099] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for identifying surface defects of steel, characterized in that: The method comprises: The backbone network is used to extract features from the steel surface image to obtain image features; Downsampling the image feature twice to obtain a first downsampling feature and a second downsampling feature; The second down-sampled features are converted into image feature vectors and then processed using an adaptive feature interactive encoder to obtain encoded features; The image feature, the coding feature and the first down-sampling feature are processed by a feature fusion module based on dynamic pooling pyramid and GhostConv to obtain a fused feature vector; the feature fusion module is used to process the image feature, the coding feature and the first down-sampling feature by using convolution, a cross-scale feature interaction module and a splicing operation to obtain a fused feature vector; the cross-scale feature interaction module is used to perform multi-granularity feature integration on the features of the input module through convolution, a dynamic pooling pyramid structure and GhostConv; Using an IoU-aware query selection strategy to filter a fixed number of image features from the fused feature vector to obtain an initial query vector; The initial query vector is processed by a decoder with an auxiliary prediction head to obtain a steel surface defect recognition result.

2. The method for identifying steel surface defects according to claim 1, characterized in that: The feature fusion module based on dynamic pooling pyramid and GhostConv includes: two first convolution modules, two second convolution modules and three cross-scale feature interaction modules; The image feature, the encoding feature and the first down-sampling feature are processed using a feature fusion module based on a dynamic pooling pyramid and GhostConv to obtain a fused feature vector, including: Passing the encoded feature through a first convolution module to obtain a first convolution feature; Concatenate the first convolution feature with the first down-sampled feature to obtain a first concatenated feature; Processing the first concatenated feature through the first cross-scale feature interaction module to obtain a cross-scale feature interaction feature; Processing the cross-scale feature interaction feature by using the second first convolution module to obtain a second convolution feature; Splicing the second convolution feature with the image feature to obtain a second splicing feature; Processing the second concatenated feature through the second cross-scale feature interaction module and the first of the second convolution modules to obtain a third convolution feature; Concatenating the third convolution feature with the second convolution feature to obtain a third concatenated feature; Processing the third concatenated feature through the third cross-scale feature interaction module and the second second convolution module to obtain a fourth convolution feature; Concatenating the fourth convolution feature with the first convolution feature to obtain a fourth concatenated feature; After splicing the fourth splicing feature, the third splicing feature and the second splicing feature, a fused feature vector is obtained.

3. The method for identifying steel surface defects according to claim 2, characterized in that: The cross-scale feature interaction module includes: two convolution modules, a dynamic pooling pyramid module and a GhostConv module; Processing the first splicing feature through the first cross-scale feature interaction module to obtain a cross-scale feature interaction feature includes: Processing the first convolution feature using the first convolution module to obtain a convolution feature; The convolutional features are processed using a dynamic pooling pyramid module to obtain pooling features of three different scales; The convolution feature and the pooling features of three different scales are concatenated and processed by the second convolution module to obtain a pooled convolution feature; The pooled convolutional features are processed using a GhostConv module to obtain cross-scale feature interaction features.

4. The method for identifying steel surface defects according to claim 3, characterized in that: The dynamic pooling pyramid module includes three maximum pooling layers; The convolutional features are processed using a dynamic pooling pyramid module to obtain pooling features of three different scales, including: After the convolutional features are processed by a first maximum pooling layer, a first-scale pooling feature is obtained; After the first-scale pooling feature is processed by the second maximum pooling layer, the second-scale pooling feature is obtained; The second-scale pooling features are processed by the third maximum pooling layer to obtain third-scale pooling features.

5. The method for identifying surface defects of steel according to claim 3, characterized in that: In the GhostConv module: Mapping operations are performed on each channel feature of the features input to the GhostConv module to obtain the corresponding Ghost feature map; the first mapping operation is an identity mapping, and the remaining mappings are lightweight transformations based on depthwise separable convolutions; The features of the input GhostConv module are concatenated with the Ghost feature map in the channel dimension to obtain the output features of the GhostConv module.

6. The method for identifying steel surface defects according to claim 2, characterized in that: The first convolution module includes: the convolution kernel size is Convolutional layers, batch normalization layers, and SiLU activation functions; The second convolution module includes: the convolution kernel size is Convolutional layers, batch normalization layers, and SiLU activation functions.

7. The method for identifying steel surface defects according to claim 1, characterized in that: The backbone network is a backbone network of the ResNet structure.

8. A steel surface defect identification device, characterized in that: The device comprises: A feature extraction module is used to extract features from the steel surface image using a backbone network to obtain image features; The encoding module is used to downsample the image feature twice to obtain a first downsampled feature and a second downsampled feature; the second downsampled feature is converted into an image feature vector and then processed by an adaptive feature interaction encoder to obtain a coding feature; the image feature, the coding feature and the first downsampled feature are processed by a feature fusion module based on a dynamic pooling pyramid and GhostConv to obtain a fused feature vector; the feature fusion module is used to process the image feature, the coding feature and the first downsampled feature by using a convolution, a cross-scale feature interaction module and a splicing operation to obtain a fused feature vector; the cross-scale feature interaction module is used to perform multi-granularity feature integration on the features of the input module through convolution, a dynamic pooling pyramid structure and GhostConv; An initial query vector determination module, configured to select a fixed number of image features from the fused feature vector using an IoU-aware query selection strategy to obtain an initial query vector; The steel surface defect recognition module is used to process the initial query vector using a decoder with an auxiliary prediction head to obtain a steel surface defect recognition result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the steel surface defect identification method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the steel surface defect identification method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Construction site intrusion detection method, computer equipment and storage medium

    CN118644760A

  • Strip steel surface defect area size measuring method

    CN119313669A

  • Spacecraft thermal control film coating defect identification method based on multi-dimensional attention mechanism

    CN119515792A

  • Improved yolov5-based rockfall detection method in complex environments

    KR102783800B1

Cited By

  • Steel surface defect detection method and device, computer equipment and storage medium

    CN120726028A

  • Methods, apparatus, computer equipment and storage media for detecting defects on steel surfaces

    CN120726028B

  • Drilling crack type prediction method and system based on deep learning and medium

    CN120997673A

  • Visual inspection method and system for intelligent blank carrying line

    CN121392214A