Steel surface microscopic crack feature extraction method, device and equipment and storage medium

The cross-scale semantic fusion of dynamic sparse attention mechanism and Transformer's self-attention mechanism is solved, and the problem of fine-grained feature loss in steel surface defect detection is improved. The efficiency and accuracy of microcrack feature extraction of steel surface microscopic crack feature is improved.

CN120510135APending Publication Date: 2025-08-19JIANGXI UNIV OF SCI & TECH +3
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510644205.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional convolutional neural networks are difficult to effectively capture multi-scale and multi-morphological defects in steel surface defect detection, especially in complex scenarios where fine cracks and macroscopic oxidized patches coexist. Fixed coding mode can easily lead to the loss of fine-grained features.

Method used

The dynamic sparse attention mechanism is used to capture the long-distance feature association across the image area, and the cross-scale semantic fusion is combined with dynamic position coding and Transformer's self-attention mechanism. Position coding is performed through the dynamic sparse attention mechanism and dynamic position coding strategy, and the cross-scale semantic fusion is used to obtain subtle defect characteristics on the steel surface.

Benefits of technology

The efficiency and accuracy of microcrack feature extraction on the steel surface are improved, and the robustness of deep learning models is enhanced when dealing with long-range dependence and diversified features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510135A_ABST
    Figure CN120510135A_ABST
Patent Text Reader

Abstract

The invention relates to a steel surface microscopic crack feature extraction method and device, equipment and a storage medium. The method comprises the following steps: acquiring steel surface features, performing down-sampling on the steel surface features twice, and then capturing long-distance feature association of a cross-image region by adopting a dynamic sparse attention mechanism to obtain dynamic attention features; after the dynamic attention features are coded, a dynamic position coding strategy is adopted for position coding, and a position coding result is obtained; carrying out cross-scale semantic fusion on a position coding result by adopting a self-attention mechanism of Transform, and then carrying out dimension remodeling to obtain fine defect features on the surface of the steel; and performing defect detection on the steel surface according to the fine defect characteristics of the steel surface. According to the method, a dynamic sparse attention mechanism is adopted, so that the robustness of the deep learning model in processing long-range dependence and diversified features is effectively enhanced, and the efficiency and accuracy of extracting the steel surface microscopic crack features are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of deep learning and steel defect detection, and in particular to a method, device, equipment and storage medium for extracting micro-crack features on the surface of steel. Background Art

[0002] In the industrial production system, steel is an indispensable basic material, and its quality directly affects the safety and reliability of the end products. It is worth noting that surface damage is prone to occur during the manufacturing process, mainly including typical defect types such as corrosion deformation, mechanical scratches and impurity embedding. These surface abnormalities will not only significantly weaken the mechanical properties of the material, but also cause potential structural failure risks. Therefore, building an accurate defect detection system plays a key role in ensuring the quality of steel and has become an important quality control link that cannot be ignored in the modern manufacturing process. Driven by the dual upgrades in the industrial manufacturing field and the improvement of material quality standards, material quality control has become a key technical indicator. Traditional manual detection methods have gradually exposed their shortcomings of insufficient efficiency and lack of stability in large-scale production scenarios, and are difficult to adapt to the precision needs of modern industry.

[0003] Traditional convolutional neural networks (CNNs), limited by their local receptive field mechanisms, face inherent limitations in modeling long-range spatial dependencies due to the multi-scale and multi-morphological challenges commonly encountered in steel surface defect detection scenarios. By building a cross-level feature correlation map, a dynamic attention mechanism based on window interactions can effectively capture defect morphological features from microscopic to macroscopic scales. In industrial vision applications of various Transformer architecture variations, the position encoding mechanism, as a core component for modeling geometric relationships in feature space, has been crucial for defect detection. While traditional absolute position encoding can provide basic coordinate information, it suffers from insufficient spatial resolution adaptability when dealing with multi-scale defects on steel surfaces. In particular, in complex scenarios where fine cracks coexist with macroscopic oxidation patches, fixed encoding schemes can easily lead to the loss of fine-grained features. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, equipment and storage medium for extracting micro-crack features on the surface of steel to address the above technical problems.

[0005] A method for extracting micro crack features on a steel surface, the method comprising: The surface features of the steel are obtained, and the surface features of the steel are downsampled twice to obtain downsampled features.

[0006] The downsampled features are used to capture long-range feature correlations across image regions using a dynamic sparse attention mechanism to obtain dynamic attention features.

[0007] After encoding the dynamic attention features, the dynamic position encoding strategy is used to perform position encoding to obtain the position encoding result.

[0008] The position encoding results are subjected to cross-scale semantic fusion using the Transformer’s self-attention mechanism to obtain cross-scale semantic fusion features.

[0009] The cross-scale semantic fusion features are reshaped to obtain subtle defect features on the steel surface.

[0010] Defect detection is performed on the steel surface based on the subtle defect characteristics of the steel surface.

[0011] In one embodiment, a dynamic sparse attention mechanism is used to capture long-range feature correlations across image regions to obtain dynamic attention features, including: The downsampled features are divided into discrete semantic units, and then projected through high-dimensional space to obtain the Q matrix, K matrix, and V matrix.

[0012] Based on the Q matrix and the K matrix, an adjacency matrix of region correlations to the Q matrix and the K matrix is determined.

[0013] The top k relevant regions of each region are dynamically screened according to the adjacency matrix and the correlation threshold to obtain the Top-K key regions.

[0014] The Top-K key regions are clustered with the discrete semantic units of the K matrix and the V matrix respectively to obtain the clustered key and value vectors.

[0015] The aggregated key and value vectors are processed using the self-attention mechanism and then integrated through a linear transformation layer to obtain the attention output.

[0016] The attention output is reshaped to obtain dynamic attention features.

[0017] In one embodiment, the surface features of the steel material are extracted from a steel material surface image extracted by a ResNet structural feature extraction network.

[0018] A device for extracting microscopic crack features on a steel surface, comprising: The dynamic attention feature extraction module is used to obtain the surface features of steel and downsample the surface features of steel twice to obtain downsampled features; the downsampled features are used to capture long-range feature associations across image regions using a dynamic sparse attention mechanism to obtain dynamic attention features.

[0019] The position encoding module is used to encode the dynamic attention features and then use the dynamic position encoding strategy to perform position encoding to obtain the position encoding result.

[0020] The module for extracting subtle defect features on the steel surface is used to perform cross-scale semantic fusion on the position encoding results using the Transformer's self-attention mechanism to obtain cross-scale semantic fusion features; the cross-scale semantic fusion features are then reshaped to obtain subtle defect features on the steel surface.

[0021] The steel surface defect detection module is used to detect defects on the steel surface based on the subtle defect characteristics of the steel surface.

[0022] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0023] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.

[0024] The above-mentioned method, device, equipment, and storage medium for extracting microcrack features from steel surfaces include: obtaining steel surface features, downsampling them twice, and then using a dynamic sparse attention mechanism to capture long-range feature correlations across image regions to obtain dynamic attention features; encoding the dynamic attention features and then position-encoding them using a dynamic position encoding strategy to obtain position-encoded results; performing cross-scale semantic fusion on the position-encoded results using a Transformer self-attention mechanism and then reshaping the dimensions to obtain subtle defect features on the steel surface; and performing defect detection on the steel surface based on the subtle defect features. This method utilizes a dynamic sparse attention mechanism to effectively enhance the robustness of deep learning models in processing long-range dependencies and diverse features, thereby improving the efficiency and accuracy of extracting microcrack features from steel surfaces. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Schematic diagram of a flow chart of a method for extracting micro-crack features on a steel surface according to an embodiment; Figure 2 Schematic diagram of the process of extracting micro crack features from steel surface in one embodiment; Figure 3 Schematic diagram of the self-attention mechanism flow of Transformer in one embodiment; Figure 4 is a diagram showing the internal structure of a computer device in another embodiment; Figure 5Schematic diagrams of six typical industrial defects in another embodiment, wherein (a) is a pit example, (b) is an inclusion example, (c) is a plaque example, (d) is a depression example, (e) is a rolling scale example, and (f) is a scratch example. Figure 6 Schematic diagram of experimental results on the NEU-DET dataset in another embodiment, where (a) is a schematic diagram of training GIoU loss, (b) is a schematic diagram of training L1 loss, (c) is a schematic diagram of precision index, (d) is a schematic diagram of recall index, (e) is a schematic diagram of verification GIoU loss, (f) is a schematic diagram of verification L1 loss, (g) is a schematic diagram of mAP50 index, and (h) is a schematic diagram of mAP@0.5:0.95 index; Figure 7 A confusion matrix diagram in another embodiment; Figure 8 Schematic diagram of experimental results on the GC10-DE dataset in another embodiment, where (a) is a schematic diagram of training GIoU loss, (b) is a schematic diagram of training L1 loss, (c) is a schematic diagram of accuracy index, (d) is a schematic diagram of recall index, (e) is a schematic diagram of verification GIoU loss, (f) is a schematic diagram of verification L1 loss, (g) is a schematic diagram of mAP50 index, and (h) is a schematic diagram of mAP@0.5:0.95 index. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0027] In one embodiment, Figure 1 As shown, a method for extracting micro crack features on the surface of steel is provided, and the method comprises the following steps: Step 100: Acquire surface features of the steel material, and downsample the surface features of the steel material twice to obtain downsampled features.

[0028] Specifically, the surface features of the steel are obtained by extracting features from the steel surface image using a convolutional neural network.

[0029] Step 102: Use a dynamic sparse attention mechanism to capture long-range feature associations across image regions using the downsampled features to obtain dynamic attention features.

[0030] Specifically, sparse attention is an optimized attention mechanism that can map a query vector and a set of key-value pairs to an output vector. However, unlike single-head attention and multi-head attention, it does not calculate the similarity between the query vector and all key vectors, but only calculates the similarity between the query vector and some key vectors, thereby reducing the amount of computation and memory consumption.

[0031] The dynamic sparse attention mechanism is a dynamic, query-aware sparse attention mechanism. The key idea of the dynamic sparse attention mechanism is to filter out most of the irrelevant key-value pairs at the coarse region level so that only a small part of the routing region is retained.

[0032] Step 104: After encoding the dynamic attention features, position encoding is performed using a dynamic position encoding strategy to obtain a position encoding result.

[0033] Specifically, as a core component of feature space geometric relationship modeling, optimizing the performance of the position encoding mechanism is crucial for defect detection. While traditional absolute position encoding can provide basic coordinate information, it suffers from insufficient spatial resolution adaptability when dealing with multi-scale defects on steel surfaces. This is especially true in complex scenarios where microcracks coexist with macroscopic oxidation patches, where fixed encoding schemes can easily lead to loss of fine-grained features. Therefore, this method employs a dynamic position encoding strategy for position encoding.

[0034] Step 106: The position encoding result is subjected to cross-scale semantic fusion using the Transformer self-attention mechanism to obtain a cross-scale semantic fusion feature.

[0035] Specifically, the Transformer's self-attention mechanism can effectively enrich the existing feature description with more detailed features of local and overall features, supplementing the key parts of subtle defects in steel cracks that are ignored due to the coarse-grained screening of dynamic sparse attention.

[0036] Step 108: Reshape the cross-scale semantic fusion features to obtain subtle defect features on the steel surface.

[0037] Step 110: Defect detection is performed on the steel surface according to the characteristics of subtle defects on the steel surface.

[0038] Specifically, the dynamic sparse attention mechanism, dynamic position encoding strategy and Transformer self-attention mechanism are combined to form a steel surface microcrack feature extraction model. The structure of the steel surface microcrack feature extraction model is as follows: Figure 2 shown.

[0039] The above-mentioned method for extracting microcrack features from steel surfaces includes: obtaining steel surface features, downsampling them twice, and then using a dynamic sparse attention mechanism to capture long-range feature correlations across image regions to obtain dynamic attention features; encoding the dynamic attention features and then position-encoding them using a dynamic position encoding strategy to obtain position-encoded results; performing cross-scale semantic fusion on the position-encoded results using the Transformer self-attention mechanism and then reshaping the dimensions to obtain subtle defect features on the steel surface; and performing defect detection on the steel surface based on the subtle defect features on the steel surface. This method uses a dynamic sparse attention mechanism to effectively enhance the robustness of deep learning models in processing long-range dependencies and diverse features, thereby improving the efficiency and accuracy of extracting microcrack features from steel surfaces.

[0040] In one embodiment, step 102 includes: dividing the downsampled features into discrete semantic units, and then projecting them through a high-dimensional space to obtain a Q matrix, a K matrix, and a V matrix; determining an adjacency matrix of regional correlations to the Q matrix and the K matrix based on the Q matrix and the K matrix; dynamically screening the top k related regions of each region based on the adjacency matrix and the correlation threshold to obtain the Top-K key regions; clustering the Top-K key regions with the discrete semantic units of the K matrix and the V matrix respectively to obtain clustered key and value vectors; processing the clustered key and value vectors using a self-attention mechanism and then integrating them through a linear transformation layer to obtain an attention output; and performing a dimension reshaping operation on the attention output to obtain a dynamic attention feature.

[0041] Specifically, the core innovation of the dynamic sparse attention mechanism lies in its two-stage feature focusing strategy: first, redundant background regions in the feature map are filtered out through a coarse-grained filtering mechanism, followed by fine-grained correlation analysis within the retained regions of interest. This architecture customizes the terminal features of the CNN output from the backbone network, effectively combining the local feature extraction advantages of convolution operations with the global context modeling capabilities of the attention mechanism. Traditional convolution kernels are limited by their local receptive field, while the global dependencies established by the dynamic sparse attention mechanism can capture long-range feature correlations across image regions.

[0042] The dynamic sparse attention mechanism employs a staged feature focusing strategy: first, the input feature map is divided into discrete semantic units (tokens), and a global correlation matrix is constructed through high-dimensional spatial projection. Second, feature aggregation is performed by dynamically selecting the top-K key regions based on a correlation threshold. This architecture innovatively combines the local perception characteristics of convolution with the global modeling advantages of attention. Traditional sliding window mechanisms (such as the Swin Transformer) result in redundant parameters due to fixed window divisions, while this scheme significantly reduces the required parameters through dynamic path planning. Images extracted by dynamic sparse image attention undergo multi-scale sequence reconstruction, expanding the convolutional feature tensor (RB×C×H×W) along the spatial dimension and reconstructing it into a temporally parsable vector sequence (RB×N×C).

[0043] Input a feature map, X ∈R(H×W×C), first divide it into S×S different regions, each of which contains eigenvectors. That is, X becomes ∈R( × ×C). Then, the Q, K, and V matrices are obtained through linear mapping; the Q, K, and V matrices are: Q =

[0044] K =

[0045] V =

[0046] in, 、 、 ∈R (C×C) are the projection weights of query, key, and value respectively, and the Q, K, and V matrices ∈R ( × ×C). We construct a vector matrix to find the regions that each given region should participate in. We then calculate the adjacency matrix of the region correlations of Q and K: A=Q

[0047] The correlation graph is then pruned by retaining only the top k connections of each region. I contains the indices of the top k most correlated regions.

[0048] I=topIndex(A) However, it is not easy to implement this step efficiently because these feature regions will be scattered throughout the feature map, and the GPU relies on coalesced memory operations to load blocks of dozens of consecutive bytes at a time. Therefore, we first aggregate the tensors of K and V, i.e. K=gather(K,I) V=gather(V,I) Among them, the K matrix and the V matrix are the vectors of the clustered key and value, and then the attention operation is used on the clustered KV pairs: O=Attention(Q,K,V) Based on this, a selective feature processing mechanism is employed during the encoding process, processing only the top-K high-order features filtered by dynamic sparse attention. This not only significantly reduces the computational load and improves processing speed, but also maintains performance. In the Transformer encoding architecture, due to the lack of the inherent sequence modeling capabilities of convolutional or recurrent structures, explicit positional representation becomes a key design element. This method employs a dynamic positional encoding strategy, injecting spatial coordinate information into the feature vector, effectively addressing the encoding mechanism's insensitivity to spatial topology. To improve feature extraction performance, global semantic information is further explored based on the local receptive field. Subsequently, after filtering feature maps with high feature correlation coefficients, cross-scale semantic fusion is achieved through the Transformer's multi-head self-attention mechanism. Dynamic sparse attention ignores local features with low correlation coefficients. Therefore, the self-attention mechanism can effectively enrich the existing feature representation with more detailed local and global features, supplementing the critical components of subtle defects such as steel cracks that are overlooked by the coarse-grained filtering of dynamic sparse attention. First, in the feature serialization stage, the input feature vector is decoupled into three projection spaces: query, key, and value. A cross-position correlation matrix is generated by calculating global correlation (i.e., the dot product of the query vector and the key vector). This mechanism probabilistically normalizes the correlation matrix using the Softmax function to form attention distribution weights. Finally, the feature representation of the feature map is reconstructed through weighted fusion of the value vectors. A self-attention mechanism (SA(Q, K, V)) is used for this processing. Specifically, the transformed features are embedded in the encoder using relative positional encoding. Attention is then calculated on the feature map to mine global semantics. This mechanism uses multiple independent self-attention heads to extract diverse dependencies in different feature subspaces. The results of each attention head are integrated through a linear transformation layer to form the final attention output. The output feature tensor undergoes a dimension reshaping operation to adapt it to the processing requirements of subsequent cross-scale feature fusion tasks. By parallelizing the processed features across multiple attention heads, the model can simultaneously capture feature interactions at different semantic levels, significantly improving its ability to represent complex patterns. The goal of this step is to mine global semantics and combine the extracted feature maps with local features to achieve a more robust representation of steel crack defects. This process effectively models the topological correlation characteristics within the sequence (such as the texture continuity of adjacent pixels). Compared to the local constraints of traditional convolution (N×N receptive field), global attention expands the range of feature interactions to the full scale of the image (H×W). This mechanism effectively enhances the robustness of deep learning models when dealing with long-range dependencies and diverse features.

[0049]

[0050]

[0051]

[0052] in, It is a parameter variable in the training process.

[0053] Transformer's multi-head self-attention mechanism is as follows Figure 3 shown.

[0054] In one embodiment, the surface features of the steel material in step 100 are extracted from a steel material surface image extracted by a ResNet structural feature extraction network.

[0055] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0056] In one embodiment, a device for extracting micro-crack features on a steel surface is provided, comprising: a dynamic attention feature extraction module, a position encoding module, a steel surface subtle defect feature extraction module, and a steel surface defect detection module, wherein: The dynamic attention feature extraction module is used to obtain the surface features of steel and downsample the surface features of steel twice to obtain downsampled features; the downsampled features are used to capture long-range feature associations across image regions using a dynamic sparse attention mechanism to obtain dynamic attention features.

[0057] The position encoding module is used to encode the dynamic attention features and then use the dynamic position encoding strategy to perform position encoding to obtain the position encoding result.

[0058] The module for extracting subtle defect features on the steel surface is used to perform cross-scale semantic fusion on the position encoding results using the Transformer's self-attention mechanism to obtain cross-scale semantic fusion features; the cross-scale semantic fusion features are then reshaped to obtain subtle defect features on the steel surface.

[0059] The steel surface defect detection module is used to detect defects on the steel surface based on the subtle defect characteristics of the steel surface.

[0060] In one embodiment, the dynamic attention feature extraction module is also used to divide the downsampled features into discrete semantic units, and then project them through high-dimensional space to obtain Q matrix, K matrix, and V matrix; determine the adjacency matrix of the regional correlation to the Q matrix and K matrix based on the Q matrix and K matrix; dynamically screen the top k related regions of each region according to the adjacency matrix and the correlation threshold to obtain the Top-K key regions; cluster the Top-K key regions with the discrete semantic units of the K matrix and the V matrix respectively to obtain clustered key and value vectors; process the clustered key and value vectors (tensor) using the self-attention mechanism and then integrate them through the linear transformation layer to obtain the attention output; and perform a dimension reshaping operation on the attention output to obtain the dynamic attention feature.

[0061] In one embodiment, the steel surface features in the dynamic attention feature extraction module are extracted from the steel surface image extracted by the ResNet structural feature extraction network.

[0062] The specific definition of the steel surface microcrack feature extraction device can be found in the definition of the steel surface microcrack feature extraction method mentioned above, and will not be repeated here. The various modules in the above-mentioned steel surface microcrack feature extraction device can be implemented in whole or in part through software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0063] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for extracting micro-crack features from the surface of steel is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0064] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0065] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiment when executing the computer program.

[0066] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.

[0067] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0068] In a specific embodiment, the steel surface microcrack feature extraction model proposed in this method is used in a steel surface defect recognition model to extract microdefect features. The steel surface defect recognition model includes a backbone network, a steel surface microcrack feature extraction model, a feature fusion module based on dynamic pooling pyramid and GhostConv, an IoU-aware query selection strategy, and a decoder with an auxiliary prediction head.

[0069] The backbone network primarily implements feature extraction. This module integrates convolutional layers, batch normalization, and activation functions to expand the receptive field while reducing computational complexity. Its core structure utilizes a residual connection design. This residual structure not only alleviates the vanishing gradient problem of deep networks but also improves small target detection accuracy through cross-layer feature fusion. The backbone network of this study uses ResNet18, a basic residual module of the ResNet architecture series. While extracting features through convolution, it also forms residual connections through cross-layer identity mapping. The residual connection mechanism of ResNet18 allows the network to directly transmit underlying feature information. This dual-convolutional layer configuration reduces the number of parameters through a parameter sharing strategy, while cross-layer connections speed up training convergence.

[0070] The feature fusion module based on dynamic pooling pyramid and GhostConv includes: two first convolution modules, two second convolution modules and three cross-scale feature interaction modules; the first convolution module includes: a convolution layer with a convolution kernel size of 1*1, a batch normalization layer and a SiLU activation function; the second convolution module includes: a convolution layer with a convolution kernel size of 3*3, a batch normalization layer and a SiLU activation function; the specific process of feature fusion performed by the feature fusion module includes: passing the encoded feature through the first first convolution module to obtain the first convolution feature; splicing the first convolution feature with the first down-sampling feature to obtain the first splicing feature; processing the first splicing feature through the first cross-scale feature interaction module to obtain the cross-scale feature interaction features; the cross-scale feature interaction features are processed by the second first convolution module to obtain the second convolution feature; the second convolution feature is spliced with the image feature to obtain the second splicing feature; the second splicing feature is processed by the second cross-scale feature interaction module and the first second convolution module to obtain the third convolution feature; the third convolution feature is spliced with the second convolution feature to obtain the third splicing feature; the third splicing feature is processed by the third cross-scale feature interaction module and the second second convolution module to obtain the fourth convolution feature; the fourth convolution feature is spliced with the first convolution feature to obtain the fourth splicing feature; after splicing the fourth splicing feature, the third splicing feature and the second splicing feature, a fusion feature vector is obtained.

[0071] The innovative feature fusion module based on dynamic pooling pyramids and GhostConv lies in its hierarchical feature fusion mechanism and cross-level feature reuse strategy. By constructing a cascaded pooling structure, it captures multi-scale information and utilizes pooling kernels of varying granularity to extract target shape, size, and spatial distribution features. To address the issues of increased model complexity and weakened features of small targets caused by channel redundancy, the GhostConv module, with its low computational load and high feature generation efficiency, employs a dynamic feature channel screening mechanism to enhance the discriminative characterization of complex texture defects (such as oxidation spots) and geometric defects (such as microcracks) on steel surfaces.

[0072] This feature fusion module utilizes a serial-parallel hybrid pooling approach, reducing the number of parameters while maintaining multi-scale perception capabilities. It also combines feature reuse strategies to enhance the response strength to small-scale defects. Traditional approaches typically use bilinear interpolation to upsample feature maps and then perform channel cascade fusion. This operation results in high redundancy, and fine-grained information in shallow, high-resolution features is easily overwritten by deep semantic features. Therefore, our model overcomes these shortcomings while significantly reducing the number of parameters caused by previous convolutions. This module proposes a three-stage feature enhancement process: first, upsampling the feature maps enhances cross-scale feature interaction and detail information fusion. Next, a spatial pyramid pooling layer (SPPF) is used to aggregate multi-granular features, and the number of parameters is reduced by concatenating features from parallel pooling branches. Finally, the fused features are implicitly expanded through the GhostConv module, utilizing a redundant feature generation mechanism to enhance feature expression capabilities.

[0073] The cross-scale feature interaction module includes: two convolution modules, a dynamic pooling pyramid module and a GhostConv module; the first spliced feature is processed by the first cross-scale feature interaction module to obtain a cross-scale feature interaction feature, including: the first convolution feature is processed by the first convolution module to obtain a convolution feature; the convolution feature is processed by the dynamic pooling pyramid module to obtain pooled features of three different scales; the convolution feature and the pooled features of three different scales are spliced and processed by the second convolution module to obtain a pooled convolution feature; the pooled convolution feature is processed by the GhostConv module to obtain a cross-scale feature interaction feature.

[0074] The dynamic pooling pyramid module includes three maximum pooling layers; the specific process of feature extraction based on the dynamic pooling pyramid module includes: processing the convolution feature through the first maximum pooling layer to obtain the first-scale pooling feature; processing the first-scale pooling feature through the second maximum pooling layer to obtain the second-scale pooling feature; processing the second-scale pooling feature through the third maximum pooling layer to obtain the third-scale pooling feature.

[0075] Specifically, the specific workflow of the GhostConv module is as follows: First, the eigenvalue map F is generated, which is completed by conventional convolution operation as the basic feature representation. For each channel feature of the eigenvalue map F, perform i Secondary mapping operation: The first mapping operation is identity mapping, and the remaining i-1 times use lightweight transformations such as depthwise separable convolution to generate Ghost feature maps Finally, the original intrinsic feature map and the Ghost feature map are spliced in the channel dimension to form the output result The specific expression of the GhostConv module is:

[0076]

[0077]

[0078] Among them, F represents the input features of the GhostConv module, The mapping of the original feature map, Represents the convolution operation of the nth channel alone, Indicates the n A lightweight transform based on depthwise separable convolution, Represents the output features of the GhostConv module.

[0079] The loss function of the steel surface defect recognition model during training uses L1 loss and GIoU loss.

[0080] L1 loss (mean absolute error, MAE) is the core evaluation indicator in regression tasks. and the true value y i The mean absolute error (MSE) is used to quantify model bias. Compared to the L2 loss (MSE) that uses squared error, the linear penalty mechanism of the L1 loss effectively avoids error amplification and exhibits greater stability when processing noisy data. This robustness stems from the constant gradient of the loss function, making the model training process less susceptible to interference from extreme values, thereby improving generalization capabilities in scenarios with complex data distributions.

[0081]

[0082] in, represents the L1 loss, 、 Represent the true value and the corresponding predicted value respectively, and n represents the number of training samples.

[0083] The L1 loss function constructs an error measurement system based on the absolute difference between the predicted value and the true value. Its linear penalty mechanism makes the contribution of outliers to the overall loss significantly lower than the L2 loss of the squared error, thus having a natural anti-interference ability against outliers. This function maintains a fixed gradient characteristic of ±1 during the optimization process, resulting in sparse parameter updates, which is suitable for modeling scenarios that require feature screening. From the perspective of error interpretation, the mean absolute error output by the L1 loss remains consistent with the original data unit, enhancing the readability of the model evaluation results. In the image reconstruction task, its balanced optimization of pixel-level errors avoids overfitting extreme noise points and effectively balances the dialectical relationship between denoising effect and detail retention.

[0084] The Generalized Intersection over Union (GIoU) loss, an improved loss function for object detection, demonstrates significant advantages over the traditional IoU loss in bounding box regression tasks. Its core mechanism effectively addresses the vanishing gradient issue of the original IoU loss when there is no overlap between bounding box shapes by introducing a calculation parameter called the minimum enclosing rectangle (C)—the minimum enclosed area that covers both the predicted and ground-truth bounding boxes.

[0085] The GIoU loss extends its value range to the interval [-1, 1]. When the predicted box and the ground-truth box completely overlap, the metric reaches an upper limit of 1; if the two boxes are completely separated and the distance between them increases, the metric approaches -1. Compared to the traditional IoU loss, where the gradient is zeroed when the boxes do not overlap, the GIoU loss, through a compensation term designed to account for the area difference of the closure region, can still generate an effective gradient signal in the non-overlapping state. This feature enables the model to continuously obtain parameter correction directions during the bounding box regression process, significantly accelerating convergence. Mathematical derivation shows that the GIoU loss effectively improves the bounding box coordinate regression accuracy through a stable gradient propagation mechanism, demonstrating stronger optimization capabilities, especially when dealing with small object positioning and complex spatial relationships.

[0086] The steel surface defect recognition model is verified using a phased verification process: first, parameter optimization is completed based on the training set, and then verification is performed on an independent test set. In scenarios where the sample size is limited, this method implements a retention verification strategy, distributing the original samples in a 4:1 ratio, of which 80% are used for model training and 20% are used as a validation set. This embodiment will conduct experiments based on the NEU-DET and GC10-DET datasets to verify the effectiveness of the model. The AdamW optimizer is used for parameter optimization in model training, with the basic learning rate configured as 1e-4 and the momentum coefficient set to 0.9. In order to comprehensively evaluate the performance of the model, a multi-dimensional evaluation system is constructed covering core indicators such as classification accuracy, recall rate, and mean average precision (mAP@0.5 and mAP@0.5:0.95).

[0087] (1) Dataset NEU-DET is a benchmark dataset dedicated to steel surface anomaly detection, primarily serving the research fields of computer vision and deep learning algorithms. The dataset contains 1,800 standardized industrial image samples, with each frame image size standardized to 200×200 pixels. It covers six typical industrial defects: cracks, inclusions, plaques, pitted surfaces, crazing, and scratches, providing a standardized evaluation benchmark for defect classification and location algorithms. Examples of the six typical industrial defects are as follows: Figure 5 As shown, Figure 5 (a) is a sample image of a pit. Figure 5 (b) is the inclusion sample diagram, Figure 5 (c) is a sample image of a patch. Figure 5 (d) is a concave sample diagram. Figure 5 (e) is a sample diagram of rolled oxide scale. Figure 5 (f) is a scratch sample image.

[0088] GC10-DET, an open-source industrial steel inspection dataset, contains ten typical surface defect types: punch holes (Pu), weld marks (Wl), crescents (Cg), water spots (Ws), oil spots (Os), silk spots (Ss), inclusions (In), mill pits (Rp), creases (Cr), and folds (Wf). This dataset is professionally annotated, encompassing a diverse range of morphological features, from micron-scale point defects to centimeter-scale planar damage. Each defect type exhibits significant differences in geometry, size distribution, and texture complexity. Known for its high image quality and detailed annotations, this dataset provides precise bounding box defect annotations, facilitating the training and performance validation of detection models. Due to its diverse defect morphology, it is of great value in the development of industrial defect recognition algorithms, particularly in the validation of deep learning-driven detection models, where it effectively evaluates the generalization capabilities of algorithms for complex defect patterns. The distribution of samples by category is visualized in a chart. Table 1 shows the distribution of samples by category.

[0089] Table 1 Distribution of sample size in each category

[0090] (2) Experimental results 1) Model performance on the NEU-DET dataset In the NEU-DET dataset, dynamic lighting conditions and variations in material surface properties cause defect samples to exhibit grayscale characteristics. This phenomenon leads to significant morphological differences between defect samples of the same type, while defects across different classes exhibit textural similarities. This dual nature presents a dual challenge for defect recognition models: they must overcome the discretization of sample distribution within a class while also enhancing their ability to capture subtle differences between classes. Model training under this complex data distribution effectively improves industrial quality inspection systems' resistance to lighting interference and material adaptability, providing a technical validation foundation for the engineering deployment of online steel surface inspection systems. The model was trained for 250 iterations, and within these 250 training cycles, the model performance exhibited typical convergence characteristics. Rapid optimization occurred in the early stages of training (the first 50 cycles), with the GIoU loss and L1 loss dropping to 0.3757 and 0.3469, respectively. After entering the mid-training phase, the loss function plateaued, and accuracy metrics were continuously optimized through parameter fine-tuning, ultimately reaching convergence. The model maintained a detection accuracy of 92.55% and a recall of 0.7952. This training trajectory verifies the effectiveness of the gradient optimization strategy, especially in balancing the precision and recall indicators in the object detection task, showing stable performance. The experimental results on the NEU-DET dataset are as follows Figure 6 As shown, Figure 6 (a) is a diagram of training GIoU loss. Figure 6 (b) is a diagram of training L1 loss. Figure 6 (c) is a schematic diagram of the accuracy index. Figure 6 (d) is a schematic diagram of the recall rate indicator. Figure 6 (e) is the verification GIoU loss lose Schematic diagram, Figure 6 (f) is a schematic diagram for verifying L1 loss. Figure 6 (g) is a schematic diagram of the mAP50 indicator. Figure 6 (h) is a schematic diagram of the mAP@0.5:0.95 indicator.

[0091] The model demonstrated differentiated performance across six defect categories on the NEU-DET dataset, as shown in Table 2. In the classification task, the inclusion category achieved the lowest recognition accuracy (60.2%), while the recall rates for the patches, inclusions, and pitted surfaces categories peaked at over 0.9. However, the recall rate for the crazing category plummeted to 0.377 due to the fusion of light-colored features with the background. The dataset size significantly correlated with model performance, with mAP@0.5 exceeding 90% for the patches, pitted surfaces, and scratches categories, pushing the overall mAP@0.5 to 83.14%. The mAP@0.5:0.95 ratio remained stable at 47.37%. High-contrast crack samples achieved high detection accuracy thanks to their distinguishable features. Confusion matrix analysis reveals that the model has strong discrimination in fine-grained classification tasks, but is still sensitive to slight perturbations of background textures. This not only reflects the algorithm's advantage in capturing defect features, but also exposes optimization space in complex industrial scenarios. Figure 7 shown.

[0092] Table 2 Recognition accuracy of 6 categories

[0093] 2) Model performance on the GC10-DE dataset To further evaluate the performance of the proposed model, the model was validated on the GC10-DET steel defect dataset, which contains 3570 high-resolution industrial images (2048×1000 pixels) divided into training and validation sets at a ratio of 4:1. After 300 training cycles of optimization, the model achieved a detection accuracy of 79.61% and a recall rate of 0.6459. The GIOU loss and L1 loss converged to 0.5542 and 0.3359, respectively, verifying the effectiveness of the multi-task optimization mechanism. The experimental results on the GC10-DE dataset are shown in Figure 2. Figure 8 As shown, Figure 8 (a) is a diagram of training GIoU loss. Figure 8 (b) is a diagram of training L1 loss. Figure 8 (c) is a schematic diagram of the accuracy index. Figure 8 (d) is a schematic diagram of the recall rate indicator. Figure 8 (e) is a diagram for verifying GIoU loss. Figure 8 (f) is a schematic diagram for verifying L1 loss. Figure 8 (g) is a schematic diagram of the mAP50 indicator. Figure 8(h) is a schematic diagram of the mAP@0.5:0.95 metric. Experimental results demonstrate that the proposed detection framework has stable feature extraction capabilities in complex industrial scenarios, and its multi-scale feature fusion mechanism effectively balances positioning accuracy and classification performance.

[0094] Verification on the GC10-DET industrial dataset shows that the model can exhibit differentiated performance characteristics in real-world applications. In the detection of specific defect types, the weld (Wl) category leads with an accuracy of 90.6%, while the mAP@0.5 of the inclusion (In) and rolling pit (Rp) categories are less than 30%, at 29.6% and 27.7% respectively. It is worth noting that although the overall mAP@0.5 of the dataset reached 67.27%, the detection accuracy of the three types of defects, punching (Pu), weld (Wl), and crescent (Cg), exceeded 90%, confirming the model's advantage in feature extraction for high-contrast defects. This performance difference reveals that in industrial inspection scenarios, the visual separability of the target and background has a decisive influence on the performance of the model.

[0095] The sample size of the GC10-DET dataset shows a significant correlation with model detection performance. While detection accuracy for rolling pits (Rp) and creases (Cr) is limited due to sample scarcity, punch holes (Pu) and welds (Wl) achieve mAP@0.5:0.95 of 52.8% and 52.9%, respectively, significantly outperforming the overall mean of 34.09% (6). The distribution of mAP@0.5:0.95 indicates that visual separability between target and background is a key constraint. Complex textures in industrial scenes can easily cause feature confusion, particularly for low-contrast defects. This performance difference reveals that the balance of sample distribution between classes in the dataset and background complexity jointly influence model generalization. Therefore, dataset size significantly influences model parameter optimization. By increasing the data size and optimizing the distribution structure of low-sample categories such as "cracks," the model's feature separation ability in complex backgrounds is enhanced. Among ten defect detection categories, our framework achieves 79.61% accuracy on crack identification. The accuracy of inclusions (In) and pits (Rp) shows that due to the small size of the dataset, the accuracy is not significantly improved, confirming the positive effect of data distribution optimization on model robustness. The recognition accuracy for the 10 categories is shown in Table 3, and the recognition accuracy for the NEU-DET and GC10-DE datasets is shown in Table 4.

[0096] Table 3 Recognition accuracy of 10 categories

[0097] Table 4 Recognition accuracy on two datasets

[0098] (3) Comparative experiment The steel surface defect recognition module, which replaces the adaptive feature interaction encoder with a dynamic sparse attention mechanism-based feature encoding module, was compared with current state-of-the-art models on the NEU-DET dataset. Experiments show that the proposed model exhibits significant advantages in accuracy. The experimental results on the NEU-DET dataset are shown in Table 5.

[0099] Table 5 Experimental results on the NEU-DET dataset

[0100] After comparing the NEU-DET dataset, we conducted supplementary testing on the GC10-DET dataset. The experimental results on the GC10-DET dataset are shown in Tables 6 and 7. The stability advantage of the GC10-DET dataset further verifies its comprehensive reliability.

[0101] Table 6 Recognition accuracy of each series of YOLO models

[0102] Table 7 Recognition accuracy of YOLO5 series models

[0103] (3) Ablation and visualization experiments In this embodiment, an ablation experiment is conducted on NEU-DET to verify the performance contribution of each module. The results of the ablation experiment on NEU-DET are shown in Table 8.

[0104] Baseline Model A combines a ResNet18 backbone with the DETR architecture, achieving 79.40% mAP@0.5 and 43.55% mAP@0.5:0.95. Model B builds on Model A by introducing the DAS module proposed in this study and integrating a dynamic sparse attention mechanism to further optimize detection performance. While increasing the parameter size by 0.21M, this solution improves mAP@0.5 and mAP@0.5:0.95 by 3.27% and 1.45%, respectively, compared to the baseline model by enhancing the image information of the associated toke. This structural improvement significantly improves performance while maintaining detection accuracy without significantly increasing the overall computational load.

[0105] Model C introduces the GhostConv mechanism into the architecture of Model B, integrating the dynamic pooling pyramid mechanism with GhostConv. This design, through the synergistic effect of the dynamic pooling pyramid mechanism and GhostConv, effectively suppresses parameter growth while strengthening feature interactions, achieving simultaneous optimization of detection accuracy and model efficiency. This solution improves mAP@0.5 and mAP@0.5:0.95 to 83.14% and 47.37%, respectively. Ablation experiment data shows that removing any core component causes a step-by-step decrease in accuracy, confirming that the synergy between the dynamic pooling pyramid mechanism and GhostConv has a decisive impact on model performance.

[0106] Model D, built on the architecture of Model C, was designed to verify the impact of stacked encoding layers on performance. Experiments found that excessively stacking Transformer encoding layers can lead to performance degradation, manifested as a significant 0.77% decrease in mAP@0.5 and a 2.57 percentage point decrease in mAP@0.5:0.95. This reverse optimization phenomenon not only confirms the rationale of the original architecture but also provides critical experimental evidence for subsequent optimization paths.

[0107] In the verification of CNN-based module E, we used ResNet50 as the underlying architecture of the convolutional embedding framework. Experiments found that while this solution significantly increased the number of parameters and computational overhead on the NEU-DET dataset, the improvement in detection accuracy was limited. Quantitative analysis showed that the detection accuracy metric (mAP@0.5:0.95) only improved by 1.23% compared to baseline model A, failing to achieve simultaneous performance optimization. This also indicates that excessive parameter count and computational complexity will undermine the model's value for engineering deployment. This study ultimately selected ResNet18 as the feature extraction backbone, as this architecture achieves a better balance between computational efficiency and detection accuracy.

[0108] Table 8 Ablation experiment results

[0109] After completing the module-level accuracy assessment, a visual comparison of the detection results is conducted on the NEU-DET dataset. The model after module replacement in this example not only achieves accurate classification and location of crack defects, but also demonstrates superior performance on typical samples. Taking pit detection as an example, while the baseline model can complete defect classification, this solution offers significant improvements in location accuracy and feature discrimination. The introduction of a dynamic sparse attention mechanism significantly expands the detection coverage. In the GhostConv ensemble model, this combined architecture demonstrates enhanced environmental adaptability, particularly in inclusion scenarios, where crack recognition accuracy is systematically improved. Experimental data shows that the added modules effectively enhance the ability to suppress complex background interference, enabling the algorithm to successfully capture previously missed micro-crack features. For patch detection scenarios, the baseline model suffers from the dual flaws of detection redundancy and insufficient area coverage. In pitted surface detection, the optimized modules demonstrate a beneficial effect: the number of redundant candidate boxes is reduced while maintaining detection accuracy. Visualization confirms that the architecture has the ability to accurately characterize the geometric characteristics of defects, especially the reduction of recognition errors in crack distribution range and morphological characteristics, highlighting the breakthrough progress of the module in feature space optimization.

[0110] To more clearly verify the recognition capabilities of our model, we also visualized its recognition capabilities on the GC10-DET dataset. Through the visualization of water stains (Ws), creases (Cr), and a mixture of different defects, we can see that with the superposition of our modules, the recognition of different types of defects becomes more accurate. This can be seen from the following aspects: as the modules are superimposed, the recognition redundant boxes of several defects are reduced, and the recognition scores are significantly increased, finally achieving the final excellent results. This progressive optimization mechanism not only enhances the localization accuracy of small defects, but also effectively suppresses background noise interference in industrial scenarios through cross-layer feature fusion, ultimately verifying the model's robust generalization ability in multi-source data scenarios.

[0111] Comprehensive experimental results confirm that the model after module replacement in this embodiment demonstrates significant advantages in complex industrial scenarios: enhanced robustness to background noise and improved sensitivity to identifying subtle defects. Visual analysis validates the algorithm's effectiveness in detecting multiple defect types, particularly reducing the localization error of weak texture defects.

[0112] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0113] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for extracting micro crack features on the surface of steel, characterized in that: The method comprises: Acquire surface features of the steel material, and downsample the surface features of the steel material twice to obtain downsampled features; The downsampled features are subjected to a dynamic sparse attention mechanism to capture long-range feature correlations across image regions, thereby obtaining dynamic attention features; After encoding the dynamic attention feature, position encoding is performed using a dynamic position encoding strategy to obtain a position encoding result; The position encoding result is subjected to cross-scale semantic fusion using the self-attention mechanism of Transformer to obtain a cross-scale semantic fusion feature; Reshaping the cross-scale semantic fusion features to obtain subtle defect features on the steel surface; Defect detection is performed on the steel surface based on the subtle defect characteristics of the steel surface.

2. The method for extracting micro crack features on steel surface according to claim 1, characterized in that: The dynamic sparse attention mechanism captures long-range feature correlations across image regions and obtains dynamic attention features, including: Divide the downsampled features into discrete semantic units, and then project them through high-dimensional space to obtain Q matrix, K matrix, and V matrix; According to the Q matrix and the K matrix, an adjacency matrix of regional correlations to the Q matrix and the K matrix is determined; Dynamically screening the top k relevant regions of each region according to the adjacency matrix and the correlation threshold to obtain the Top-K key regions; Aggregating the Top-K key regions with the discrete semantic units of the K matrix and the V matrix respectively to obtain vectors of aggregated keys and values; The aggregated key and value vectors are processed using the self-attention mechanism and then integrated through a linear transformation layer to obtain the attention output; The attention output is subjected to a dimension reshaping operation to obtain a dynamic attention feature.

3. The method for extracting micro crack features on steel surface according to claim 1, characterized in that: The steel surface features are extracted from the steel surface image extracted by the ResNet structural feature extraction network.

4. A device for extracting micro crack features on the surface of steel, characterized in that: The device comprises: A dynamic attention feature extraction module is used to obtain steel surface features and downsample the steel surface features twice to obtain downsampled features; the downsampled features are used to capture long-range feature correlations across image regions using a dynamic sparse attention mechanism to obtain dynamic attention features; A position encoding module is used to encode the dynamic attention feature and then perform position encoding using a dynamic position encoding strategy to obtain a position encoding result; A steel surface subtle defect feature extraction module is used to perform cross-scale semantic fusion on the position encoding results using the Transformer's self-attention mechanism to obtain cross-scale semantic fusion features; and to reshape the cross-scale semantic fusion features to obtain subtle defect features on the steel surface; The steel surface defect detection module is used to detect defects on the steel surface based on the subtle defect characteristics of the steel surface.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Cited By

  • Steel plate surface defect detection system based on space-time mutual attention and sparse space-time perception attention

    CN121213558A

  • A steel plate surface defect detection system based on space-time mutual attention and sparse space-time perception attention

    CN121213558B

  • Efficient two-step multi-scale feature extraction method

    CN121214196A