A steel surface defect detection method and system based on multi-scale feature extraction
By combining local and global receptive field feature extraction modules, large and small scale feature enhancement modules, and morphological enhancement modules, the problems of insufficient accuracy in small-scale defect detection and insufficient efficiency in large-scale defect recognition in traditional methods are solved, and efficient and accurate recognition of steel surface defects is achieved, thereby improving production efficiency and quality.
Patent Information
- Application Number
- CN202511050837.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-29
AI Technical Summary
While traditional steel surface defect detection methods based on multi-scale feature extraction improve the accuracy of small-scale defect detection, they ignore the recognition efficiency and real-time performance of large-scale defects, resulting in poor scale robustness of the detection model. This makes it difficult to achieve rapid real-time detection on industrial production lines, affecting production efficiency.
The system adopts feature extraction modules based on local and global receptive fields, large and small scale feature enhancement modules and morphological enhancement modules. Through multi-scale feature extraction and enhancement, it takes into account the recognition of small-scale and large-scale defect features and improves the balance between detection accuracy and speed.
It achieves efficient identification of steel surface defects, improves the efficiency of identifying large-scale defects, reduces the false detection rate in the production process, and improves the overall quality and production efficiency of steel products.
Smart Images

Figure CN120563502B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, and in particular to a steel surface defect detection method and system based on multi-scale feature extraction. BACKGROUND
[0002] Traditional steel surface defect detection methods based on multi-scale feature extraction mostly focus on improving the detection accuracy of small-scale defects, while relatively ignoring the recognition efficiency of large-scale defects, resulting in poor scale robustness of the defect detection model. In fact, large-scale defects (such as cracks, indentations, etc.) often have a more significant impact on the overall quality of steel products, and their missed detection will cause serious quality risks and substantial economic losses to enterprises.
[0003] At the same time, traditional methods generally focus on improving detection accuracy, while relatively ignoring real-time considerations, making it difficult to achieve fast real-time defect detection on production lines in industrial actual environments, thereby restricting the improvement of production efficiency. SUMMARY
[0004] The present application provides a steel surface defect detection method and system based on multi-scale feature extraction, which can balance the recognition of small-scale defect features and large-scale defect features, and achieve the balance between production efficiency and recognition accuracy.
[0005] The first aspect of the present application provides a steel surface defect detection method based on multi-scale feature extraction, the method comprising: obtaining a first feature map generated according to defect features of a steel defect image; converting the first feature map into a first shallow feature map that captures multi-scale steel defect features through a feature extraction module based on local and global receptive fields; converting the first shallow feature map into a second shallow feature map that enhances large and small scale steel defect features through a large and small scale feature enhancement module; extracting multi-scale steel defect features of the second shallow feature map through a feature extraction module based on local and global receptive fields; and converting the second shallow feature map after feature extraction into a deep feature map that enhances morphological features through a morphological enhancement module, wherein the first shallow feature map, the second shallow feature map and the deep feature map are all used for detecting steel surface defects.
[0006] In some embodiments of the first aspect, converting the first feature map into a first shallow feature map that captures multi-scale steel defect features through a feature extraction module based on local and global receptive fields comprises:
[0007] The first feature map is input into two parallel 1x1 convolution layers for dimension reduction after being extracted by the depth separable convolution layer to obtain a first feature map after dimension reduction; the first feature map after dimension reduction is passed through a plurality of parallel filter groups to obtain an original feature map, a small-scale defect feature map, a medium-scale defect feature map and a large-scale defect feature map corresponding to the first feature map after dimension reduction; the original feature map, the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map are connected in residual connection with the first feature map after dimension reduction to obtain a second feature map; the first feature map after dimension reduction and the second feature map are fused and connected in residual connection with the first feature map to obtain a first shallow feature map.
[0008] In some embodiments of the first aspect, the original feature map, the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map corresponding to the first feature map after dimension reduction are obtained by a plurality of parallel filter groups, including: an original feature map is generated by a first filter by retaining the global context relationship of the original feature information of the first feature map after dimension reduction; a small-scale defect feature map is generated by a second filter by performing a 3x3 convolution operation on the first feature map after dimension reduction; a medium-scale defect feature map is generated by a third filter by performing a 3x3 convolution operation on the first feature map after dimension reduction and the small-scale defect feature map after element-wise addition; and a large-scale defect feature map is generated by a fourth filter by performing a 3x3 convolution operation on the first feature map after dimension reduction and the medium-scale defect feature map after element-wise addition.
[0009] In some embodiments of the first aspect, the first shallow feature map is converted into a second shallow feature map with enhanced size-scale steel defect features by the size-scale feature enhancement module, including: a third feature map capturing semantic information of large-scale defect features and a fourth feature map capturing small-scale defect features are respectively generated by inputting the first shallow feature map into parallel reparameterization convolution layers and receptive field attention convolution layers; a fifth feature map focusing on large-scale defect features is obtained by passing the third feature map through a channel squeeze excitation module with a first compression ratio value; two first sub-feature maps are obtained by uniformly dividing the fourth feature map, and the two first sub-feature maps are respectively input into channel squeeze excitation modules with second and third compression ratio values, and two second sub-feature maps focusing on small-scale defect features are output; and the second shallow feature map is obtained by fusing the fifth feature map and the two second sub-feature maps.
[0010] In some embodiments of the first aspect, the compression ratio values are in descending order of the first compression ratio value, the second compression ratio value and the third compression ratio value.
[0011] In some embodiments of the first aspect, before the channel extrusion excitation module inputting the two channels with the second compression ratio value and the third compression ratio value respectively, the method further comprises: inputting the two first sub-feature maps into two 1X1 convolution layers respectively to enhance the representation ability of the two first sub-feature maps to the underlying texture information of small-scale defect features.
[0012] In some embodiments of the first aspect, the second shallow feature map after feature extraction is converted into a deep feature map with enhanced morphological features by the morphological enhancement module, comprising: inputting the second shallow feature map after feature extraction into a plurality of feature extraction modules to output a morphological enhancement feature map, the morphological enhancement feature map being used to capture scratch edge features of small-scale defect features and crack direction features of large-scale defect features; element-wise adding the second shallow feature map after feature extraction and the morphological enhancement feature map to obtain an added feature map; performing self-adaptive calibration on the channel weights of the added feature map by the high-efficiency channel attention module to obtain a calibrated feature map; element-wise multiplying the calibrated feature map and the added feature map to obtain the deep feature map with enhanced morphological features.
[0013] In some embodiments of the first aspect, inputting the second shallow feature map after feature extraction into a plurality of feature extraction modules to output a morphological enhancement feature map comprises: extracting scratch edge features and crack direction features of the second shallow feature map after feature extraction by a 7X7 depth separable convolution layer; enhancing global context information and spatial relationships of the scratch edge features and the crack direction features by a multi-layer perceptron module constructed by a linear transformation layer and a GELU function to generate an enhanced second shallow feature map; and element-wise adding the enhanced second shallow feature map and the second shallow feature map after feature extraction to generate the morphological enhancement feature map.
[0014] In some embodiments of the first aspect, the method further comprises: enhancing the size-scale defect features of the deep feature map by a size-scale feature enhancement module; and further enhancing the morphological features of the deep feature map by a morphological enhancement module.
[0015] The second aspect of the present application provides a steel surface defect detection system, comprising: a processor and a memory; the memory is coupled with the processor, and the memory is used to store computer program code, and the processor invokes the computer program code to enable the steel surface defect detection system to execute the method of the first aspect.
[0016] It can be understood that the steel surface defect detection method and system based on multi-scale feature extraction can improve the overall quality of steel products by setting lightweight modules for extracting and enhancing large-scale defect features and small-scale defect features of the steel surface, such as a feature extraction module based on local and global receptive fields, a size-scale feature enhancement module, and a morphology enhancement module, taking into account the identification of large-scale and small-scale defects, and improving the identification efficiency of large-scale defects. At the same time, since the three modules are lightweight designs, the defect detection model can improve the real-time detection efficiency on the basis of maintaining high detection accuracy, achieving a balance between accuracy and speed. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.
[0018] Figure 1 An application scenario diagram of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0019] Figure 2 A flowchart of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0020] Figure 3 An application scenario diagram of the feature extraction module based on local and global receptive fields of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0021] Figure 4 An application scenario diagram of the size-scale feature enhancement module of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0022] Figure 5 An application scenario diagram of the morphology enhancement module of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0023] Figure 6 Another application scenario diagram of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0024] Figure 7 A detection result diagram of the steel surface defect detection method based on multi-scale feature extraction provided by the embodiments of the present application.
[0025] Figure 8 A structure diagram of the steel surface defect detection system provided by the embodiments of the present application.
[0026] The specific embodiments of the present application have been shown by the above-described drawings, and will be described in more detail hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0027] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. In the following description, the same numbers refer to the same elements throughout the drawings, unless otherwise represented. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application.
[0028] The terms "first", "second", and the like used herein are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated.
[0029] Referring to Figure 1 , Figure 1 The feature extraction branch of the defect detection model applied by the steel surface defect detection method based on multi-scale feature extraction provided by the present application comprises at least a plurality of convolution layers, a feature extraction module based on local and global receptive fields, a size scale feature enhancement module, and a morphology enhancement module. The steel defect image sequentially passes through the above-mentioned modules of the feature extraction branch, and the first shallow feature map, the second shallow feature map, and the deep feature map can be extracted. The first shallow feature map, the second shallow feature map, and the deep feature map can be used to detect the steel surface defects.
[0030] The technical solutions of the present application and how the technical solutions of the present application solve the technical problems will be described in detail in the following specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments.
[0031] Referring to Figure 1 and Figure 2 , Figure 2 A flowchart of the steel surface defect detection method based on multi-scale feature extraction provided by the present application is shown. The execution subject of the detection method can be a steel surface defect detection system. As Figure 2 shown, the detection method can include the following steps:
[0032] Step S110: obtaining a first feature map generated according to defect features of a steel defect image.
[0033] Specifically, the steel defect image is obtained, and then the defect features of the steel defect image are extracted by a plurality of convolution layers to generate a first feature map.
[0034] Step S120: converting the first feature map into a first shallow feature map capturing multi-scale steel defect features by the local and global receptive field based feature extraction module.
[0035] Specifically, the local and global receptive field based feature extraction module is as shown in FIG. 2. Figure 3 Firstly, the first feature map is input into two parallel 1x1 convolution layers after extracting multi-scale defect features by a 3x3 depth separable convolution layer, to obtain a reduced dimension first feature map. Then, the reduced dimension first feature map is input into a full-channel parallel residual module to output a second feature map. Finally, the second feature map and the reduced dimension first feature map are fused by a concatenation layer and then connected with the first feature map by a 1x1 convolution layer to obtain a first shallow feature map.
[0036] It can be understood that, in sequence, the 3x3 depth separable convolution with a hollow convolution can be directly used on the first feature map to obtain multi-scale defect features such as scratches, inclusions and cracks. Then, through a multi-branch parallel path (the multi-branch parallel path includes two parallel 1x1 convolution layers), the gradient can be effectively propagated, and the representation ability of small-scale defect features such as scratches can be enhanced. Among them, the full-channel parallel residual module is further used to extract multi-scale defect features. It can be understood that the full-channel parallel residual module can enhance the perception ability of different scale defects (small-scale defect fine scratches and large-scale defect patch regions). Finally, the residual connection is used to retain the shallow features (scratch edges, inclusion outlines, etc.) of the original input defect map (i.e. the first feature map), and then a high-quality feature map (i.e. the first shallow feature map) with multi-scale defect information is generated.
[0037] In one embodiment, after the reduced dimension first feature map is input into the full-channel parallel residual module, steps a and b are performed in the full-channel parallel residual module.
[0038] a. After the reduced dimension first feature map is input into the full-channel parallel residual module, steps a and b are performed in the full-channel parallel residual module.
[0039] In one example of step a, there are four groups of filters in the full-channel parallel residual module, a first filter, a second filter, a third filter and a fourth filter. As shown in FIG. 3, the first filter, the second filter, the third filter and the fourth filter are respectively connected with the reduced dimension first feature map, the original feature map corresponding to the first feature map, the small-scale defect feature map and the large-scale defect feature map. Figure 3As shown, the global context relationship of the original feature information of the first feature map after dimensionality increasing is reserved through the first filter X1 to generate the original feature map. The first feature map after dimensionality increasing is subjected to 3x3 convolution operation through the second filter X2 to generate the small-scale defect feature map. The first feature map after dimensionality increasing and the small-scale defect feature map are added element by element through the third filter X3 and then subjected to 3x3 convolution operation to generate the medium-scale defect feature map. The first feature map after dimensionality increasing and the medium-scale defect feature map are added element by element through the fourth filter X4 and then subjected to 3x3 convolution operation to generate the large-scale defect feature map.
[0040] The expressions of the four groups of filters are as follows:
[0041]
[0042] wherein, represents the output feature map of each group of filters, and the output feature map includes the original feature map, the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map. represents the input feature map of each group of filters, and the input feature map is the first feature map after dimensionality reduction. .
[0043] As can be understood, the multi-scale defect features are extracted in parallel through the parallel design of the four groups of filters, such as Figure 3 As shown in FIG. 3, X1 is directly output as Y1 to reserve the global context information of the original feature (inclusions, etc.), X2 is subjected to 3x3 convolution to generate Y2 for capturing small-scale defect features such as scratch edges. Y2 is added element by element with X3, and then subjected to 3x3 convolution to generate Y3 to obtain medium-scale defect features (such as punching). Y3 is added element by element with X4, and then subjected to 3x3 convolution to generate Y4 for extracting large-scale defect features such as cracks.
[0044] b. The original feature map, the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map are subjected to residual connection with the first feature map after dimensionality reduction to obtain the second feature map.
[0045] As can be understood, the multi-scale feature maps (i.e. the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map) after fusion are added element by element with the original input feature map (i.e. the first feature map) in a residual connection manner to output the enhanced feature map (i.e. the second feature map). Compared with the first feature map, the edge information in the second feature map is clearer, and the defect morphological features are more accurate, thereby reducing the false detection rate in the production process.
[0046] It can be understood that the design of the feature extraction module based on local and global receptive fields can capture global feature information such as cracks of steel surface defects and local information such as inclusion profiles while reducing the parameter amount, thereby enhancing the recognition ability of the defect detection model to multi-scale defects, and is a lightweight and efficient feature extraction module.
[0047] Step S130: converting the first shallow feature map into a second shallow feature map with enhanced size and scale steel defect features through the size and scale feature enhancement module.
[0048] Specifically, the module schematic of the size and scale feature enhancement module is as shown in Figure 4 That is, step S130 includes the following steps:
[0049] Step S131: inputting the first shallow feature map into the parallel re-parameterized convolution layer and receptive field attention convolution layer to respectively generate a third feature map capturing semantic information of large-scale defect features and a fourth feature map capturing small-scale defect features.
[0050] Exemplarily, the kernel size of the re-parameterized convolution layer (Re-parameterized Convolution, RefConv) is 3, and the stride is 2. The kernel size of the receptive field attention convolution layer (Receptive Field Attention Convolution, RFAConv) is 1, and the stride is 1. It can be understood that the re-parameterized convolution layer is introduced to refocus the original convolution kernel weight distribution to expand the effective receptive field and capture semantic information of large-scale defects such as patch regions, thereby improving the learning ability of the defect detection model to defect features in a complex background. The receptive field attention convolution layer is introduced to focus on local features of small-scale defects such as fine scratches with a smaller receptive field.
[0051] Step S132: transforming the third feature map into a fifth feature map focusing on large-scale defect features through the channel squeeze and excitation module set to a first compression ratio value.
[0052] Step S133: uniformly dividing the fourth feature map into two first sub-feature maps, and inputting the two first sub-feature maps into two 1X1 convolution layers respectively to enhance the representation ability of the two first sub-feature maps to the underlying texture information of small-scale defect features.
[0053] Step S134: inputting the two first sub-feature maps into the channel squeeze and excitation modules set to a second compression ratio value and a third compression ratio value respectively, and outputting two second sub-feature maps focusing on small-scale defect features.
[0054] Exciation (SE) module sequentially performs global average pooling, connection through a fully connected layer (FC), uses a Relu function and a Sigmoid function operation on the first sub-feature map, enhances the first sub-feature map, and outputs a second sub-feature map after element-wise multiplication of the enhanced first sub-feature map and the first sub-feature map.
[0055] Exemplarily, the first compression ratio is 8, the second compression ratio value is 7, and the third compression ratio value is 6. By setting the channel squeeze and excitation module with different compression ratios, large-scale and small-scale defect features are focused respectively. It can be understood that the RefConv branch focuses on large-scale defect features by refocusing the original convolution kernel weight distribution, and the compression ratio value is set to be relatively large to reduce the computational complexity, so as to balance the performance and efficiency (the computational complexity of RefConv itself is greater than that of RFAConv). The compression ratio of the two branches of RFAConv is set to be relatively small in order to output more bottom texture information. Specifically, the computational complexity of RFAConv itself is not large, and the compression ratio is set to be relatively small so that it can output more bottom texture information, and the two branches are set to focus on small-scale defects while reducing the computational complexity, thereby balancing the performance and efficiency.
[0056] Step S135: fuse the fifth feature map and the two second sub-feature maps to obtain a second shallow feature map through 1X1 convolution operation.
[0057] It can be understood that the related algorithm significantly loses the edge information of the steel defect in the down-sampling process. The large and small scale feature enhancement module of the present application, on the basis of lightweight design, not only maintains the focusing ability of the high-level semantic information of the large-scale defect features such as plaque regions in the down-sampling process, but also effectively compensates for the loss of bottom texture information by enhancing the edge information of small-scale defects such as fine scratches, thereby improving the accuracy of the defect detection model in positioning different scale defect features of the steel, and effectively reducing the false detection rate in the production process.
[0058] Step S140: extract the multi-scale steel defect features of the second shallow feature map through the feature extraction module based on local and global receptive fields. For specific extraction methods, refer to step S120, which will not be described here.
[0059] Step S150: convert the second shallow feature map extracted through the feature extraction module into a deep feature map with enhanced morphological features through the morphological enhancement module.
[0060] Specifically, the module schematic diagram of the morphological enhancement module is as follows Figure 5As shown, the second shallow feature map after feature extraction is input into multiple feature extraction modules (the multiple feature extraction modules can exist in a series mode), and a morphological enhancement feature map is output, which is used to capture scratch edge features of small-scale defect features and crack direction features of large-scale defect features. Then, the second shallow feature map after feature extraction and the morphological enhancement feature map are added element by element to obtain an added feature map, and the channel weights of the added feature map are adaptively calibrated by an efficient channel attention (ECA) module to obtain a calibrated feature map. Finally, the calibrated feature map and the added feature map are multiplied element by element to obtain a deep feature map.
[0061] It can be understood that the morphological enhancement module is a lightweight ConvNeXt module, which can effectively simulate the powerful feature extraction of ConvNeXt, and further reduce the complexity of the defect detection model through lightweight design, and effectively balance the performance and efficiency with only a small increase in the number of parameters.
[0062] In an embodiment, inputting the second shallow feature map after feature extraction into multiple feature extraction modules to output a morphological enhancement feature map includes: extracting scratch edge features and crack direction features of the second shallow feature map after feature extraction by a 7x7 depth separable convolution layer. Then, a multilayer perceptron (MLP) module constructed by a linear transformation layer and a GELU function is used to enhance the global context information and spatial relationship of the scratch edge features and the crack direction features, and generate an enhanced second shallow feature map. Finally, the enhanced second shallow feature map and the second shallow feature map after feature extraction are added element by element to generate the morphological enhancement feature map.
[0063] It can be understood that using a 7x7 depth separable convolution layer in the feature extraction module expands the effective receptive field while keeping the computational complexity low, and simultaneously captures defect features of different scales such as scratch edges and crack directions. Then, the morphological features of different defects are enhanced by the multilayer perceptron module constructed by the linear transformation layer and the GELU function, thereby improving the recognition ability of multi-scale defects (scratches and cracks). In order to preserve the shallow feature (scratch edge) information of the original input, the second shallow feature map after feature extraction is added element by element with the morphological enhancement feature map output by the feature extraction module and the enhanced second shallow feature map output by the multilayer perceptron module. Finally, the efficient channel attention module is introduced to adaptively calibrate the channel weights. While keeping the integrity of the features, the recognition ability of the defect detection model for key defect feature channel information is significantly improved, thereby achieving higher detection accuracy in complex texture backgrounds.
[0064] It can be understood that the design of the morphology enhancement module can globally inject semantic information of large-scale defect features such as cracks while clearly retaining small-scale defect feature characteristics such as scratch edge pixels, enhance the morphological characteristics of different defects, and further reduce the false detection rate in the production process.
[0065] In the technical solution, the local and global receptive field based feature extraction module, the size and scale feature enhancement module, and the morphology enhancement module are used to extract and enhance the large-scale and small-scale defect features of the steel surface, the identification of the large-scale and small-scale defects is considered, the identification efficiency of the large-scale defects is improved, and the overall quality of the steel product is improved. At the same time, since the three modules are lightweight designs, the defect detection model can improve the real-time detection efficiency on the basis of maintaining high detection accuracy, and balance the accuracy and speed.
[0066] Please refer to Figure 6 In some embodiments, the number of the local and global receptive field based feature extraction module, the size and scale feature enhancement module, and the morphology enhancement module can be set according to actual needs. The following further illustrates the steel surface defect detection method based on multi-scale feature extraction in an application scenario. The steel surface defect detection method based on multi-scale feature extraction includes the following steps:
[0067] Step S1: obtaining a steel defect image.
[0068] Step S2: extracting defect features of the steel defect image through two convolution layers (Conv) to generate a first feature map.
[0069] Step S3: converting the first feature map into a first shallow feature map that captures multi-scale steel defect features through the local and global receptive field based feature extraction module (also named as LGF), and inputting the first shallow feature map into a feature fusion branch;
[0070] Step S4: converting the first shallow feature map into a second shallow feature map that enhances the size and scale steel defect features through the size and scale feature enhancement module (also named as UID), and inputting the second shallow feature map into the feature fusion branch;
[0071] Step S5: extracting multi-scale steel defect features of the second shallow feature map through the local and global receptive field based feature extraction module (LGF) and inputting the second shallow feature map into the feature fusion branch;
[0072] Step S6: converting the second shallow feature map extracted through the feature extraction into a deep feature map that enhances the morphological features through the morphology enhancement module (also named as LCN), and inputting the deep feature map into the feature fusion branch;
[0073] Step S7: Enhance the size scale defect features of the deep feature map by a size scale feature enhancement module (UID) and input the feature fusion branch.
[0074] Step S8: Further enhance the morphological features of the deep feature map by a morphological enhancement module (LCN) and input the feature fusion branch.
[0075] Step S9: After the first shallow feature map, the second shallow feature map and the deep feature map input into the feature fusion branch are fused, input the detection module for detection.
[0076] It can be understood that the feature maps output by the LGF module, the UID module or the LCN module can all enter the feature fusion branch and then enter the detection module for detection. (That is, the feature maps output by each module can be used for detection separately.) The optimal way is to sequentially pass the feature maps of the LGF module, the UID module and the LCN module (i.e. the feature maps input into the feature fusion branch in step S6 and subsequent steps). The feature maps sequentially passing through the LGF, UID and LCN modules are deep feature maps after feature extraction, and the deep feature maps are fused with the shallow feature maps of the previous LGF and LCN, so as to better retain the defect information.
[0077] As shown in Figure 7 , the defect detection model applying the method of the present application scenario can recognize the image as shown in Figure 7 .
[0078] Table 1 is the ablation experiment result of the method shown in the present application scenario on the NEU-DET dataset:
[0079]
[0080] From the above experimental results, it can be seen that the LGF, UID and LCN modules can effectively improve the model scale robustness when used alone, combined in pairs and combined as a whole. Among them, the overall combination of LGF, UID and LCN achieves the optimal balance between precision and speed.
[0081] Table 2 is the comparison experiment result of the method shown in the present application scenario with the mainstream SOTA model on the NEU-DET dataset:
[0082]
[0083] From the above experimental results, it can be observed that the method of the present application is compared with the existing mainstream SOTA model, and the optimal level is obtained on the mAP50, mAP50:95, mAPs and mAPl indicators. After comprehensively considering the detection accuracy, reasoning speed and actual application deployment feasibility of the three factors, the method of the present application performs best in comprehensive performance, can significantly enhance the model scale robustness, and achieve the optimal balance of precision and speed.
[0084] Table 3 Comparison experimental results of the method shown in the application scenario with the mainstream SOTA model on the GC10-DET data set:
[0085]
[0086] From the above experimental results, it can be seen that the method of the present application is compared with the existing mainstream SOTA model on the more challenging GC10-DET data set, and it is far ahead of the SOTA model in mAP50 and mAPl indicators; the difference between the remaining indicators and the optimal value is small, after comprehensively considering the detection accuracy, reasoning speed and actual application deployment feasibility of the three factors, the method of the present application performs best in comprehensive performance, significantly improves the model scale robustness, and achieves the optimal balance of precision and speed. This experimental result also proves that the method of the present application has stronger generalization ability than the mainstream SOTA model in the actual application of steel defect detection.
[0087] In the above table, mAP50 and mAP50:95 represent the average detection accuracy indicators. Among them, mAP50 represents the mAP value calculated when the intersection over union (IOU) of the predicted box and the real box is greater than 0.5, which belongs to a relatively loose evaluation standard. And mAP50:95 is the average value calculated between IOU threshold from 0.5 to 0.95 (step size is 0.05), which is a more strict evaluation standard, which can more comprehensively reflect the detection ability of the defect detection model applied based on the multi-scale feature extraction steel surface defect detection method of the present application in different scenarios. mAPs, mAPm and mAPl correspond to the average detection indicators of small, medium and large scale defects respectively. Params is the parameter quantity of the model, and FPS is the reasoning speed.
[0088] Figure 8 The structure diagram of the steel surface defect detection system provided by the present application is shown. As shown in Figure 8 The steel surface defect detection system 10 comprises:
[0089] A processor 11, a memory 12 and a bus 13;
[0090] The memory 12 is used to store the computer program code of the processor 11;
[0091] The processor 11 is configured to execute the steel surface defect detection method based on multi-scale feature extraction in any of the preceding method embodiments by executing the computer program code.
[0092] Optionally, the memory 12 can be independent or integrated with the processor 11.
[0093] The memory 12 is connected with the processor 11 through the bus 13 and completes mutual communication.
[0094] Optionally, the memory 12 can include random access memory (RAM) and can also include non-volatile memory such as at least one disk memory.
[0095] The bus 13 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, only one thick line is shown in the figure, but it does not mean that there is only one bus or only one type of bus.
[0096] The processor described above can be a general-purpose processor including a central processing unit CPU, a network processor NP, etc.; can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0097] The steel surface defect detection system 10 is used to execute the technical solutions provided in any of the preceding method embodiments, and has similar implementation principles and technical effects, which will not be described here.
[0098] The application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by the processor, the steel surface defect detection method based on multi-scale feature extraction is realized.
[0099] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The foregoing program can be stored in a computer readable storage medium. The program executes to perform the steps of the above-mentioned method embodiments; and the foregoing storage medium includes various storage media that can store program codes, such as ROM, RAM, magnetic disk or optical disk.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A steel surface defect detection method based on multi-scale feature extraction, characterized in that the method include: Acquire a first feature map generated according to defect features of the steel defect image; After extracting multi-scale defect features from the first feature map through a depthwise separable convolutional layer, the first feature map is input into two parallel 1×1 convolutional layers for dimensionality reduction to obtain the first feature map after dimensionality reduction. The channel dimensionally upgrades the first feature map after dimensionality reduction and passes through multiple sets of parallel filters to obtain the original feature map, the small-scale defect feature map, the medium-scale defect feature map and the large-scale defect feature map corresponding to the first feature map after dimensionality upgrade; Performing residual connection on the original feature map, the small-scale defect feature map, the medium-scale defect feature map, and the large-scale defect feature map with the first feature map after dimensionality reduction to obtain a second feature map; After fusing the first and second feature maps after dimensionality reduction, a residual connection is performed with the first feature map to obtain a first shallow feature map; Inputting the first shallow feature map into a parallel reparameterized convolutional layer and a receptive field attention convolutional layer to generate a third feature map capturing the semantic information of large-scale defect features and a fourth feature map capturing small-scale defect features, respectively; Transforming the third characteristic map into a fifth characteristic map focusing on large-scale defect features through a channel squeezing excitation module set to a first compression ratio value; Evenly dividing the fourth feature map into two first sub-feature maps, and inputting the two sub-feature maps into a channel squeezing excitation module with compression ratios set to a second compression ratio value and a third compression ratio value, respectively, to output two second sub-feature maps focusing on small-scale defect features; Fusing the fifth feature map and the two second sub-feature maps to obtain a second shallow feature map; The multi-scale steel defect features of the second shallow feature map are extracted by a feature extraction module based on local and global receptive fields; Inputting the second shallow feature map after feature extraction into multiple feature extraction modules to output a morphologically enhanced feature map, wherein the morphologically enhanced feature map is used to capture scratch edge features of small-scale defect features and crack direction features of large-scale defect features; Adding the second shallow feature map after feature extraction and the morphological enhancement feature map element by element to obtain an added feature map; Adaptively calibrating the channel weights of the added feature map through an efficient channel attention module to obtain a calibrated feature map; The calibrated feature map and the added feature map are multiplied element by element to obtain a deep feature map with enhanced morphological features, wherein the first shallow feature map, the second shallow feature map and the deep feature map are all used to detect surface defects of steel.
2. The method according to claim 1, characterized in that The method of obtaining the original feature map, the small-scale defect feature map, the medium-scale defect feature map, and the large-scale defect feature map corresponding to the first feature map after dimensionality increase by using multiple sets of filters in parallel includes: The first filter is used to retain the global contextual relationship of the original feature information of the first feature map after dimensionality increase, thereby generating an original feature map; Performing a 3×3 convolution operation on the first feature map after dimension increase through a second filter to generate a small-scale defect feature map; Through the third filter, the first feature map after dimension increase is added to the small-scale defect feature map element by element, and then a 3×3 convolution operation is performed to generate a medium-scale defect feature map; Through the fourth filter, the first feature map after dimensionality increase and the medium-scale defect feature map are added element by element and then a 3×3 convolution operation is performed to generate a large-scale defect feature map.
3. The method according to claim 1, characterized in that The ratio values are, from large to small, a first compression ratio value, a second compression ratio value, and a third compression ratio value.
4. The method according to claim 1, wherein Before respectively inputting the channel squeezing excitation module with the compression ratios set to the second compression ratio value and the third compression ratio value, the method further includes: The two first sub-feature maps are respectively input into two 1×1 convolutional layers to enhance the representation capabilities of the two first sub-feature maps for the underlying texture information of small-scale defect features.
5. The method according to claim 1, wherein The step of inputting the second shallow feature map after feature extraction into multiple feature extraction modules and outputting a morphologically enhanced feature map comprises: The scratch edge features and crack direction features of the second shallow feature map after feature extraction are extracted through a 7×7 depth-separable convolution layer; A multi-layer perceptron module constructed through linear transformation layers and GELU functions enhances the global contextual information and spatial relationship of scratch edge features and crack direction features, generating an enhanced second shallow feature map. The enhanced second shallow feature map and the feature-extracted second shallow feature map are added element by element to generate a morphologically enhanced feature map.
6. The method according to claim 1, characterized in that The method further comprises, Enhance the large and small scale defect features of the deep feature map through the large and small scale feature enhancement module; The morphological features of the deep feature map are further enhanced through the morphological enhancement module.
7. A steel surface defect detection system, characterized in that: include: processor and memory; The memory is coupled to the processor, and the memory is used to store computer program code. The processor calls the computer program code to enable the steel surface defect detection system to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Steel surface defect detection method and system, storage medium and product
CN118781107A
Steel surface defect detection method based on improved model
CN119048503A
Cited By
Steel surface defect detection method and device based on improved YOLOv11
CN122367890A