Steel surface defect detection method and related device
Through the improved object detection model, combined with focus modulation, multi-scale convolution and robust feature downsampling modules, the Yolov8n model has solved the problem of high computational complexity and low micro defect recognition accuracy in steel surface defect detection, achieving more efficient defect detection.
Patent Information
- Application Number
- CN202510310367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing Yolov8n model has problems with high computational complexity and low recognition accuracy in steel surface defect detection, especially in the detection of multi-detail objects and small objects, which are not effective in feature extraction.
The improved object detection model is adopted, including the focus modulation module, the large core multi-scale convolution module and the robust feature downsampling module, and the fine feature representation is generated through aggregation, and parallel calculations are performed with convolution kernels of different sizes, and key feature information is retained during the downsampling process.
This reduces the computational complexity of the model, improves the recognition accuracy of tiny defects, and enhances the model's ability to detect complex backgrounds and multi-detail objects.
Smart Images

Figure CN120387979A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of cell image recognition, and in particular, to a method and related device for detecting steel surface defects. Background Art
[0002] With the continuous development of manufacturing and processing technology levels, various metal products (such as stainless steel) refined from iron ore have characteristics such as corrosion resistance, high hardness, high toughness, and high compressive strength, and have become indispensable items in people's daily lives. Therefore, how to quickly and batch-screen qualified and high-quality metal products has naturally become an important topic that cannot be ignored in industrial manufacturing. However, due to the influence of various factors such as mechanical equipment, processing methods, and material quality during the processing of steel, various defects are likely to occur on the surface of metal products during the actual production, processing, storage, and transportation in industry, such as: scratches, burns, dense_rust, and pitting, etc. These defects not only reduce the appearance quality of the products, but may also cause damage to machine processing equipment, and even reduce the safety performance and service life of derivative products.
[0003] Currently, the most used model in Yolov8 is Yolov8n, which is the smallest version of Yolov8. Due to its lightweight characteristics, it is easier to load onto embedded devices and is currently commonly used in the detection of industrial target detection fields.
[0004] However, in the face of surface defect detection, the Yolov8n model still has many significant problems: in the SPPF module, there is a situation where pooling layers are stacked in large numbers. Although the pooling layer has a fast calculation speed, when facing objects with many details, the feature extraction effect of the pooling layer will still be affected to a certain extent. And when facing small target detection, the ordinary downsampling convolution module is not very effective in extracting features of small targets and is prone to feature loss.
[0005] Secondly, in deep learning, in order to improve the model's ability to learn features, the number of network layers is often stacked, and the size and number of convolutional kernels are increased. Although these practices can slightly improve the test accuracy of the model to a certain extent, in the face of exponentially growing parameter quantities, the tiny improvement in accuracy will not be able to offset this huge cost, and current models mostly have these problems. Summary of the Invention
[0006] The embodiments of the present invention provide a method and related device for detecting steel surface defects, which can reduce the computational complexity of the model and improve the recognition accuracy of the model for tiny defects.
[0007] First aspect, an embodiment of the present invention provides a method for detecting surface defects of steel, including:
[0008] Obtain a surface image of the steel to be detected;
[0009] Input the surface image of the steel into a trained object detection model for defect detection to obtain a surface defect detection result of the steel. Among them, the object detection model is improved based on the Yolov8n model. The object detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module dynamically adjusts the importance of the multiple feature maps using spatially aware weights and finally generates a single output feature map; the large kernel multi-scale convolution module is used for parallel computing by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process.
[0010] In some embodiments, the focus modulation module uses a group of depthwise separable convolutions to encode the short-range to long-range visual context, encodes spatial contexts of different granularities as summary tokens, and then selectively fuses them into the query according to the query content.
[0011] In some embodiments, the hierarchical contextualization module projects the input feature map into a new feature space, and the gated aggregation module is also used to aggregate linear convolution layers at different levels.
[0012] In some embodiments, the hierarchical contextualization module projects the input feature map into a new feature space, including:
[0013] The hierarchical contextualization module linearly projects the input feature map to generate an initial feature map;
[0014] The hierarchical contextualization module gradually constructs a hierarchical feature representation by stacking multiple layers of depthwise separable convolutions.
[0015] In some embodiments, the robust feature downsampling module includes a shallow robust feature downsampling module and a deep robust feature downsampling module.
[0016] In some embodiments, in the shallow robust feature downsampling module, the input feature map X is processed by a 7x7 convolution with a stride of 1, expanding the number of channels from 3 to 1 / 4C; then, the feature map X is copied into y1 and y2. Among them, y1 retains the original feature information through slicing downsampling DCut and generates a matrix with halved size at the same time; y2 respectively passes through 3x3 downsampling convolution blocks with a stride of 1 and a stride of 2, and the 3x3 downsampling convolution blocks double the number of output channels; then, y1 and y2 are merged into a total feature map with the number of channels C, and the number of channels is reduced to 1 / 2C through a 1x1 convolution with a stride of 1; finally, after performing a convolution operation on the previous output, it is sent to the second half of the shallow robust feature downsampling module, where the structure of the second half of the shallow robust feature downsampling module is the same as that of the deep robust feature downsampling module.
[0017] In some embodiments, in the deep robust feature downsampling module, the input feature map Y is downsampled by DConv, DCut, and DMax to obtain output feature maps X1, X2, and X3 respectively; subsequently, X1, X2, and X3 are concatenated together to form a feature map with 6C channels; finally, a 1x1 convolutional layer with a stride of 1 is used to reduce the number of channels to 2C.
[0018] In a second aspect, an embodiment of the present invention further provides a steel surface defect detection device, the device includes:
[0019] An acquisition module, configured to acquire a steel surface image to be detected;
[0020] A detection module, configured to input the steel surface image into a trained target detection model for defect detection to obtain a steel surface defect detection result. Among them, the target detection model is improved based on the Yolov8n model. The target detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module dynamically adjusts the importance of the multiple feature maps using spatially aware weights and finally generates a single output feature map; the large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process.
[0021] In a third aspect, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steel surface defect detection method as described in the first aspect when executing the computer program.
[0022] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for executing the steel surface defect detection method as described in the first aspect.
[0023] According to the steel surface defect detection method and related device provided by the embodiments of the present invention, the steel surface defect detection method includes: obtaining a steel surface image to be detected; inputting the steel surface image into a trained object detection model for defect detection to obtain a steel surface defect detection result, wherein the object detection model is improved based on the Yolov8n model, and the object detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module dynamically adjusts the importance of multiple feature maps using spatially aware weights and finally generates a single output feature map; the large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process. First, the fusion focus modulation module solves to some extent the problem that multiple stacked max pooling layers in the SPPF module will reduce the network's processing ability for perceiving complex backgrounds or multi-detail objects, and at the same time can also increase the feature extraction ability of the original model. Second, the large kernel multi-scale convolution module with large kernel convolution reduces the computational complexity of the model and improves the detection accuracy of the model to some extent. Finally, by integrating the robust feature downsampling module, the YOLOv8n model can better handle defect targets with small objects and few features, further improving the recognition accuracy of the network model for tiny defects. Based on this, the embodiments of the present invention can reduce the computational complexity of the model and improve the recognition accuracy of the model for tiny defects. Description of the Drawings
[0024] Figure 1 is a flowchart of the steel surface defect detection method provided by an embodiment of the present invention;
[0025] Figure 2A is a schematic structural diagram of the SPPF module used in the YOLOv8 backbone network;
[0026] Figure 2BIt is a schematic structural diagram of the Focal Modulation module provided by an embodiment of the present invention;
[0027] Figure 3A It is the C2f module used in the YOLOv8 backbone network;
[0028] Figure 3B It is a schematic structural diagram of the Multi-Scale Convolution with Large Kernel (MSC) module provided by an embodiment of the present invention;
[0029] Figure 4A It is a schematic structural diagram of the Shallow Robust Feature Downsampling (SRFD) module provided by an embodiment of the present invention;
[0030] Figure 4B It is a schematic structural diagram of the Deep Robust Feature Downsampling (DRFD) module provided by an embodiment of the present invention;
[0031] Figure 5 It is a schematic structural diagram of the steel surface defect detection device provided by an embodiment of the present invention;
[0032] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0034] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. Terms such as "first", "second", etc. in the specification, claims and the following drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0035] In the embodiments of the present invention, words such as "furthermore", "exemplarily" or "optionally" are used to represent examples, illustrations or explanations, and should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Using words such as "furthermore", "exemplarily" or "optionally" is intended to present related concepts in a specific manner.
[0036] In order to more conveniently describe the working principle of the embodiments of the present invention hereinafter, an introduction to the related technical scenarios will be given first.
[0037] With the continuous development of manufacturing and processing technology, various metal products refined from iron ore (such as stainless steel) have characteristics such as corrosion resistance, high hardness, high toughness, and high compressive strength, and have become indispensable items in people's daily lives. Therefore, how to quickly and batch-screen qualified and high-quality metal products has naturally become an important topic that cannot be ignored in industrial manufacturing. However, due to the influence of various factors such as mechanical equipment, processing methods, and material quality during the processing of steel, various defects are likely to occur on the surface of metal products during the actual production, processing, storage, and transportation in industry, such as: scratches, burns, dense rust, and pitting, etc. These defects not only reduce the appearance quality of the products, but may also cause damage to machine processing equipment, and even reduce the safety performance and service life of derivative products.
[0038] Currently, the most widely used model in Yolov8 is Yolov8n, which is the smallest version of Yolov8. Due to its lightweight characteristics, it is easier to load onto embedded devices and is currently commonly used in the detection of industrial target detection fields.
[0039] However, in the face of surface defect detection, there are still many significant problems with the Yolov8n model: in the SPPF module, there is a large stack of pooling layers. Although the pooling layer has a fast calculation speed, when facing objects with many details, the feature extraction effect of the pooling layer will still be affected to a certain extent. And when facing small target detection, the ordinary downsampling convolution module is not very effective in extracting features of small targets and is prone to feature loss.
[0040] Secondly, in deep learning, in order to improve the model's ability to learn features, people often choose to stack the number of network layers, increase the size and number of convolution kernels, etc. Although these practices can slightly improve the test accuracy of the model to a certain extent, in the face of the exponentially increasing number of parameters, the tiny improvement in accuracy will not be able to offset this huge cost, and current models mostly have these problems.
[0041] Based on this, the present invention provides a method and related device for detecting steel surface defects. Among them, the method for detecting steel surface defects includes: obtaining a steel surface image to be detected; inputting the steel surface image into a trained object detection model for defect detection to obtain a steel surface defect detection result. The object detection model is improved based on the Yolov8n model and includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures contextual information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module dynamically adjusts the importance of multiple feature maps using spatially aware weights and finally generates a single output feature map; the large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process. First, the fused focus modulation module solves to some extent the problem that multiple stacked max pooling layers in the SPPF module reduce the network's processing ability for perceiving complex backgrounds or multi-detail objects, and at the same time can also increase the feature extraction ability of the original model. Second, the large kernel multi-scale convolution module with large kernel convolution reduces the computational complexity of the model and improves the detection accuracy of the model to some extent. Finally, by fusing the robust feature downsampling module, the YOLOv8n model can better handle defect targets with small objects and few features, further improving the recognition accuracy of the network model for tiny defects. Based on this, the embodiments of the present invention can reduce the computational complexity of the model and improve the recognition accuracy of the model for tiny defects.
[0042] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.
[0043] As Figure 1 shown, Figure 1 is a flowchart of a method for detecting steel surface defects provided by an embodiment of the present invention. The method for detecting steel surface defects may include but is not limited to steps S101 to S102.
[0044] Step S101: Obtain a steel surface image to be detected;
[0045] Step S102: Input the steel surface image into the trained object detection model for defect detection to obtain the steel surface defect detection result. The object detection model is improved based on the Yolov8n model. The object detection model includes a focal modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focal modulation module is used to generate a fine feature representation through aggregation. The focal modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps. The gated aggregation module dynamically adjusts the importance of multiple feature maps using spatially aware weights and finally generates a single output feature map. The large kernel multi-scale convolution module is used to perform parallel computing by combining convolution kernels of different sizes. The robust feature downsampling module is used to retain key feature information during the downsampling process.
[0046] It can be understood that as an improved version of SPP, SPPF can process the features of multi-scale objects, ensuring that the model can obtain rich information when facing targets of various sizes. This characteristic enables the SPPF module to greatly improve the accuracy of the model when small or large targets frequently appear, such as Figure 2A shown in the SPPF module used in the YOLOv8 backbone network, where MaxPool is a stacked max pooling layer, Conv is a convolutional layer, and Concat is a concatenation operation. However, although the multiple stacked max pooling operations in SPPF speed up the processing speed of the module, they reduce the network's processing ability to perceive complex backgrounds and face multi-detail objects. The self-attention mechanism can enable the global information of the image to interact and enhance the network's ability to extract features. However, the self-attention calculation complexity is relatively high and it cannot run well on embedded devices. At this time, the focal modulation module has become a new method that can replace the self-attention mechanism. The focal modulation module is as shown in Figure 2B where Linear is a linear layer, Aggregate is an aggregation layer, Interact is an interaction layer, Hlerarchical Contextualization is a hierarchical contextualization module, and Gated Aggregation is a gated aggregation module.
[0047] The focal modulation module uses a set of depthwise separable convolutions to encode the short-range to long-range visual context, encoding spatial contexts of different granularities as summarized tokens, and then selectively fusing them into the query according to the query content. As shown in Equation (1),
[0048] yi = q(x i ) ⊙ M(x i , X) (1)
[0049] Among them, M represents the aggregation process, q(·) is a query mapping function, and ⊙ is the element-wise multiplication operator.
[0050] Compared with the self-attention SA, this method of aggregating first and then interacting significantly improves the operation efficiency of the Focal Net. The core idea of the Focal Modulation Network is to generate a fine-grained feature representation through aggregation. Its aggregation process includes two key steps: Hierarchical Contextualization and Gated Aggregation.
[0051] Hierarchical Contextualization gradually captures context information at different levels through depth convolution operations to generate multiple feature maps. Subsequently, Gated Aggregation dynamically adjusts the importance of these feature maps using spatially aware weights, and finally generates a single output feature map. This method can not only capture fine-grained local information but also effectively integrate large-scale spatial context. This way of using hierarchical context-aware input images can project the input feature map into a new feature space. Specifically, the input feature map first undergoes a linear projection to generate the initial feature map Z 0 , as shown in Equation (2). Subsequently, by stacking multiple layers of depthwise separable convolutions (dw convolutions), a hierarchical feature representation is gradually constructed. Compared with the stacked max pooling layers in SPPF, the l depthwise separable convolutions used in the Focal Modulation Network are learnable, and its structure has perceptual ability, enabling the Focal Modulation Network to obtain a larger receptive field to capture global context information. Among them, at the l-th layer, the output feature map Z l is calculated as shown in Equation (3).
[0052]
[0053] In addition, a gated aggregation mechanism is also used to aggregate linear convolutional layers at different levels, as shown in the following Equation (4).
[0054]
[0055] Among them, is the weight of the l-th layer. Finally, an additional computational layer h(·) is also used to obtain the final output modulator M = h(Z out ).
[0056] It can be understood that the embodiment of the present invention uses a large kernel multi-scale convolution (MSC) module to replace C2f for feature extraction. Figure 3AIt is the C2f module used in the YOLOv8 backbone network. Among them, BottleNeck is the bottleneck layer, Split is the splitting operation, Conv is the convolutional layer, and Concat is the merging operation. Figure 3B It is the large kernel multi-scale convolution (MSC) module used for replacement.
[0057] In deep learning, in order to improve the model's ability to learn features, one often chooses to stack the number of network layers and increase the size and number of convolutional kernels, etc. Although these practices can slightly improve the model's test accuracy to a certain extent, when faced with exponentially growing parameter quantities, the tiny improvement in accuracy will not be able to offset this huge cost. Based on this, the present invention turns to flattened multi-scale convolution operations. By combining convolutional kernels of different sizes for parallel computing, it is able to extract rich features under the condition of reducing parameters.
[0058] For this reason, the present invention further improves the MSC module and adds a large kernel convolution of 7x7. This module will reduce the computational complexity of the model to a certain extent. At the same time, experiments also show that this module can increase the network's feature extraction ability and can, to a certain extent, reduce the accuracy loss problem caused by a large number of stacked BottleNecks in C2f. The reduction in the number of module parameters not only helps to reduce the complexity of the model but also can further reduce overfitting on small datasets and improve the performance of the model on small datasets, making the model more robust in practical applications.
[0059] It can be understood that the embodiment of the present invention uses a robust feature downsampling (RFD) module to replace the convolutional module.
[0060] When the network receives the original image, the original image usually contains a large amount of pixel-level noise and redundant information, and this noise will affect feature extraction. Therefore, retaining image features as much as possible and reducing the impact of noise on the image are things that the downsampling module in the backbone network needs to solve. Although the traditional 3x3 convolutional module in YOLOv8 has a relatively fast calculation speed, however, this downsampling method is not very effective for the feature description of small targets. This operation is likely to cause the loss of features of small targets and increase the possibility of interference from the original noise on the feature map, thus affecting the final performance.
[0061] Therefore, the present invention introduces a robust feature downsampling (RFD) module, which is divided into two forms: namely, shallow robust feature downsampling (SRFD) and deep robust feature downsampling (DRFD), used to replace the ordinary convolutional module of the YOLOv8 network. Among them, shallow RFD (SRFD) replaces the first convolutional module in the network, while DRFD is used to replace other convolutional modules.
[0062] Figure 4A Shown is the Shallow Robust Feature Downsampling (SRFD) module. In the SRFD module, the input feature map X is processed by a 7x7 convolution with a stride of 1, expanding the number of channels from 3 to 1 / 4C. This convolution operation can fuse local features and filter redundant information, and enhance important features in the original image. Then, the feature map x is copied into y1 and y2, where y1 retains the original feature information through slicing downsampling (DCut), while generating a matrix with half the size. y2 is respectively passed through 3x3 downsampling convolution blocks with a stride of 1 and a stride of 2. This convolution block doubles the number of output channels, thus integrating local features and further reducing the size of the feature map. Then, y1 and y2 are merged into a total feature map with the number of channels C, and passed through a 1x1 convolution with a stride of 1 to reduce the number of channels to 1 / 2C. Finally, after convolution on the previous output, it is sent to the second half of SRFD. The module structure of the second half is the same as that of the DRFD module. All in all, the SRFD module tries to avoid the operation of max pooling during downsampling, ensuring the feature expression ability at different scales, and at the same time reducing the amount of noise in the original image while retaining key features.
[0063] Figure 4B Shown is the Deep Robust Feature Downsampling (DRFD) module. In DRFD, the input feature map Y is downsampled through DConv, DCut, and DMax to obtain output feature maps X1, X2, and X3 respectively. Subsequently, X1, X2, and X3 are concatenated together to form a feature map with 6C channels. Finally, a 1x1 convolution layer with a stride of 1 is used to reduce the number of channels to 2C. This design ensures that during downsampling, the loss of semantic information is reduced.
[0064] It can be understood that the present invention technically combines the YOLOv8 architecture and the RFD module, enabling the model to more accurately identify small defects. And it innovatively proposes the Large Kernel Multi-Scale Convolution (MSC) module, which can reduce the computational complexity of the model to a certain extent. The Focal Modulation module is also introduced to make the model detection accuracy higher.
[0065] The steel surface defect detection method according to the embodiment of the present invention. First, the fusion of Focal Modulation solves to a certain extent the problem that multiple stacked max-pooling layers in the SPPF module will reduce the network's processing ability for perceiving complex backgrounds or multi-detail objects, and at the same time can also increase the feature extraction ability of the original model. Secondly, the MSC module with a 7x7 large kernel convolution reduces the computational complexity of the model to a certain extent and improves the detection accuracy of the model. Finally, by fusing the RFD module, the YOLOv8n model can better handle defect targets with smaller objects and fewer features, further improving the recognition accuracy of the network model for micro-defects. Based on this, the embodiment of the present invention can reduce the computational complexity of the model and improve the recognition accuracy of the model for micro-defects.
[0066] In addition, as Figure 5 shown, an embodiment of the present invention also discloses a steel surface defect detection device, which includes:
[0067] An acquisition module 110, configured to acquire a steel surface image to be detected;
[0068] A detection module 120, configured to input the steel surface image into a trained object detection model for defect detection to obtain a steel surface defect detection result, where the object detection model is improved based on the Yolov8n model, and the object detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module dynamically adjusts the importance of multiple feature maps using spatially aware weights, and finally generates a single output feature map; the large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process.
[0069] The steel surface defect detection device according to the embodiment of the present invention is used to execute the steel surface defect detection method in the above embodiment, and its specific processing process is the same as that of the steel surface defect detection method in the above embodiment, and will not be elaborated here one by one.
[0070] In addition, as Figure 6 shown, an embodiment of the present invention also discloses an electronic device, including: at least one processor 210; at least one memory 220, configured to store at least one program; when at least one program is executed by at least one processor 210, it implements the steel surface defect detection method in any of the previous embodiments.
[0071] In addition, an embodiment of the present invention also discloses a computer-readable storage medium storing computer-executable instructions for executing the steel surface defect detection method in any of the previous embodiments.
[0072] The system architecture and application scenarios described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. As known to those skilled in the art, with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0073] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0074] In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0075] As used in this specification, the terms "component", "module", "system", etc. are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside within a process or execution thread, and a component can be located on one computer or distributed between two or more computers. In addition, these components can execute from various computer-readable media having various data structures stored thereon. A component can communicate, for example, by way of signals with one or more data packets (e.g., data from two components interacting with each other from a local system, a distributed system, or a network, such as the Internet interacting with other systems via signals) through local or remote processes.
Claims
1. A method for detecting surface defects of steel, comprising: Obtaining a surface image of the steel to be detected; Inputting the surface image of the steel into a trained object detection model for defect detection to obtain a surface defect detection result of the steel. Among them, the object detection model is improved based on the Yolov8n model. The object detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps; the gated aggregation module uses spatially aware weights to dynamically adjust the importance of the multiple feature maps and finally generates a single output feature map; the large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes; the robust feature downsampling module is used to retain key feature information during the downsampling process.
2. The method according to claim 1, wherein The focus modulation module uses a group of depthwise separable convolutions to encode the short-range to long-range visual context, encode the spatial context of different granularities into summary tokens, and then selectively fuse them into the query according to the query content.
3. The method according to claim 1, wherein The hierarchical contextualization module projects the input feature map into a new feature space, and the gated aggregation module is also used to aggregate linear convolutional layers of different levels.
4. The method according to claim 3, wherein The hierarchical contextualization module projects the input feature map into a new feature space, including: The hierarchical contextualization module linearly projects the input feature map to generate an initial feature map; The hierarchical contextualization module gradually constructs a hierarchical feature representation by stacking multiple layers of depthwise separable convolutions.
5. The method according to claim 1, characterized in that The robust feature downsampling module includes a shallow robust feature downsampling module and a deep robust feature downsampling module.
6. The method according to claim 5, wherein In the shallow robust feature downsampling module, the input feature map X is processed by a 7x7 convolution with a stride of 1 to expand the number of channels from 3 to 1 / 4C; then, the feature map X is copied into y1 and y2. Among them, y1 retains the original feature information through slicing downsampling DCut and generates a matrix with a size halved at the same time; y2 passes through 3x3 downsampling convolution blocks with a stride of 1 and a stride of 2 respectively, and the 3x3 downsampling convolution blocks double the number of output channels; then, y1 and y2 are merged into a total feature map with the number of channels C, and the number of channels is reduced to 1 / 2C through a 1x1 convolution with a stride of 1; finally, after performing a convolution operation on the previous output, it is sent to the second half of the shallow robust feature downsampling module, where the structure of the second half of the shallow robust feature downsampling module is the same as that of the deep robust feature downsampling module.
7. The method according to claim 5, wherein In the deep robust feature downsampling module, the input feature map Y is downsampled by DConv, DCut, and DMax to obtain output feature maps X1, X2, and X3 respectively. Subsequently, X1, X2, and X3 are concatenated together to form a feature map with 6C channels. Finally, a 1x1 convolutional layer with a stride of 1 is used to reduce the number of channels to 2C.
8. A steel surface defect detection device, characterized in that, The device includes: An acquisition module for acquiring the steel surface image to be detected; A detection module for inputting the steel surface image into a trained object detection model for defect detection to obtain the steel surface defect detection result. The object detection model is improved based on the Yolov8n model. The object detection model includes a focus modulation module, a large kernel multi-scale convolution module, and a robust feature downsampling module. The focus modulation module is used to generate a fine feature representation through aggregation. The focus modulation module includes a hierarchical contextualization module and a gated aggregation module. The hierarchical contextualization module gradually captures context information at different levels through depth convolution operations to generate multiple feature maps. The gated aggregation module dynamically adjusts the importance of the multiple feature maps using spatially aware weights and finally generates a single output feature map. The large kernel multi-scale convolution module is used to perform parallel calculations by combining convolution kernels of different sizes. The robust feature downsampling module is used to retain key feature information during the downsampling process.
9. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steel surface defect detection method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing computer-executable instructions for executing the steel surface defect detection method according to any one of claims 1 to 7.