Steel surface defect detection method and system

By improving the YOLOv11 network model, including adding detection heads, replacing the C3k2 module and using a lightweight detection head with parameters, the real-time and efficiency problems of steel surface defect detection in the prior art are solved, and more efficient and accurate detection effects are achieved.

CN119941724AActive Publication Date: 2025-05-06NANCHANG GENGXIANG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510422293.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing machine learning methods are difficult to meet the real-time and efficiency requirements of industrial scenarios in steel surface defect detection.

Method used

Using the improved YOLOv11 network model, the recognition ability and computing efficiency of the model are improved by adding a detection head on the detection layer, replacing the C3k2 module with a C3k2-MSM module, and replacing the detection head with a parameter-focused and lightweight detection head.

Benefits of technology

It improves the accuracy and efficiency of steel surface defect detection, can maintain high detection accuracy while reducing computing resource requirements, and is suitable for real-time detection in industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941724A_ABST
    Figure CN119941724A_ABST
Patent Text Reader

Abstract

The invention discloses a steel surface defect detection method and system, and belongs to the technical field of computer vision, and the method comprises the steps: obtaining a steel image data set, and carrying out the preprocessing, and obtaining a target data set; then, a YOLOv11 network model is constructed, a detection head is added to a detection layer of the YOLOv11 network model, and meanwhile, a C3k2-MSM module is constructed to replace a C3k2 module in the backbone network; and then, all detection heads are replaced by heavy parameter lightweight detection heads, and an improved YOLOv11 network model is obtained. And performing model training by using part of data of the target data set, and testing the other part of data. And finally, inputting a to-be-detected image into the tested model, and outputting a steel surface defect detection result. By optimizing the network structure and the detection head, the detection precision and efficiency are improved, and automatic efficient real-time detection of steel surface defects is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer vision technology, and specifically relates to a steel surface defect detection method and system. Background Art

[0002] As the core basic material of the modern industrial system, steel occupies an important position in the field of advanced equipment manufacturing. Its application scope covers key scenarios such as aerospace structural parts, rail transportation equipment components, automobile body and chassis systems. However, during the rolling process, heat treatment and service, various types of defects are prone to occur on the surface of steel. For example, thermal stress-induced cracks, mechanical damage scratches, oxidation pitting, fatigue pitting, etc. These surface defects will significantly affect the service performance of the material, causing systemic risks such as decreased tensile strength, attenuated ductility, and reduced corrosion resistance. What's more serious is that undetected critical size defects may cause catastrophic failures, such as pressure vessel bursts, drive shaft fractures and other Level II industrial accidents. Therefore, it is of great significance to detect defects on the surface of steel.

[0003] The existing machine learning methods for steel surface defect detection have high computing resource consumption due to their complex network structure and large number of parameters, making it difficult to meet the real-time and efficiency of defect detection in industrial scenarios. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method and system for detecting steel surface defects, which can solve the technical problem that existing machine learning methods are difficult to meet the real-time and efficiency requirements of steel surface defect detection in industrial scenarios.

[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a method for detecting surface defects of steel, the method comprising: Acquire a steel image dataset, and preprocess the steel image dataset to obtain a target dataset; Building a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; Construct a C3k2-MSM module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module; Replacing all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtaining an improved YOLOv11 network model; The improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; The image to be detected is input into the tested improved YOLOv11 network model, and the defect detection result is output.

[0006] As an optional implementation of the first aspect of the present application, the steel image dataset is acquired, and the steel image dataset is preprocessed to obtain a target dataset; specifically: Acquire a large number of steel images with defects on the surface to construct a steel image dataset according to each of the steel images; Perform defect marking on each of the steel images in the steel image data set, and mark the defect position of each of the steel images; The steel image data set after the defect annotation is subjected to data expansion and enhancement processing to obtain the target data set.

[0007] As an optional implementation of the first aspect of the present application, the process of data processing of the C3k2-MSM module is specifically as follows: The feature vector is input into the C3k2-MSM module, and the convolution layer in the C3k2-MSM module performs convolution processing on the feature vector to obtain an initial feature map; Performing feature extraction and transformation processing on the initial feature map according to the deep pooling layer in the C3k2-MSM module to obtain a deep fusion feature map; The initial feature map and the deep fusion feature map are residually connected according to the residual connection layer in the C3k2-MSM module to obtain a target feature map.

[0008] As an optional implementation of the first aspect of the present application, the deep pooling layer performs feature extraction and transformation processing on the initial feature map to obtain a deep fusion feature map; specifically: The deep pooling layer performs multi-scale pooling processing on the initial feature map to obtain multiple pooling feature maps of different scales; Performing convolution processing on each of the pooled feature maps, and then performing dynamic convolution processing on each of the pooled feature maps after the convolution processing, to obtain each dynamic convolution feature map corresponding to each of the pooled feature maps; Performing upsampling processing on each of the dynamic convolution feature maps, and performing residual connection between each of the dynamic convolution feature maps after the upsampling processing and the initial feature map to obtain a residual feature map corresponding to each of the dynamic convolution feature maps; Performing multi-scale feature fusion processing on each of the residual feature maps of different scales to obtain a fused feature map; The fused feature map is activated and convolved to obtain the deep fused feature map.

[0009] As an optional implementation of the first aspect of the present application, the process of data processing by the heavy-parameter lightweight detection head is specifically as follows: Inputting the input feature map into the heavy parameter lightweight detection head, and the diversified branch module in the heavy parameter lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; The branch fusion feature map is processed according to the conditional group normalization convolution module to obtain an output feature map.

[0010] As an optional implementation of the first aspect of the present application, the diversified branch module performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically: The standard convolution layer, 1×1 convolution layer, average pooling layer and double convolution layer in the diversified branch module process the input feature map respectively to obtain convolution features, 1×1 features, pooling features and double convolution features respectively; According to the fusion layer in the diversified branch module, the input feature map, the convolution feature, the 1×1 feature, the pooling feature and the double convolution feature are subjected to feature fusion to obtain a fused feature map; Performing activation processing on the fused feature map according to the activation layer in the diversified branch module to obtain the branch fused feature map; The conditional group normalized convolution module processes the branch fusion feature map to obtain an output feature map, specifically: Performing two-dimensional convolution processing on the branch fusion feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain a branch fusion convolution feature map; Normalizing the branch fusion convolution feature map according to the group normalization layer in the conditional group normalization convolution module; The normalized branch fusion convolution feature map is activated according to the activation layer in the conditional group normalization convolution module to obtain the output feature map.

[0011] As an optional implementation manner of the first aspect of the present application, the improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; specifically: Dividing the target data set into a training set, a validation set, and a test set according to a preset ratio; Constructing a loss function, and training the improved YOLOv11 network model according to the training set and the loss function; Verifying the trained improved YOLOv11 network model according to the verification set, and optimizing parameters of the improved YOLOv11 network model according to the difference between the verification output of the improved YOLOv11 network model and the true label; The improved YOLOv11 network model after parameter optimization is tested according to the test set to obtain test results.

[0012] In a second aspect, an embodiment of the present application provides a steel surface defect detection system, the system comprising: Acquisition module: acquires a steel image data set, preprocesses the steel image data set, and obtains a target data set; The first improvement module: constructing a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; The second improvement module: constructing a C3k2-MSM module, replacing the C3k2 module in the YOLOv11 network model backbone network with the C3k2-MSM module; The third improvement module: replace all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtain an improved YOLOv11 network model; Training module: training the improved YOLOv11 network model according to part of the data in the target data set, and testing the improved YOLOv11 network model according to another part of the data; Prediction module: inputs the image to be detected into the improved YOLOv11 network model after the test, and outputs the defect detection result.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In the embodiments of the present application, compared with the prior art, the following technical effects are achieved: (1) By constructing a YOLOv11 network model and adding a detection head to its detection layer, and subsequently replacing all detection heads on the detection layer with heavy-parameter lightweight detection heads, the model's ability to recognize steel surface defects can be enhanced, thereby improving the accuracy of detection.

[0016] (2) The introduction of the C3k2-MSM module may further enhance the model’s ability to capture defect features by improving feature extraction capabilities, which helps to more accurately identify defects of various types and sizes.

[0017] (3) The use of a heavy-parameter lightweight detection head helps to reduce the computational complexity of the model, thereby improving the detection speed and achieving rapid response while maintaining high detection accuracy.

[0018] (4) The introduction of a heavy-parameter lightweight detection head helps to reduce the complexity of the model and the model's demand for computing resources, allowing this method to be deployed and applied on a wider range of hardware platforms.

[0019] (5) This method realizes automatic real-time detection of steel surface defects by constructing and improving the YOLOv11 network model, reducing manual intervention and improving the consistency and reliability of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of a steel surface defect detection method provided by some embodiments of the present application; Figure 2 It is a structural diagram of an improved YOLOv11 network model of a steel surface defect detection method provided by some embodiments of the present application; Figure 3 It is a C3k2-MSM module structure diagram of a steel surface defect detection method provided by some embodiments of the present application; Figure 4 It is a deep pooling layer structure diagram of a steel surface defect detection method provided by some embodiments of the present application; Figure 5 It is a structural diagram of a heavy-parameter lightweight detection head of a steel surface defect detection method provided by some embodiments of the present application; Figure 6 It is a diversified branch module structure diagram of a steel surface defect detection method provided by some embodiments of the present application; Figure 7 This is a structural diagram of a conditional group normalized convolution module of a steel surface defect detection method provided in some embodiments of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0023] In the following, in conjunction with the accompanying drawings, a steel surface defect detection method and system provided by the embodiment of the present application are described in detail through specific embodiments and their application scenarios.

[0024] Example A method for detecting surface defects of steel comprises the following steps: S100: Acquire a steel image dataset, preprocess the steel image dataset, and obtain a target dataset; It should be noted that S100 is specifically: S110: Acquire a large number of steel images with surface defects to construct a steel image dataset according to each steel image; S120: marking defects of each steel image in the steel image data set, marking the defect position of each steel image; S130: Performing data expansion and enhancement processing on the steel image dataset after defect annotation to obtain a target dataset.

[0025] Furthermore, the steel image dataset in S110 uses a self-collected dataset and an open dataset on the Internet. The self-collected dataset uses a two-dimensional camera to photograph defective steel to obtain images of steel surface defects. Specifically, a Nikon Z8 micro-single camera is used to photograph 300 steel defect images at the steel production and manufacturing site to obtain the self-collected dataset; the open dataset on the Internet comes from the NEU-DET steel defect dataset publicly available on the Internet of Northeastern University; S120 first imports the self-collected dataset into the annotation tool, which is X-AnyLabeling; the defects of the 300 photographed images are marked in yolo format, and the annotation file contains information such as the category number of each defect target; in S130, the image data and labels are expanded to 3,000 using data enhancement methods such as random contrast enhancement, noise, flipping, and scaling to simulate the images recognized by the camera under various extreme conditions, so as to improve the generalization ability of the training model; the expanded self-collected dataset and the open dataset on the Internet are merged to obtain the target dataset.

[0026] S200: Build a YOLOv11 network model and add a detection head to the detection layer of the YOLOv11 network model; It should be noted that by adding a detection head, the model has four detection heads, which can extract four feature maps with sizes of 120×120, 64×64, 32×32, and 16×16 from the steel image with defects on the surface, which are used to detect defect targets of four different sizes: tiny, small, medium, and large on the steel surface, respectively.

[0027] S300: Build the C3k2-MSM module and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module; Furthermore, the data processing process of the C3k2-MSM (Multi-Scale Mixer) module is as follows: S310: The feature vector is input into the C3k2-MSM module, and the convolution layer in the C3k2-MSM module performs convolution processing on the feature vector to obtain an initial feature map; S320: extracting and transforming the initial feature map according to the deep pooling layer in the C3k2-MSM module to obtain a deep fusion feature map; S330: Perform residual connection on the initial feature map and the deep fusion feature map according to the residual connection layer in the C3k2-MSM module to obtain a target feature map.

[0028] It should be noted that the deep pooling layer in S320 performs feature extraction and transformation processing on the initial feature map to obtain a deep fusion feature map; specifically: S321: The deep pooling layer performs multi-scale pooling processing on the initial feature map to obtain multiple pooling feature maps of different scales; S322: performing convolution processing on each pooling feature map, and then performing dynamic convolution processing on each pooling feature map after the convolution processing, to obtain each dynamic convolution feature map corresponding to each pooling feature map; S323: performing upsampling processing on each dynamic convolution feature map, and performing residual connection between each dynamic convolution feature map after the upsampling processing and the initial feature map to obtain a residual feature map corresponding to each dynamic convolution feature map; S324: performing multi-scale feature fusion processing on each residual feature map of different scales to obtain a fused feature map; S325: Activate and convolve the fused feature map to obtain a deep fused feature map.

[0029] Specifically, the processing formula of the deep pooling layer is expressed as follows: , in, represents the initial feature map, represents average pooling processing, represents the convolution kernel, Indicates the stride, Indicates that the pooling kernel size is ( × ), represents the convolution process, express The corresponding convolutional feature map, express The corresponding dynamic convolution feature map, represents upsampling processing, express Dynamic convolution feature map after upsampling, express The corresponding residual feature map, represents the fused feature map, Indicates concatenation of all residual feature maps. Indicates activation processing, represents the fused feature map after activation processing, Represents the deep fusion feature map.

[0030] Furthermore, the deep pooling layer first performs multi-scale pooling on the initial feature map; multiple pooling feature maps of different scales are generated through pooling windows of different sizes (the pooling kernel sizes in this embodiment are 8×8, 4×4 and 2×2, and the shapes are (B, C, H / i, W / i)). These pooling feature maps can capture feature information at different scales and enhance the model's perception of diverse defects; convolution processing is performed on each pooling feature map. The convolution layer extracts features through a set of learnable convolution kernels to generate convolution feature maps. This step can further refine the important information in the pooling feature map and enhance the expressiveness of the features. After the convolution processing, each convolution feature map is subjected to dynamic convolution processing (the convolution kernel size is 3×3, the step size is 1, and the padding is 1). The idea of ​​dynamic convolution is to dynamically adjust the weights of the convolution kernel according to the different input feature maps to adapt to different features. This method can improve the flexibility and adaptability of the model, enabling it to better capture complex features. Upsampling is performed on each dynamic convolution feature map. The purpose of upsampling is to restore the spatial dimension of the feature map to the same size as the initial feature map. Deconvolution or interpolation methods are usually used for upsampling. Each upsampled dynamic convolution feature map is residually connected to the corresponding initial feature map. Through the addition operation, the residual connection can effectively fuse the information in the initial feature map with the information in the dynamic convolution feature map to generate a residual feature map corresponding to each dynamic convolution feature map. Multi-scale feature fusion processing is performed on each residual feature map of different scales. Through weighted averaging or splicing and other methods, the residual feature maps of different scales are fused together to form a comprehensive feature representation, which enhances the model's detection ability for various defects. Finally, the fused feature map is activated, usually using ReLU or other activation functions to introduce nonlinear features. This step can further enhance the expressiveness of the feature map and obtain the final deep fusion feature map; the deep pooling layer achieves efficient feature extraction and fusion through multi-scale pooling, dynamic convolution and residual connection; by accumulating feature maps of different scales, the model can simultaneously utilize global and local information, enhance the expressiveness of features, and improve the robustness of the model to targets of different scales; applying the GELU activation function to the accumulated feature maps enhances the expressiveness of the model.

[0031] S400: Replace all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtain an improved YOLOv11 network model; It should be noted that the data processing process of the Rep Shared Convolutional Detection Head is as follows: S410: Inputting the input feature map into the heavy parameter lightweight detection head, and the diversified branch module in the heavy parameter lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; S420: Process the branch fusion feature map according to the conditional group normalization convolution module to obtain an output feature map.

[0032] Furthermore, the heavy parameter lightweight detection head improves the decoupling head encoding so that the model can detect the function of rotating targets, solves the problem of discontinuous angle regression of traditional models, and solves the problem of poor module deployment compatibility through lightweight architecture. Compared with the detection head before improvement, it has significantly improved the feature representation and processing capabilities as well as the recognition speed and efficiency, and can be applied to more complex steel recognition scenarios. First, the input feature map is processed by branch feature extraction and fusion through the diversified branch module to obtain the branch fusion feature map; then the branch fusion feature map is processed by the conditional group normalization convolution module to obtain the output feature map. In terms of module architecture, the rotation angle parameter is introduced, and the integrated prediction of the center point, width, height and angle is realized by fusing the rotation box with the horizontal box. The angle encoding is adjusted, and the sigmoid activation function is used to constrain the angle prediction to the range of [-π / 4, 3π / 4]. Compared with the direct prediction of 0-π / 2, the boundary mutation problem is avoided. The anchor decoding is adjusted to convert the distance offset into a rotation box, and the geometric adaptation is realized by combining the anchor size and the feature map step size, making the model detection more flexible.

[0033] It should be noted that the diversified branch module in S410 performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically: S411: The standard convolution layer, 1×1 convolution layer, average pooling layer and double convolution layer in the diversified branch module process the input feature map respectively to obtain convolution features, 1×1 features, pooling features and double convolution features respectively; S412: performing feature fusion on the input feature map, convolution feature, 1×1 feature, pooling feature and double convolution feature according to the fusion layer in the diversified branch module to obtain a fused feature map; S413: Activate the fused feature map according to the activation layer in the diversified branch module to obtain a branch fused feature map.

[0034] Furthermore, the diversified branch module performs a standard convolution operation on the input feature map to extract convolution features; this process can capture local features by sliding the input feature map through the convolution kernel. The 1×1 convolution layer performs a 1×1 convolution operation on the input feature map to obtain a 1×1 feature. This convolution method is mainly used for channel compression and feature interaction of features, which can effectively reduce the amount of parameters while retaining important information. The average pooling layer performs average pooling on the input feature map to obtain pooling features. Average pooling can reduce the dimension of the feature map while retaining global information by calculating the average value of the local area of ​​the feature map. The double convolution layer performs double convolution on the input feature map to obtain double convolution features. Double convolution usually refers to two consecutive convolution operations, which can extract more complex features. The double convolution layer is composed of two 1×1 convolution layers connected in sequence. In the fusion layer in the diversified branch module, the input feature map, convolution features, 1×1 features, pooling features and double convolution features are fused. The fusion layer can use addition, concatenation or other fusion strategies to integrate features from different sources to generate a comprehensive fusion feature map. This feature fusion can effectively combine the advantages of different feature extraction methods to form a richer and more comprehensive feature representation. Finally, the fusion feature map is input to the activation layer for activation processing. The activation layer usually uses ReLU (Rectified Linear Unit) or other activation functions to introduce nonlinear features so that the model can better learn complex feature relationships. After activation processing, the final branch fusion feature map is obtained.

[0035] It should be noted that the conditional group normalized convolution module in S420 processes the branch fusion feature map to obtain an output feature map, which is specifically: S421: The conditional group normalized convolution module processes the branch fusion feature map to obtain an output feature map, specifically: S422: performing two-dimensional convolution processing on the branch fusion feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain a branch fusion convolution feature map; S423: normalizing the branch fusion convolution feature map according to the group normalization layer in the conditional group normalization convolution module; S424: Activate the normalized branch fusion convolution feature map according to the activation layer in the conditional group normalization convolution module to obtain an output feature map.

[0036] Specifically, the conditional group normalization convolution module is expressed by the following formula: , in, represents the branch fusion feature map, represents the convolution operation, represents the convolution kernel weight, represents the bias term, represents the branch fusion convolution feature map, Indicates Group output feature map, Indicates Group branch fusion convolution feature map , and Respectively represent The mean and variance of the group, and denote the learnable scaling and offset parameters, respectively, Represents a constant.

[0037] Further, First, the input branch fusion feature map is convolved using the two-dimensional convolution layer in the conditional group normalization convolution module, where the shape of the branch fusion feature map is (B, C1, H1, W1), B is the batch size, C1 is the number of input channels, H1 and W1 are the height and width of the branch fusion feature map. The purpose of this step is to extract features and generate branch fusion convolution feature maps, where the shape of the branch fusion convolution feature map is (B, C2, H′, W′), H′ and W′ are the height and width of the branch fusion convolution feature map. The convolution operation is performed on the feature map by sliding the convolution kernel, which can capture local features, where the shape of the convolution kernel is (C2, C1, K, K), C2 is the number of output channels, and K is the convolution kernel size. Next, the branch fusion convolution feature map is normalized using a group normalization layer. Group normalization is a normalization method that divides feature maps into several groups and normalizes each group independently. This helps improve the stability and convergence speed of the model, especially when training in small batches. Finally, the normalized branch fusion convolutional feature map is activated through the activation layer. The activation function (such as ReLU, Sigmoid, etc.) introduces nonlinearity, allowing the model to better fit complex functional relationships, thereby obtaining the final output feature map.

[0038] S500: training the improved YOLOv11 network model according to part of the data in the target data set, and testing the improved YOLOv11 network model according to another part of the data; It should be noted that S500 is specifically: S510: Divide the target data set into a training set, a validation set, and a test set according to a preset ratio; S520: construct a loss function, and train the improved YOLOv11 network model according to the training set and the loss function; S530: Verify the trained improved YOLOv11 network model according to the verification set, and optimize the parameters of the improved YOLOv11 network model according to the difference between the verification output of the improved YOLOv11 network model and the true label.

[0039] S540: Testing the improved YOLOv11 network model after parameter optimization according to the test set to obtain a test result.

[0040] Furthermore, first, the target dataset is divided into three parts according to a preset ratio: a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for model tuning and selection, and the test set is used to finally evaluate the performance of the model. For example, in this example, the target dataset can be divided into 80% training set, 10% validation set, and 10% test set. During the training process, a suitable loss function needs to be constructed to measure the difference between the model prediction output and the true label. For an object detection model such as YOLOv11, the loss function usually includes positioning loss (such as bounding box regression loss) and classification loss (such as cross entropy loss of the target category). The improved YOLOv11 network model is trained using the training set and the constructed loss function. Through the back propagation algorithm, the model parameters are optimized to minimize the loss function. Some techniques such as learning rate adjustment, data augmentation, etc. can be used during the training process to improve the generalization ability of the model. After each training cycle (epoch), the trained improved YOLOv11 network model is verified using the validation set. The performance of the model is evaluated by calculating the difference between the output of the model on the validation set and the true label. Based on the verification results, the model parameters can be optimized, such as adjusting the learning rate, modifying the network structure, or performing early stopping. Finally, the improved YOLOv11 network model after parameter optimization is tested using the test set. The test set is data that the model has never seen before and can effectively evaluate the actual performance of the model. By calculating the test results (such as accuracy, recall, F1-score, etc.), you can fully understand the performance of the model in real scenarios.

[0041] S600: Input the image to be detected into the tested improved YOLOv11 network model and output the defect detection result.

[0042] According to a steel surface defect detection method of the present embodiment, the present invention proposes an improved YOLOv11 model based on the steel defect application scenario, and the designed C3k2-MSM module inherits the cross-stage local connection network C3k2, and introduces multi-level pooling branches, dynamic hole convolution, and progressive feature fusion, which solves the problems of fixed structure of traditional algorithms, inability to perform feature fusion and extraction across classes, shallow network retaining details but large computational complexity, deep network with strong semantics but low resolution, and unbalanced lightness and precision. Compared with a single C3k2 feature extraction module, it can target a variety of small and complex defects in steel, capture more fine-grained features and enhance feature processing capabilities, and adapt to different complex application scenarios. The designed RepShared Convolutional Detection Head detection head module improves the decoupling head encoding so that the model can detect the function of rotating targets, solves the problem of discontinuous angle regression of traditional models, and solves the problem of poor module deployment compatibility through a lightweight architecture. Compared with the detection head before improvement, it has significantly improved feature representation and processing capabilities, as well as recognition speed and efficiency, and can be applied to more complex steel recognition scenarios.

[0043] It should be noted that the steel surface defect detection method provided in the embodiment of the present application can be executed by a steel surface defect detection system, or a control module in the steel surface defect detection system for executing and loading a steel surface defect detection method. In the embodiment of the present application, a steel surface defect detection system is used to execute and load a steel surface defect detection method as an example to illustrate a steel surface defect detection method provided in the embodiment of the present application.

[0044] A steel surface defect detection system, comprising: Acquisition module: acquires steel image data set, preprocesses the steel image data set, and obtains the target data set; The first improvement module: build a YOLOv11 network model and add a detection head to the detection layer of the YOLOv11 network model; The second improved module: build the C3k2-MSM module and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module; The third improvement module: replace all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtain an improved YOLOv11 network model; Training module: train the improved YOLOv11 network model based on part of the data in the target data set, and test the improved YOLOv11 network model based on another part of the data; Prediction module: Input the image to be detected into the tested improved YOLOv11 network model and output the defect detection result.

[0045] A steel surface defect detection system in the embodiment of the present application may be a device, or a component, integrated circuit, or chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device may be a laptop computer, a PDA, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.

[0046] A steel surface defect detection system in the embodiment of the present application may be a device having an operating system. The operating system may be a Windows operating system or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0047] The steel surface defect detection system provided in the embodiment of the present application can achieve Figures 1 to 7 The various processes of a steel surface defect detection method implemented in the method embodiment will not be described again here to avoid repetition.

[0048] According to a steel surface defect detection system of the present embodiment, the module is responsible for acquiring a steel image data set and performing preprocessing operations such as denoising, contrast enhancement, and resizing to obtain a target data set suitable for model training. This step is crucial to improving the efficiency and accuracy of subsequent model training. Adding a detection head to the detection layer of the YOLOv11 network model helps the model capture more feature information during the detection process, thereby improving the accuracy and comprehensiveness of the detection. Construct a C3k2-MSM module and replace it with the backbone network of the YOLOv11 network model. The C3k2-MSM module can improve the model's ability to extract features of steel surface defects. Replace all detection heads on the detection layer with heavy-parameter lightweight detection heads. This step is intended to reduce the computational workload and storage requirements of the model while maintaining or improving detection performance, thereby achieving more efficient detection. Use part of the data in the target data set to train the improved YOLOv11 network model. During the training process, the model learns how to extract features from images and identify defects. Use another part of the data in the target data set to test the trained model to evaluate its performance and accuracy. This step is crucial to ensure the reliability of the model in practical applications. The image to be detected is input into the tested improved YOLOv11 network model, and the model will output defect detection results. These results include information such as the location of the defect, which provides an important basis for subsequent repair and processing.

[0049] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned steel surface defect detection method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0050] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned steel surface defect detection method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0051] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0052] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0053] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0054] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for detecting surface defects of steel, characterized in that: The method comprises: Acquire a steel image dataset, and preprocess the steel image dataset to obtain a target dataset; Building a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; Construct a C3k2-MSM module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module; Replacing all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtaining an improved YOLOv11 network model; The improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; The image to be detected is input into the tested improved YOLOv11 network model, and the defect detection result is output.

2. A method for detecting surface defects of steel according to claim 1, characterized in that: The step of obtaining a steel image dataset and preprocessing the steel image dataset to obtain a target dataset is as follows: Acquire a large number of steel images with defects on the surface to construct a steel image dataset according to each of the steel images; Perform defect marking on each of the steel images in the steel image data set, and mark the defect position of each of the steel images; The steel image data set after the defect annotation is subjected to data expansion and enhancement processing to obtain the target data set.

3. A method for detecting surface defects of steel according to claim 1, characterized in that: The data processing process of the C3k2-MSM module is specifically as follows: The feature vector is input into the C3k2-MSM module, and the convolution layer in the C3k2-MSM module performs convolution processing on the feature vector to obtain an initial feature map; Performing feature extraction and transformation processing on the initial feature map according to the deep pooling layer in the C3k2-MSM module to obtain a deep fusion feature map; The initial feature map and the deep fusion feature map are residually connected according to the residual connection layer in the C3k2-MSM module to obtain a target feature map.

4. A method for detecting surface defects of steel according to claim 3, characterized in that: The deep pooling layer performs feature extraction and transformation processing on the initial feature map to obtain a deep fusion feature map; specifically: The deep pooling layer performs multi-scale pooling processing on the initial feature map to obtain multiple pooling feature maps of different scales; Performing convolution processing on each of the pooled feature maps, and then performing dynamic convolution processing on each of the pooled feature maps after the convolution processing, to obtain each dynamic convolution feature map corresponding to each of the pooled feature maps; Performing upsampling processing on each of the dynamic convolution feature maps, and performing residual connection between each of the dynamic convolution feature maps after the upsampling processing and the initial feature map to obtain a residual feature map corresponding to each of the dynamic convolution feature maps; Performing multi-scale feature fusion processing on each of the residual feature maps of different scales to obtain a fused feature map; The fused feature map is activated and convolved to obtain the deep fused feature map.

5. A method for detecting surface defects of steel according to claim 1, characterized in that: The process of data processing by the heavy parameter lightweight detection head is specifically as follows: Inputting the input feature map into the heavy parameter lightweight detection head, and the diversified branch module in the heavy parameter lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; The branch fusion feature map is processed according to the conditional group normalization convolution module to obtain an output feature map.

6. A method for detecting surface defects of steel according to claim 5, characterized in that: The diversified branch module performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically: The standard convolution layer, 1×1 convolution layer, average pooling layer and double convolution layer in the diversified branch module process the input feature map respectively to obtain convolution features, 1×1 features, pooling features and double convolution features respectively; According to the fusion layer in the diversified branch module, the input feature map, the convolution feature, the 1×1 feature, the pooling feature and the double convolution feature are subjected to feature fusion to obtain a fused feature map; Performing activation processing on the fused feature map according to the activation layer in the diversified branch module to obtain the branch fused feature map; The conditional group normalized convolution module processes the branch fusion feature map to obtain an output feature map, specifically: Performing two-dimensional convolution processing on the branch fusion feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain a branch fusion convolution feature map; Normalizing the branch fusion convolution feature map according to the group normalization layer in the conditional group normalization convolution module; The normalized branch fusion convolution feature map is activated according to the activation layer in the conditional group normalization convolution module to obtain the output feature map.

7. A method for detecting surface defects of steel according to claim 1, characterized in that: The improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; specifically: Dividing the target data set into a training set, a validation set, and a test set according to a preset ratio; Constructing a loss function, and training the improved YOLOv11 network model according to the training set and the loss function; Verifying the trained improved YOLOv11 network model according to the verification set, and optimizing parameters of the improved YOLOv11 network model according to the difference between the verification output of the improved YOLOv11 network model and the true label; The improved YOLOv11 network model after parameter optimization is tested according to the test set to obtain test results.

8. A steel surface defect detection system, implementing a steel surface defect detection method according to any one of claims 1 to 7, characterized in that: The system comprises: Acquisition module: acquires a steel image data set, preprocesses the steel image data set, and obtains a target data set; The first improvement module: constructing a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; The second improvement module: constructing a C3k2-MSM module, replacing the C3k2 module in the YOLOv11 network model backbone network with the C3k2-MSM module; The third improvement module: replace all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtain an improved YOLOv11 network model; Training module: training the improved YOLOv11 network model according to part of the data in the target data set, and testing the improved YOLOv11 network model according to another part of the data; Prediction module: inputs the image to be detected into the improved YOLOv11 network model after the test, and outputs the defect detection result.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of a method for detecting surface defects of steel as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps of a steel surface defect detection method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Modulation identification method and device based on asymmetric convolution and parameter fusion, equipment and medium

    CN119052038A

  • Obstacle detection method of visual impaired group guiding waistcoat system based on improved YOLOv11

    CN119600523A

  • PPY-YOLO-based steel surface defect detection method and system

    CN119672031A

  • Road well lid disease detection method based on edge enhanced feature aggregation

    CN119693927A

  • Student classroom behavior detection method based on deep learning

    CN119763179A

Cited By

  • Lightweight target detection method and unmanned aerial vehicle image target detection method

    CN120107572A

  • Concrete building peeling defect detection method, device and system and storage medium

    CN120235885A

  • Concrete building spalling defect detection method, device, system, and storage medium

    CN120235885B

  • Mechanical part defect detection method and system based on improved YOLOv12 model

    CN120279020A

  • Mechanical parts defect detection method and system based on improved YOLOv12 model

    CN120279020B