Steel Surface Defect Detection Method and System
By improving the YOLOv11 network model and introducing the C3k2-MSM module, the problem of difficult real-time and efficiency detection of steel surface defects in the prior art is solved, and more efficient and accurate detection results are achieved.
Patent Information
- Application Number
- CN202510422293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing machine learning methods are difficult to meet the real-time and efficiency requirements of industrial scenarios in steel surface defect detection.
By building an improved YOLOv11 network model, the detection head is added and replaced with a lightweight detection head with a parameter, and the C3k2-MSM module is introduced to optimize feature extraction and processing capabilities.
Improves the accuracy and efficiency of steel surface defect detection, reduces the demand for computing resources, and enables the method to be deployed and applied on a wider range of hardware platforms.
Smart Images

Figure CN119941724B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision technology, and particularly relates to a method and system for detecting steel surface defects. Background Art
[0002] As the core basic material of the modern industrial system, steel occupies an important position in the field of advanced equipment manufacturing. Its application scope covers key scenarios such as aerospace structural components, rail transit equipment components, automobile body and chassis systems, etc. However, during the rolling process, heat treatment, and service process, various types of defects are likely to occur on the steel surface. For example, thermally induced cracks, mechanical damage scratches, oxidation corrosion spots, fatigue pitting, etc. These surface defects will significantly affect the service performance of the material, causing systematic risks such as a decrease in tensile strength, attenuation of ductility, and reduction of corrosion resistance. More seriously, undetected critical-sized defects may lead to catastrophic failures, such as industrial accidents of grade II like pressure vessel explosion and drive shaft fracture. Therefore, it is of great significance to detect defects on the steel surface.
[0003] In the existing machine learning methods for detecting steel surface defects, due to their complex network structure and large number of parameters, the consumption of computing resources is relatively high, resulting in difficulty in meeting the real-time performance and efficiency requirements for defect detection in industrial scenarios. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a method and system for detecting steel surface defects, which can solve the technical problem that the existing machine learning methods are difficult to meet the real-time performance and efficiency requirements for detecting steel surface defects in industrial scenarios.
[0005] To solve the above technical problems, this application is implemented as follows:
[0006] In a first aspect, the embodiments of this application provide a method for detecting steel surface defects, and the method includes:
[0007] Obtain a steel image data set, and preprocess the steel image data set to obtain a target data set;
[0008] Construct a YOLOv11 network model, and add a detection head to the detection layer of the YOLOv11 network model;
[0009] Construct a C3k2-MSM module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module;
[0010] Replace all the detection heads on the detection layer with reparameterized lightweight detection heads; to obtain an improved YOLOv11 network model;
[0011] Train the improved YOLOv11 network model according to part of the data in the target data set, and test the improved YOLOv11 network model according to another part of the data;
[0012] Input the image to be detected into the tested improved YOLOv11 network model to output the defect detection result.
[0013] As an optional implementation manner of the first aspect of this application, the process of obtaining the steel image data set and preprocessing the steel image data set to obtain the target data set is as follows:
[0014] Obtain a large number of steel images with defects on the surface to construct a steel image data set according to each steel image;
[0015] Label the defects of each steel image in the steel image data set, and mark the defect positions of each steel image;
[0016] Perform data augmentation and enhancement processing on the steel image data set after the defect labeling to obtain the target data set.
[0017] As an optional implementation manner of the first aspect of this application, the process of data processing by the C3k2-MSM module is specifically as follows:
[0018] Input the feature vector into the C3k2-MSM module, and the convolutional layer in the C3k2-MSM module performs convolutional processing on the feature vector to obtain an initial feature map;
[0019] According to the depth pooling layer in the C3k2-MSM module, perform feature extraction and transformation processing on the initial feature map to obtain a depth fusion feature map;
[0020] According to the residual connection layer in the C3k2-MSM module, perform residual connection on the initial feature map and the depth fusion feature map to obtain a target feature map.
[0021] As an optional implementation manner of the first aspect of this application, the depth pooling layer performs feature extraction and transformation processing on the initial feature map to obtain a depth fusion feature map; specifically:
[0022] The depth pooling layer performs multi-scale pooling processing on the initial feature map to obtain multiple pooling feature maps with different scales;
[0023] Perform convolutional processing on each pooling feature map, and then perform dynamic convolutional processing on each pooling feature map after the convolutional processing to obtain each dynamic convolutional feature map corresponding to each pooling feature map;
[0024] Upsample each of the dynamic convolution feature maps, and perform residual connection between each of the upsampled dynamic convolution feature maps and the initial feature map to obtain a residual feature map corresponding to each of the dynamic convolution feature maps;
[0025] Perform multi-scale feature fusion processing on each of the residual feature maps of different scales to obtain a fused feature map;
[0026] Perform activation and convolution processing on the fused feature map to obtain the deep fused feature map.
[0027] As an alternative implementation manner of the first aspect of the present application, the process of the reparameterized lightweight detection head for data processing is specifically as follows:
[0028] Input the input feature map into the reparameterized lightweight detection head, and the diversified branch module in the reparameterized lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map;
[0029] Process the branch fusion feature map according to the conditional group normalization convolution module to obtain an output feature map.
[0030] As an alternative implementation manner of the first aspect of the present application, the diversified branch module performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically:
[0031] The standard convolution layer, 1×1 convolution layer, average pooling layer and dual convolution layer in the diversified branch module process the input feature map respectively to obtain convolution features, 1×1 features, pooling features and dual convolution features;
[0032] Perform feature fusion on the input feature map, convolution features, 1×1 features, pooling features and dual convolution features according to the fusion layer in the diversified branch module to obtain a fused feature map;
[0033] Perform activation processing on the fused feature map according to the activation layer in the diversified branch module to obtain the branch fusion feature map;
[0034] The conditional group normalization convolution module processes the branch fusion feature map to obtain an output feature map, specifically:
[0035] Perform two-dimensional convolution processing on the branch fusion feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain a branch fusion convolution feature map;
[0036] Perform normalization processing on the branch fusion convolution feature map according to the group normalization layer in the conditional group normalization convolution module;
[0037] The activation layer in the normalization convolutional module normalizes and activates the fused convolutional feature map of the branch to obtain the output feature map according to the described condition group.
[0038] As an optional implementation manner of the first aspect of the present application, the improved YOLOv11 network model is trained according to a part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; specifically:
[0039] The target data set is divided into a training set, a validation set, and a test set according to a preset ratio;
[0040] A loss function is constructed, and the improved YOLOv11 network model is trained according to the training set and the loss function;
[0041] The trained improved YOLOv11 network model is verified according to the validation set, and the parameters of the improved YOLOv11 network model are optimized according to the difference between the verification output of the improved YOLOv11 network model and the true label;
[0042] The improved YOLOv11 network model with optimized parameters is tested according to the test set to obtain a test result.
[0043] In a second aspect, an embodiment of the present application provides a steel surface defect detection system, the system includes:
[0044] An acquisition module: acquires a steel image data set, and preprocesses the steel image data set to obtain a target data set;
[0045] A first improvement module: constructs a YOLOv11 network model, and adds a detection head to the detection layer of the YOLOv11 network model;
[0046] A second improvement module: constructs a C3k2-MSM module, and replaces the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module;
[0047] A third improvement module: replaces all the detection heads on the detection layer with reparameterized lightweight detection heads; to obtain an improved YOLOv11 network model;
[0048] A training module: trains the improved YOLOv11 network model according to a part of the data in the target data set, and tests the improved YOLOv11 network model according to another part of the data;
[0049] Prediction module: Input the image to be detected into the improved YOLOv11 network model after testing, and output the defect detection result.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0051] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0052] In the embodiment of the present application, compared with the prior art, the following technical effects are achieved:
[0053] (1) By constructing the YOLOv11 network model and adding a detection head to its detection layer, and subsequently replacing all the detection heads on the detection layer with reparameterized lightweight detection heads, the model's ability to identify steel surface defects can be enhanced, thereby improving the detection accuracy.
[0054] (2) The introduction of the C3k2-MSM module may further enhance the model's ability to capture defect features by improving the feature extraction ability, which helps to more accurately identify various types and sizes of defects.
[0055] (3) The use of the reparameterized lightweight detection head helps to reduce the computational load of the model, thereby improving the detection speed while maintaining a high detection accuracy and achieving a fast response.
[0056] (4) The introduction of the reparameterized lightweight detection head helps to reduce the complexity of the model and the model's demand for computing resources, enabling this method to be deployed and applied on a wider range of hardware platforms.
[0057] (5) By constructing and improving the YOLOv11 network model, this method realizes the automated real-time detection of steel surface defects, reduces manual intervention, and improves the consistency and reliability of detection. Description of the Drawings
[0058] Figure 1 is a flowchart of a method for detecting steel surface defects provided by some embodiments of the present application;
[0059] Figure 2 is a structural diagram of an improved YOLOv11 network model for a method for detecting steel surface defects provided by some embodiments of the present application;
[0060] Figure 3It is the structural diagram of the C3k2 - MSM module of a steel surface defect detection method provided by some embodiments of the present application;
[0061] Figure 4 It is the structural diagram of the depth pooling layer of a steel surface defect detection method provided by some embodiments of the present application;
[0062] Figure 5 It is the structural diagram of the reparameterized lightweight detection head of a steel surface defect detection method provided by some embodiments of the present application;
[0063] Figure 6 It is the structural diagram of the diversified branch module of a steel surface defect detection method provided by some embodiments of the present application;
[0064] Figure 7 It is the structural diagram of the conditional group normalization convolution module of a steel surface defect detection method provided by some embodiments of the present application. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0066] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0067] Next, a steel surface defect detection method and system provided by the embodiments of the present application will be described in detail with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0068] Embodiment
[0069] A steel surface defect detection method includes the following steps:
[0070] S100: Obtain a steel image data set, preprocess the steel image data set to obtain a target data set;
[0071] It should be noted that S100 is specifically:
[0072] S110: Obtain a large number of steel images with defects on the surface to construct a steel image dataset based on each steel image;
[0073] S120: Perform defect annotation on each steel image in the steel image dataset, and mark the defect positions of each steel image;
[0074] S130: Perform data augmentation and enhancement processing on the steel image dataset after defect annotation to obtain a target dataset.
[0075] Furthermore, the steel image dataset in S110 uses a self-collected dataset and a publicly available network dataset. The self-collected dataset is obtained by using a two-dimensional camera to take pictures of defective steel to obtain images of steel surface defects. Specifically, a mirrorless camera with the model Nikon Z8 is used to take 300 steel defect pictures at the steel production site to obtain the self-collected dataset; the publicly available network dataset comes from the publicly available NEU-DET steel defect dataset of Northeastern University. In S120, the self-collected dataset is first imported into the annotation tool, and the annotation tool is X-AnyLabeling; the defective parts of these 300 taken pictures are marked and annotated in the yolo format. The annotation file contains information such as the class number of each defect target. In S130, data augmentation methods such as randomly enhancing contrast, noise, flipping, and scaling are used to expand the image data and labels to 3000 pictures, simulating the pictures recognized by the camera in various extreme situations, so as to improve the generalization ability of the training model; the expanded self-collected dataset and the publicly available network dataset are merged to obtain the target dataset.
[0076] S200: Construct a YOLOv11 network model, and add a detection head to the detection layer of the YOLOv11 network model;
[0077] It should be noted that adding a detection head enables the model to have four detection heads, which can extract four feature maps with sizes of 120×120, 64×64, 32×32, and 16×16 from the steel image with defects on the surface, respectively, for detecting 4 different sizes of defect targets on the steel surface, namely tiny, small, medium, and large.
[0078] S300: Construct a C3k2-MSM module, and replace the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module;
[0079] Furthermore, the process of data processing of the C3k2-MSM (Multi-Scale Mixer) module is specifically as follows:
[0080] S310: Input the feature vector into the C3k2-MSM module, and the convolutional layer in the C3k2-MSM module performs convolutional processing on the feature vector to obtain an initial feature map;
[0081] S320: Feature extraction and transformation processing are performed on the initial feature map according to the depth pooling layer in the C3k2-MSM module to obtain a depth fusion feature map;
[0082] S330: Residual connection is performed on the initial feature map and the depth fusion feature map according to the residual connection layer in the C3k2-MSM module to obtain a target feature map.
[0083] It should be noted that in S320, the depth pooling layer performs feature extraction and transformation processing on the initial feature map to obtain a depth fusion feature map; specifically:
[0084] S321: The depth pooling layer performs multi-scale pooling processing on the initial feature map to obtain multiple pooling feature maps with different scales;
[0085] S322: Convolution processing is performed on each pooling feature map, and then dynamic convolution processing is performed on each pooled feature map after convolution processing to obtain each dynamic convolution feature map corresponding to each pooling feature map;
[0086] S323: Upsampling processing is performed on each dynamic convolution feature map, and the upsampled each dynamic convolution feature map is connected with the initial feature map by residual connection to obtain a residual feature map corresponding to each dynamic convolution feature map;
[0087] S324: Multi-scale feature fusion processing is performed on each residual feature map with different scales to obtain a fusion feature map;
[0088] S325: Activation and convolution processing are performed on the fusion feature map to obtain a depth fusion feature map.
[0089] Specifically, the processing formula of the depth pooling layer is expressed as follows:
[0090] ,
[0091] Among them, represents the initial feature map, represents average pooling processing, represents the convolution kernel, represents the stride, represents the pooling feature map with a pooling kernel size of ( × ), represents convolution processing, represents corresponding convolution feature map, represents corresponding dynamic convolution feature map, represents upsampling processing, represents The dynamic convolution feature map after upsampling processing denotes the corresponding residual feature map denotes the fused feature map denotes concatenating all the residual feature maps denotes activation processing denotes the fused feature map after activation processing denotes the depth - fused feature map
[0092] Furthermore, the depth pooling layer first performs multi - scale pooling processing on the initial feature map; by using pooling windows of different sizes (in this embodiment, the pooling kernel sizes are 8×8, 4×4, and 2×2 respectively, and the shape is (B, C, H / i, W / i)), multiple pooling feature maps of different scales are generated. These pooling feature maps can capture feature information at different scales, enhancing the model's perception ability of diverse defects; performing convolution processing on each pooling feature map. The convolution layer extracts features through a set of learnable convolution kernels to generate convolution feature maps. This step can further refine the important information in the pooling feature maps and enhance the feature expression ability. After convolution processing, dynamic convolution processing (convolution kernel size is 3×3, stride is 1, padding is 1) is performed on each convolution feature map. The idea of dynamic convolution is to dynamically adjust the weights of the convolution kernel according to different input feature maps to adapt to different features. This method can improve the flexibility and adaptability of the model, enabling it to better capture complex features. Upsampling processing is performed on each dynamic convolution feature map. The purpose of upsampling is to restore the spatial dimension of the feature map to the same size as the initial feature map, usually using transposed convolution or interpolation methods for upsampling. Residual connection is performed between each upsampled dynamic convolution feature map and the corresponding initial feature map. Through the addition operation, the residual connection can effectively fuse the information in the initial feature map and the information in the dynamic convolution feature map to generate the residual feature map corresponding to each dynamic convolution feature map. Multi - scale feature fusion processing is performed on each residual feature map of different scales. By methods such as weighted average or concatenation, the residual feature maps of different scales are fused together to form a comprehensive feature representation, enhancing the model's detection ability for various defects. Finally, activation processing is performed on the fused feature map, usually using ReLU or other activation functions to introduce non - linear features. This step can further enhance the expression ability of the feature map to obtain the final depth - fused feature map; through multi - scale pooling, dynamic convolution, and residual connection, the depth pooling layer realizes efficient feature extraction and fusion; by accumulating feature maps of different scales, the model can simultaneously utilize global and local information, enhancing the feature expression ability and improving the model's robustness to targets of different scales; applying the GELU activation function to the accumulated feature map enhances the model's expression ability.
[0093] S400: Replace all detection heads on the detection layer with reparameterized lightweight detection heads; obtain an improved YOLOv11 network model;
[0094] It should be noted that the process of the reparameterized lightweight detection head (Rep Shared Convolutional DetectionHead) for data processing is specifically as follows:
[0095] S410: Input the input feature map into the reparameterized lightweight detection head, and the diversified branch module in the reparameterized lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map;
[0096] S420: Process the branch fusion feature map according to the conditional group normalization convolutional module to obtain an output feature map.
[0097] Furthermore, the reparameterized lightweight detection head enables the model to detect rotating targets by improving the decoupled head encoding, and at the same time solves the problem of discontinuous angle regression in traditional models. The problem of poor module deployment compatibility is solved through a lightweight architecture. Compared with the detection head before improvement, there has been a significant improvement in feature representation and processing capabilities, as well as recognition speed and efficiency, and it can be applied to more complex steel identification scenarios. First, the input feature map is processed through a diversified branch module for branch feature extraction and fusion to obtain a branch fusion feature map; then the branch fusion feature map is processed through a conditional group normalization convolutional module to obtain an output feature map. In terms of the module architecture, rotation angle parameters are introduced, and the integration prediction of the center point, width, height, and angle is realized through the fusion of the rotated box and the horizontal box. The angle encoding is adjusted, and the sigmoid activation function is used to constrain the angle prediction to the range of [-π / 4, 3π / 4]. Compared with the scheme of directly predicting 0-π / 2, the problem of boundary mutation is avoided. The anchor point decoding is adjusted, the distance offset is converted into a rotated box, and geometric adaptability is achieved by combining the anchor point size and the feature map stride, making the model detection more flexible.
[0098] It should be noted that in S410, the diversified branch module performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically:
[0099] S411: The standard convolutional layer, 1×1 convolutional layer, average pooling layer, and dual convolutional layer in the diversified branch module process the input feature map respectively to obtain convolutional features, 1×1 features, pooling features, and dual convolutional features;
[0100] S412: According to the fusion layer in the diversified branch module, feature fusion is performed on the input feature map, convolutional features, 1×1 features, pooling features, and dual convolutional features to obtain a fusion feature map;
[0101] S413: Activate the fused feature map according to the activation layer in the diversification branch module to obtain the branch fused feature map.
[0102] Furthermore, the diversification branch module performs a standard convolution operation on the input feature map to extract convolution features; this process slides and calculates the input feature map through a convolution kernel, which can capture local features. The 1×1 convolution layer performs a 1×1 convolution operation on the input feature map to obtain 1×1 features. This convolution method is mainly used for channel compression and feature interaction of features, which can effectively reduce the number of parameters while retaining important information. The average pooling layer performs average pooling on the input feature map to obtain pooled features. Average pooling can reduce the dimension of the feature map while retaining global information by calculating the average value of the local area of the feature map. The double convolution layer performs double convolution on the input feature map to obtain double convolution features. Double convolution usually refers to two consecutive convolution operations, which can extract more complex features. The double convolution layer is composed of two 1×1 convolution layers connected in sequence. In the fusion layer of the diversification branch module, the input feature map, convolution features, 1×1 features, pooled features, and double convolution features are fused. The fusion layer can adopt addition, concatenation, or other fusion strategies to integrate features from different sources to generate a comprehensive fused feature map. This feature fusion can effectively combine the advantages of different feature extraction methods to form a richer and more comprehensive feature representation. Finally, the fused feature map is input into the activation layer for activation processing. The activation layer usually uses ReLU (Rectified Linear Unit) or other activation functions to introduce non-linear features, enabling the model to better learn complex feature relationships. After activation processing, the final branch fused feature map is obtained.
[0103] It should be noted that in S420, the conditional group normalization convolution module processes the branch fused feature map to obtain the output feature map, specifically:
[0104] S421: The conditional group normalization convolution module processes the branch fused feature map to obtain the output feature map, specifically:
[0105] S422: Perform a two-dimensional convolution on the branch fused feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain the branch fused convolution feature map;
[0106] S423: Normalize the branch fused convolution feature map according to the group normalization layer in the conditional group normalization convolution module;
[0107] S424: Activate the branch fused convolution feature map after normalization according to the activation layer in the conditional group normalization convolution module to obtain the output feature map.
[0108] Specifically, the conditional group normalization convolution module is represented by the following formula:
[0109] ,
[0110] where represents the branch fusion feature map, represents the convolution operation, represents the convolution kernel weight, represents the bias term, represents the branch fusion convolution feature map, represents the group of output feature maps, represents the group of branch fusion convolution feature maps , and respectively represent the mean and variance of the group, and respectively represent the learnable scaling and offset parameters, represents a constant.
[0111] Furthermore, first, the two-dimensional convolutional layer in the conditional group normalization convolution module is used to perform a convolution operation on the input branch fusion feature map, where the shape of the branch fusion feature map is (B, C1, H1, W1), B is the batch size, C1 is the number of input channels, and H1 and W1 are the height and width of the branch fusion feature map. The purpose of this step is to extract features and generate the branch fusion convolution feature map, where the shape of the branch fusion convolution feature map is (B, C2, H′, W′), and H′ and W′ are the height and width of the branch fusion convolution feature map. The convolution operation is performed by sliding the convolution kernel on the feature map, which can capture local features, where the shape of the convolution kernel is (C2, C1, K, K), C2 is the number of output channels, and K is the convolution kernel size. Next, the group normalization layer is used to normalize the branch fusion convolution feature map. Group normalization is a normalization method that divides the feature map into several groups and performs independent normalization on each group. This helps to improve the stability and convergence speed of the model, especially during small-batch training. Finally, the normalized branch fusion convolution feature map is activated through the activation layer. Activation functions (such as ReLU, Sigmoid, etc.) introduce non-linearity, enabling the model to better fit complex functional relationships and thus obtain the final output feature map.
[0112] S500: Train the improved YOLOv11 network model according to part of the data in the target dataset, and test the improved YOLOv11 network model according to another part of the data;
[0113] It should be noted that S500 specifically is as follows:
[0114] S510: Divide the target data set into a training set, a validation set, and a test set according to a preset ratio;
[0115] S520: Construct a loss function, and train the improved YOLOv11 network model according to the training set and the loss function;
[0116] S530: Validate the trained improved YOLOv11 network model according to the validation set, and optimize the parameters of the improved YOLOv11 network model according to the difference between the validation output of the improved YOLOv11 network model and the true label.
[0117] S540: Test the improved YOLOv11 network model with optimized parameters according to the test set to obtain test results.
[0118] Furthermore, first, divide the target data set into three parts according to a preset ratio: a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for model tuning and selection, and the test set is used for final evaluation of the model's performance. For example, in this instance, the target data set can be divided into an 80% training set, a 10% validation set, and a 10% test set. During the training process, a suitable loss function needs to be constructed to measure the difference between the model's predicted output and the true label. For a target detection model like YOLOv11, the loss function usually includes a localization loss (such as a bounding box regression loss) and a classification loss (such as a cross-entropy loss for target classes). Use the training set and the constructed loss function to train the improved YOLOv11 network model. Through the backpropagation algorithm, optimize the model parameters to minimize the loss function. Some techniques can be used during the training process, such as learning rate adjustment, data augmentation, etc., to improve the model's generalization ability. After each training epoch, use the validation set to validate the trained improved YOLOv11 network model. By calculating the difference between the model's output on the validation set and the true label, evaluate the model's performance. According to the validation results, the model's parameters can be optimized, such as adjusting the learning rate, modifying the network structure, or performing early stopping, etc. Finally, use the test set to test the improved YOLOv11 network model after parameter optimization. The test set is data that the model has not seen before and can effectively evaluate the model's actual performance. By calculating the test results (such as accuracy, recall, F1-score, etc.), the performance of the model in a real scenario can be comprehensively understood.
[0119] S600: Input the image to be detected into the tested improved YOLOv11 network model and output the defect detection result.
[0120] A method for detecting surface defects of steel. According to the application scenarios of steel defects, the present invention proposes an improved YOLOv11 model. The designed C3k2-MSM module inherits the cross-stage partial connection network C3k2 and introduces multi-level pooling branches, dynamic dilated convolution, and progressive feature fusion, solving the problems of fixed structure in traditional algorithms, inability to perform cross-stage feature fusion and extraction, large computational volume in shallow networks while retaining details, strong semantics but low resolution in deep networks, and imbalance between lightweight and accuracy. Compared with a single C3k2 feature extraction module, it can target various small and complex defects in steel, capture finer-grained features, enhance feature processing capabilities, and adapt to different complex application scenarios. The designed RepShared Convolutional Detection Head module enables the model to detect rotated targets by improving the decoupled head encoding, and at the same time solves the problem of discontinuous angle regression in traditional models. It also solves the problem of poor module deployment compatibility through a lightweight architecture. Compared with the detection head before improvement, it has significantly improved in feature representation and processing capabilities, as well as recognition speed and efficiency, and can be applied to more complex steel recognition scenarios.
[0121] It should be noted that for a method for detecting surface defects of steel provided in an embodiment of the present application, the execution subject can be a system for detecting surface defects of steel, or a control module in the system for detecting surface defects of steel that is used to execute the method for detecting surface defects of steel. In an embodiment of the present application, taking a system for detecting surface defects of steel that executes the method for detecting surface defects of steel as an example, a method for detecting surface defects of steel provided in an embodiment of the present application is described.
[0122] A system for detecting surface defects of steel, comprising:
[0123] An acquisition module: acquiring a steel image data set, preprocessing the steel image data set to obtain a target data set;
[0124] A first improvement module: constructing a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model;
[0125] A second improvement module: constructing a C3k2-MSM module, and replacing the C3k2 module in the backbone network of the YOLOv11 network model with the C3k2-MSM module;
[0126] A third improvement module: replacing all the detection heads on the detection layer with reparameterized lightweight detection heads; obtaining an improved YOLOv11 network model;
[0127] Training module: Train the improved YOLOv11 network model based on part of the data in the target dataset, and test the improved YOLOv11 network model based on another part of the data;
[0128] Prediction module: Input the image to be detected into the tested improved YOLOv11 network model and output the defect detection result.
[0129] The steel surface defect detection system in the embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a laptop computer, a palmtop computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiment of the present application does not make specific limitations.
[0130] The steel surface defect detection system in the embodiment of the present application can be a device with an operating system. The operating system can be a windows operating system or other possible operating systems. The embodiment of the present application does not make specific limitations.
[0131] The steel surface defect detection system provided by the embodiment of the present application can implement Figures 1 to 7 each process implemented by a steel surface defect detection method in the method embodiment. To avoid repetition, it will not be elaborated here.
[0132] A steel surface defect detection system according to this embodiment. This module is responsible for obtaining a steel image dataset and performing preprocessing operations such as denoising, enhancing contrast, and adjusting the size, etc., to obtain a target dataset suitable for model training. This step is crucial for improving the efficiency and accuracy of subsequent model training. Add a detection head to the detection layer of the YOLOv11 network model, which helps the model capture more feature information during the detection process, thereby improving the accuracy and comprehensiveness of detection. Construct a C3k2-MSM module and replace it into the backbone network of the YOLOv11 network model. The C3k2-MSM module can improve the model's ability to extract features of steel surface defects. Replace all the detection heads on the detection layer with reparameterized lightweight detection heads. This step aims to reduce the computational amount and storage requirements of the model while maintaining or improving the detection performance, thus achieving more efficient detection. Use a part of the data in the target dataset to train the improved YOLOv11 network model. During the training process, the model will learn how to extract features from images and identify defects. Use another part of the data in the target dataset to test the trained model to evaluate its performance and accuracy. This step is crucial for ensuring the reliability of the model in practical applications. Input the image to be detected into the improved YOLOv11 network model after testing, and the model will output the defect detection results. These results include information such as the location of the defects, providing an important basis for subsequent repair and processing.
[0133] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above embodiment of the method for detecting steel surface defects and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0134] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above embodiment of the method for detecting steel surface defects and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0135] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.
[0136] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0138] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A method for detecting surface defects of steel, characterized in that: The method comprises: Acquire a steel image dataset, and preprocess the steel image dataset to obtain a target dataset; Building a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; Construct a C3k2-MSM module, and replace the C3k2 module in the YOLOv11 network model backbone network with the C3k2-MSM module, wherein the data processing process of the C3k2-MSM module is specifically as follows: The feature vector is input into the C3k2-MSM module, and the convolution layer in the C3k2-MSM module performs convolution processing on the feature vector to obtain an initial feature map; According to the deep pooling layer in the C3k2-MSM module, the initial feature map is subjected to feature extraction and transformation processing to obtain a deep fusion feature map, wherein specifically: the deep pooling layer performs multi-scale pooling processing on the initial feature map to obtain a plurality of pooling feature maps of different scales; convolution processing is performed on each of the pooling feature maps, and then dynamic convolution processing is performed on each of the pooling feature maps after the convolution processing to obtain each dynamic convolution feature map corresponding to each of the pooling feature maps; upsampling processing is performed on each of the dynamic convolution feature maps, and residual connection is performed between each of the dynamic convolution feature maps after the upsampling processing and the initial feature map to obtain a residual feature map corresponding to each of the dynamic convolution feature maps; multi-scale feature fusion processing is performed on each of the residual feature maps of different scales to obtain a fused feature map; activation and convolution processing are performed on the fused feature map to obtain the deep fusion feature map; Performing a residual connection on the initial feature map and the deep fusion feature map according to the residual connection layer in the C3k2-MSM module to obtain a target feature map; Replacing all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtaining an improved YOLOv11 network model; The improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; The image to be detected is input into the tested improved YOLOv11 network model, and the defect detection result is output.
2. A method for detecting surface defects of steel according to claim 1, characterized in that: The step of obtaining a steel image dataset and preprocessing the steel image dataset to obtain a target dataset is as follows: Acquire a large number of steel images with defects on the surface to construct a steel image dataset according to each of the steel images; Perform defect marking on each of the steel images in the steel image data set, and mark the defect position of each of the steel images; The steel image data set after the defect annotation is subjected to data expansion and enhancement processing to obtain the target data set.
3. A method for detecting surface defects of steel according to claim 1, characterized in that: The process of data processing by the heavy parameter lightweight detection head is specifically as follows: Inputting the input feature map into the heavy parameter lightweight detection head, and the diversified branch module in the heavy parameter lightweight detection head performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; The branch fusion feature map is processed according to the conditional group normalization convolution module to obtain an output feature map.
4. A method for detecting surface defects of steel according to claim 3, characterized in that: The diversified branch module performs branch feature extraction and fusion processing on the input feature map to obtain a branch fusion feature map; specifically: The standard convolution layer, 1×1 convolution layer, average pooling layer and double convolution layer in the diversified branch module process the input feature map respectively to obtain convolution features, 1×1 features, pooling features and double convolution features respectively; According to the fusion layer in the diversified branch module, the input feature map, the convolution feature, the 1×1 feature, the pooling feature and the double convolution feature are subjected to feature fusion to obtain a fused feature map; Performing activation processing on the fused feature map according to the activation layer in the diversified branch module to obtain the branch fused feature map; The conditional group normalized convolution module processes the branch fusion feature map to obtain an output feature map, specifically: Performing two-dimensional convolution processing on the branch fusion feature map according to the two-dimensional convolution layer in the conditional group normalization convolution module to obtain a branch fusion convolution feature map; Normalizing the branch fusion convolution feature map according to the group normalization layer in the conditional group normalization convolution module; The normalized branch fusion convolution feature map is activated according to the activation layer in the conditional group normalization convolution module to obtain the output feature map.
5. A method for detecting surface defects of steel according to claim 1, characterized in that: The improved YOLOv11 network model is trained according to part of the data in the target data set, and the improved YOLOv11 network model is tested according to another part of the data; specifically: Dividing the target data set into a training set, a validation set, and a test set according to a preset ratio; Constructing a loss function, and training the improved YOLOv11 network model according to the training set and the loss function; Verifying the trained improved YOLOv11 network model according to the verification set, and optimizing parameters of the improved YOLOv11 network model according to the difference between the verification output of the improved YOLOv11 network model and the true label; The improved YOLOv11 network model after parameter optimization is tested according to the test set to obtain test results.
6. A steel surface defect detection system, implementing a steel surface defect detection method according to any one of claims 1 to 5, characterized in that: The system comprises: Acquisition module: acquires a steel image data set, preprocesses the steel image data set, and obtains a target data set; The first improvement module: constructing a YOLOv11 network model, and adding a detection head to the detection layer of the YOLOv11 network model; The second improved module: construct a C3k2-MSM module, replace the C3k2 module in the YOLOv11 network model backbone network with the C3k2-MSM module, wherein the data processing process of the C3k2-MSM module is specifically as follows: The feature vector is input into the C3k2-MSM module, and the convolution layer in the C3k2-MSM module performs convolution processing on the feature vector to obtain an initial feature map; According to the deep pooling layer in the C3k2-MSM module, the initial feature map is subjected to feature extraction and transformation processing to obtain a deep fusion feature map, wherein specifically: the deep pooling layer performs multi-scale pooling processing on the initial feature map to obtain a plurality of pooling feature maps of different scales; convolution processing is performed on each of the pooling feature maps, and then dynamic convolution processing is performed on each of the pooling feature maps after the convolution processing to obtain each dynamic convolution feature map corresponding to each of the pooling feature maps; upsampling processing is performed on each of the dynamic convolution feature maps, and residual connection is performed between each of the dynamic convolution feature maps after the upsampling processing and the initial feature map to obtain a residual feature map corresponding to each of the dynamic convolution feature maps; multi-scale feature fusion processing is performed on each of the residual feature maps of different scales to obtain a fused feature map; activation and convolution processing are performed on the fused feature map to obtain the deep fusion feature map; Performing a residual connection on the initial feature map and the deep fusion feature map according to the residual connection layer in the C3k2-MSM module to obtain a target feature map; The third improvement module: replace all detection heads on the detection layer with heavy-parameter lightweight detection heads; obtain an improved YOLOv11 network model; Training module: training the improved YOLOv11 network model according to part of the data in the target data set, and testing the improved YOLOv11 network model according to another part of the data; Prediction module: inputs the image to be detected into the improved YOLOv11 network model after the test, and outputs the defect detection result.
7. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of a method for detecting surface defects of steel as described in any one of claims 1 to 5 are implemented.
8. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps of a steel surface defect detection method as described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
PPY-YOLO-based steel surface defect detection method and system
CN119672031A
Vision-based enhanced omni-directional defect detection apparatus and method
US20250021086A1
Cited By
Industrial workpiece surface defect detection method and system based on TSK fuzzy mapping
CN122115331A