Steel defect detection method and system, electronic equipment and storage medium
By introducing ContextGuided convolution, PConv convolution, SPD-Conv module, EMA module and lightweight cross-scale feature fusion module in steel defect detection, the dependency problem of manual feature extraction in traditional methods is solved, and efficient and accurate detection of multi-scale defects is achieved to adapt to complex backgrounds and small object detection.
Patent Information
- Application Number
- CN202510741163.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The prior art relies on manual feature extraction in steel defect detection, making it difficult to process high-dimensional, nonlinear complex data, and has limited detection capabilities for multiple types and multi-scale defects. The traditional methods are relatively low in efficiency, making it difficult to meet the efficient and accurate detection needs of modern production lines.
ContextGuided convolution and PConv convolution are used to process steel surface defect data, combine the YOLOv8 feature extraction SPD-Conv module and EMA module in the backbone network, fuse the lightweight cross-scale feature fusion module of the Neck part, and introduce the Inner-IoU loss function to construct a steel defect detection model.
It improves the detection accuracy and robustness of the model in complex contexts, can better deal with defects of different scales, especially small target defects, adapt to multi-scale defect detection, improves the accuracy and adaptability of the detection, while maintaining a high inference speed.
Smart Images

Figure CN120259831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a steel defect detection method, system, electronic device and storage medium. Background Art
[0002] Traditional steel defect detection methods, such as manual visual inspection, ultrasonic testing, magnetic particle inspection and radiographic testing, have long played an important role in industrial production. Manual visual inspection relies on the experience and visual ability of inspectors. Although intuitive, it is easily affected by subjective factors and it is difficult to ensure the consistency of detection. Ultrasonic testing uses the characteristics of high-frequency sound waves propagating inside materials to identify defects. It is suitable for detecting hidden defects such as internal cracks, but has limited ability to identify surface defects. Magnetic particle inspection shows surface and near-surface defects by adsorbing ferromagnetic powder with a magnetic field. It is suitable for crack detection of ferromagnetic materials, but cannot detect non-ferromagnetic materials. Radiographic testing (X-ray or γ-ray) detects internal defects through the transmission imaging of materials. It is suitable for high-precision detection, but the equipment cost is high, the operation is complex, and strict radiation safety protection is required. Although these traditional methods still play a role in industrial inspection, most of them rely on manual or specific equipment, and the detection results are greatly affected by operation experience. It is difficult to achieve standardization and automation, and the efficiency is low, making it difficult to meet the requirements of modern production lines for efficient and accurate detection.
[0003] With the development of machine learning technology, methods based on statistical learning have begun to be applied to steel defect detection, improving the automation level of detection. Support Vector Machine (SVM) realizes precise classification of defects by constructing the optimal classification hyperplane in a high-dimensional feature space and performs well in small-sample learning scenarios. Random Forest (RF) enhances the generalization ability of the model through the ensemble learning of multiple decision trees and has certain advantages in dealing with complex defect patterns. The K-Nearest Neighbor (KNN) method classifies based on the similarity between samples and is suitable for defect detection tasks with relatively concentrated feature distributions. These traditional machine learning methods perform well in the classification of steel surface defects, but usually rely on manual feature extraction, are difficult to process high-dimensional and non-linear complex data, and have limited detection capabilities for various types and multi-scale defects.
[0004] In recent years, the rapid development of deep learning has further promoted the progress of steel defect detection technology. Deep neural networks can automatically learn multi-level feature representations of data, avoiding the dependence on manual feature design in traditional methods. In particular, object detection technology extracts spatial features through deep convolutional neural networks (CNNs), improving the accuracy and efficiency of defect detection. Currently, object detection methods are mainly divided into two categories: two-stage detectors and one-stage detectors. Two-stage detectors (such as Faster R-CNN) first generate candidate regions and then perform classification and regression, with high detection accuracy but slow inference speed, making it difficult to meet the high real-time requirements. One-stage detectors (such as the YOLO series and SSD) directly perform end-to-end detection on the input image, with low computational cost and fast detection speed, suitable for real-time detection in steel production lines.
[0005] In summary, in the existing technology, traditional machine learning methods still rely on manual feature extraction in the classification of steel surface defects, making it difficult to process high-dimensional and non-linear complex data, and having limited detection capabilities for various types and multi-scale defects. Summary of the Invention
[0006] Based on this, the purpose of the present invention is to provide a steel defect detection method, system, electronic device, and storage medium to solve the above deficiencies in the existing technology.
[0007] In the first aspect, the present invention provides a steel defect detection method, which includes: Collect a steel surface defect data set and sequentially organize and transform the surface defect data set; Process the organized and transformed surface defect data set based on ContextGuided convolution and PConv convolution in sequence to obtain a processed surface defect data set; Introduce an SPD-Conv module into the YOLOv8 feature extraction backbone network and embed an EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model; Fuse the Neck part in the feature extraction model, add a lightweight cross-scale feature fusion module, and introduce an Inner-IoU loss function into the feature extraction model to obtain an updated feature extraction model; Divide the processed surface defect data set into a training set and a test set, and train the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file; Detect the test set based on the weight file and the steel defect detection model to obtain a detection result.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: Through ContextGuided convolution and PConv convolution, the detailed features on the steel surface can be extracted, so that the extraction of detailed features does not need to rely on manual work. By introducing the SPD-Conv module and embedding the EMA module into the YOLOv8 feature extraction backbone network, the detection effect of the model in complex backgrounds is effectively improved, and different scales of defects can be better processed, improving the detection accuracy and robustness. By fusing the Neck part and adding a lightweight cross-scale feature fusion module, the feature information of different scales can be fully integrated, improving the adaptability of the model in multi-scale defect detection, especially significantly improving the detection of small target defects. And by introducing the Inner-IoU loss function, the target localization ability can be optimized.
[0009] Further, the steps of sorting and converting the surface defect dataset in sequence include: Sort the surface defect dataset to label the defects in each image in the surface defect dataset; Convert the format of the sorted surface defect dataset.
[0010] Further, the steps of processing the sorted and converted surface defect dataset based on ContextGuided convolution and PConv convolution in sequence include: Extract local features from the sorted and converted surface defect dataset based on the ContextGuided convolution to obtain local features; Extract the context information of the local features based on the PConv convolution, jointly fuse features to fuse the co-occurrence relationship between the local features and the surrounding environment, and dynamically adjust the co-occurrence relationship through global semantic information.
[0011] Further, the steps of introducing the SPD-Conv module into the YOLOv8 feature extraction backbone network and embedding the EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network include: Replace the CBS module in the YOLOv8 feature extraction backbone network and introduce the SPD-Conv neural network building block; Design the C2f_EMA module with an attention mechanism and introduce the C2f_EMA module into the YOLOv8 feature extraction backbone network to replace the C2f module in the YOLOv8 feature extraction backbone network; Embed the C2f_EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network.
[0012] Further, the step of fusing the Neck part in the feature extraction model and adding a lightweight cross-scale feature fusion module includes: Adding a lightweight cross-scale feature fusion module to the Neck part, and fusing the Neck part after adding the lightweight cross-scale feature fusion module into the feature extraction model; Adding convolution operations before and after each upsampling in the feature extraction model, and connecting the upsampling with the concatenation operation.
[0013] Further, the step of dividing the processed surface defect dataset into a training set and a test set includes: Dividing the processed surface defect dataset based on a stratified sampling strategy to obtain a training set, a validation set, and a test set; Storing the training set, the validation set, and the test set in corresponding folders respectively.
[0014] Further, after the step of detecting the test set based on the weight file and the steel defect detection model to obtain a detection result, the method further includes: Evaluating the detection result based on an object detection model and evaluation metrics, where the evaluation metrics include accuracy, recall, and mean average precision; The calculation expression of the accuracy is: ; In the formula, represents the accuracy, represents the number of samples correctly classified as positive by the steel defect detection model, represents the number of samples misclassified as positive by the steel defect detection model but actually negative; The calculation expression of the recall is: ; In the formula, represents the recall, represents the number of samples misclassified as negative by the steel defect detection model but actually positive; The calculation expression of the mean average precision is: ; In the formula, represents the mean average precision, represents the total number of categories, represents the average precision of the i-th category, represents the average precision.
[0015] In a second aspect, the present invention also provides a steel defect detection system, which includes: A collection module, configured to collect a steel surface defect data set, and sequentially organize and transform the surface defect data set; A processing module, configured to process the surface defect data set after being organized and transformed based on ContextGuided convolution and PConv convolution in sequence to obtain a processed surface defect data set; An introduction module, configured to introduce an SPD-Conv module into the YOLOv8 feature extraction backbone network, and embed an EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model; A fusion module, configured to fuse the Neck part in the feature extraction model, add a lightweight cross-scale feature fusion module, and introduce an Inner-IoU loss function in the feature extraction model to obtain an updated feature extraction model; A training module, configured to divide the processed surface defect data set into a training set and a test set, and train the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file; A detection module, configured to detect the test set based on the weight file and the steel defect detection model to obtain a detection result.
[0016] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned steel defect detection method is implemented.
[0017] In a fourth aspect, the present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned steel defect detection method is implemented. Description of the Drawings
[0018] Figure 1 It is a flowchart of the steel defect detection method in the first embodiment of the present invention; Figure 2 It is a schematic diagram of the feature extraction backbone network process in the first embodiment of the present invention; Figure 3 It is a schematic diagram of the SPD-Conv process in the first embodiment of the present invention; Figure 4 It is a schematic diagram of the C2f_EMA process in the first embodiment of the present invention; Figure 5Schematic diagram of the CCFM process introduced into the feature fusion module in the first embodiment of the present invention; Figure 6 Block diagram of the steel defect detection system in the second embodiment of the present invention; Figure 7 Schematic diagram of the hardware structure of the electronic device in the third embodiment of the present invention.
[0019] Description of the main component symbols: 10. Collection module; 20. Processing module; 30. Introduction module; 40. Fusion module; 50. Training module; 60. Detection module; 70. Bus; 71. Processor; 72. Memory; 73. Communication interface.
[0020] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific embodiments
[0021] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0022] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0024] Embodiment 1 Please refer to Figure 1 , which shows the steel defect detection method in the first embodiment of the present invention. The method includes steps S1 to S6: S1. Collect a dataset of steel surface defects, and sequentially organize and transform the surface defect dataset; Specifically, step S1 includes steps S11 to S12: S11, sorting the surface defect dataset to mark defects in each image in the surface defect dataset; It can be understood that a public surface defect dataset of steel is collected. In this embodiment, the surface defect dataset includes six typical surface defects, specifically rolling scale (Rs), plaque (Pa), crack (Cr), pitting surface (Ps), inclusion (In) and scratch (Sc), and the surface defect dataset includes 1800 grayscale images, six different types of typical surface defects, and each type of surface defect includes 300 samples; When collating the surface defect dataset, the original image data and the corresponding annotation files are first sorted to ensure that each image has the corresponding defect annotation and to remove damaged or invalid data; S12, converting the format of the sorted surface defect data set; It is understandable that the surface defect dataset is uniformly converted into YOLO format, and each image corresponds to a .txt label file with the same name, and the annotation content is<class_id> ,<x_center> ,<y_center> , <width> 、 <height>, all coordinates are normalized values relative to the width and height of the image. During the specific conversion process, the top-left and bottom-right coordinates of the target box are extracted from the original annotation, the center point coordinates and width and height of the target are calculated, and the normalization process is completed. Finally, the images and corresponding labels are organized into the folders of train / images, train / labels, val / images, val / labels, test / images, and test / labels according to the directory structure required by the YOLO model, and the data.yaml file is configured to specify the class names and the paths of each dataset.
[0025] In addition, it is worth noting that the relevant experimental environment needs to be configured. For the specific experimental environment configuration, please refer to Table 1: Table 1
[0026] In the default.yaml file of the model, according to system resources, model, and dataset size, etc., the batch size is set to 16, the input image resolution size is adjusted to 640×640 pixels, the number of worker threads for data loading is 8, the initial learning rate is 0.01, the weight decay of the optimizer is adjusted to 0.0005 to prevent overfitting, the intersection over union (IoU) threshold for non-maximum suppression (NMS) is 0.7, and in addition, the total number of epochs for all model training is 300.
[0027] S2. Based on the ContextGuided convolution and the PConv convolution, the sorted and transformed surface defect dataset is processed in sequence to obtain the processed surface defect dataset; Specifically, the step S2 includes steps S21 to S22: S21. Based on the ContextGuided convolution, local features are extracted from the sorted and transformed surface defect dataset to obtain local features; S22. Based on the PConv convolution, the context information of the local features is extracted, and joint feature fusion is used to fuse the co-occurrence relationship between the local features and the surrounding environment, and the co-occurrence relationship is dynamically adjusted through global semantic information; It can be understood that extracting local features from the surface defect dataset based on the ContextGuided convolution is used for the local detail features of the steel surface, and extracting the context information of the local features through the PConv convolution can extract the information of the surrounding area of the layout features, solve the problem of local occlusion, and effectively enhance the expression ability of the features by fusing the co-occurrence relationship between the local features and the surrounding environment through joint feature fusion. Moreover, using global semantic information to achieve global context weighting can dynamically adjust the importance of the features.
[0028] S3. Introduce the SPD-Conv module into the YOLOv8 feature extraction backbone network, and embed the EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model; Specifically, step S3 includes steps S31 to S33: S31. Replace the CBS module in the YOLOv8 feature extraction backbone network and introduce the SPD-Conv neural network building block; S32. Design the C2f_EMA module with an attention mechanism and introduce the C2f_EMA module into the YOLOv8 feature extraction backbone network to replace the C2f module in the YOLOv8 feature extraction backbone network; S33. Embed the C2f_EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network; It can be understood that introducing the SPD-Conv module into the YOLOv8 feature extraction backbone network to replace the original CBS module can enhance the perception ability of low-resolution regions and small targets. Design the C2f_EMA module that integrates the EMA attention mechanism on the basis of the original C2f module. By embedding the C2f_EMA module between the Split division layer and the Bottleneck neck layer, redundant feature information can be effectively suppressed, the attention of the network to key target regions can be improved, and a more efficient feature extraction model can be constructed. The network structure diagram is as Figure 2 shown.
[0029] In specific implementation, it includes the following steps 101 to 113: Step 101. Input the training set images into the feature extraction model and process them through the SPD-Conv module. The process of the SPD-Conv module is as shown in Figure 3 , Figure 3 In (a) of which, the size of the input feature map is (C1, S, S), that is, (64, 640, 640), where C1 is the number of channels and S is the spatial resolution; Step 102. Then, Figure 3 In (b) of which, through the Space-to-Depth transformation operation, Figure 3 In (c), it is rearranged into the channel dimension for each 2×2 spatial region. Each 2×2 spatial block is split into 4 pixel points, and the information of these 4 pixel points is placed in 4 positions in the channel dimension. One 2×2 small region in each original channel now becomes one pixel point corresponding in 4 channels. Each original channel contributes 4 new channels. In the width and height directions, every 2 pixels are merged into 1, that is, the spatial size is reduced to half of the original, so that Figure 3 In (d), the size of the feature map becomes (4C1, S / 2, S / 2), that is, (256, 320, 320). Figure 3 (a) to (e) in it represent the process of SPD-Conv; Step S103. Subsequently, these rearranged feature maps are divided into four sub-regions, which respectively represent the local information of the original feature map at different spatial positions. The size of each sub-region is (C1, S / 2, S / 2), that is, (64, 320, 320). These four sub-feature maps are concatenated in the channel dimension to obtain a new feature map with a size of (4C1, S / 2, S / 2), that is, (256, 320, 320); Step S104. Finally, a standard convolution operation with a stride of 1 is applied to this feature map, and a feature map with a size of (C2, S / 2, S / 2), that is, (64, 320, 320), is output. The semantic information of the output feature map is richer and the receptive field is larger. The value of C2 is set according to requirements. Through this "space-for-channel" strategy, SPD-Conv can retain local structural information while reducing the spatial resolution, thereby enhancing the feature expression ability of low-resolution images and small objects; Step S105. The output feature map size (64, 320, 320) in Step 104 is output as a feature map size (128, 160, 160) after being processed by SPD-Conv, and then it is used as the input to the C2f_EMA module that embeds EMA between the Split division layer and the Bottleneck neck layer based on C2f. Specifically, as Figure 4 shown, the C2f_EMA module has an original feature map size of 128×160×160, indicating that the number of channels is 128 and the spatial size is 160×160, containing the feature information of the original local region; Step S106. The input feature map passes through a convolution operation, that is, the initial CBS (Conv+BN+SiLU): A 1×1 convolution is used to perform a linear combination on the input feature map, which helps the information reorganization and feature compression between channels. Then it is normalized through BatchNorm, and finally the SiLU activation function is used to enhance the non-linear expression ability. The output is still 128×160×160, which is beneficial for subsequent feature extraction; Step S107, then perform channel splitting (Split): Split the 128 channels along the channel dimension into two parts: 64×160×160 and 64×160×160, named respectively as and . Among them, is the shallow - layer information for subsequent splicing, serving as the input to the feature enhancement path; Step S108, then enter the EMA attention module: First, perform global average pooling to extract channel - level global context information (output 64×1×1), then generate compressed channel attention weights (1×1×64) through two - layer fully - connected networks, and then perform channel - by - channel weighted operation with to enhance important channels and suppress redundant channels. The output is 64×160×160, and the enhanced feature map has stronger discriminative ability; Step S109, after the feature map 64×160×160, pass through the bottleneck layer (Bottleneck) ×n: The enhanced further enters the bottleneck layer residual module n times for deep - level feature extraction. Each block consists of two layers of convolution, combined with residual connections, which can extract multi - scale semantic features while retaining the spatial distribution information of the original features. The output of each time is 64×160×160, with a total of n; Step S110, then perform feature concatenation (Concat): Concatenate the original shallow - layer feature , the attention - enhanced feature, and the outputs of n bottleneck layers in the channel dimension to form a high - dimensional feature map of (n + 2)×64×160×160, realizing the fusion of shallow - layer + deep - layer + enhanced features, effectively improving the expression ability of small and medium - sized targets and edge information in detection; Step S111, the finally generated high - dimensional feature map is further processed by a CBS (Conv+BN+SiLU): Perform 1×1 convolution on the high - dimensional concatenated feature map to compress the number of channels back to the original input channel number (128), and ensure feature stability and non - linear expression through BN and SiLU. The output size is 128×160×160. This operation integrates the fused features and maintains the original size, facilitating subsequent connection to the backbone network. The output feature map has the same size as the input, which is 128×160×160, but contains multi - scale and multi - dimensional fused features extracted from the shallow - layer, deep - layer, and attention mechanism, with stronger expression ability and discriminative ability; Step S112: After processing the output feature map with a size of 128×160×160 in Step 111 through 3 SPD-Convs and 3 C2f_EMAs, the representational ability of the feature map is enhanced. Each time the C2f_EMA module processes, the number of channels of the feature map gradually increases, and then higher-dimensional feature maps are formed through feature concatenation, containing more information. Finally, the dimension of the output feature map usually remains the same as the input. After being processed by 3 SPD-Convs and 3 C2f_EMAs, both the number of channels and the feature representational ability of the feature map are enhanced, enabling better handling of features in low-resolution, small-object, and complex-background scenarios, improving the detection accuracy of details and small objects. The finally output feature map with a size of 1024×20×20 has stronger expressive ability and the ability to adapt to complex scenarios; Step S113: Finally, after the Spatial Pyramid Pooling (SPPF) operation, SPPF keeps the spatial size and the number of channels of the feature map unchanged (the input and output are of the same dimension), but enhances the multi-scale spatial context information, especially significantly improving multi-scale object detection. The feature map will become more diverse and contain more spatial hierarchical information.
[0030] S4: Incorporate the Neck part in the feature extraction model, add a lightweight cross-scale feature fusion module, and introduce the Inner-IoU loss function in the feature extraction model to obtain an updated feature extraction model; Specifically, the said Step S4 includes Step S41 to Step S42: S41: Add a lightweight cross-scale feature fusion module to the Neck part, and fuse the Neck part with the added lightweight cross-scale feature fusion module into the feature extraction model; S42: Add convolution operations before and after each upsampling in the feature extraction model, and connect the upsampling with the concatenation operation; It can be understood that adding a lightweight cross-scale feature fusion module to the feature fusion Neck part combines multi-scale extraction, convolution fusion, and channel compression; In specific implementation, in the feature fusion Neck part of the original YOLOv8 model, a multi-scale feature fusion module CCFM is introduced. For details, please refer to Figure 5 , a convolutional operation is added before and after each upsampling, and the upsampling is connected to the concatenation operation. After the high-dimensional feature map 1024×20×20 is output from the feature extraction network, it enters the feature fusion network. First, a CBS convolutional operation is performed, and then an upsampling operation. If directly upsampled, two problems will occur. One is that the number of channels is too large, which will increase the computational burden. The other is semantic information redundancy, and directly concatenating the underlying features is not conducive to fusion. After the upsampling operation, a feature map with a larger spatial scale is output, which contains the features extracted from the input feature map and has an improvement in spatial resolution; In addition, the upsampling operation is connected to the concatenation operation. The upsampling operation aligns the high-level feature map in the spatial dimension with the low-level features, and the concatenation operation realizes the aggregation of multi-scale information. During the concatenated fusion process, the CBS convolutional operation is further combined to further integrate and reconstruct the features, and the output feature map is 512×40×40, alleviating the problem of inconsistent feature distribution; After that, through a series of feature fusion (C2f, Conv, Upsample, and Concat) operations, the model converts the output feature maps after passing through the last three C2f layers into the input feature maps for defect detection respectively. Among the outputs after these three C2f layers, the first output is converted into the input feature map for detecting small targets, the second output is converted into the input feature map for medium targets, and the last output is converted into the input feature map for detecting large targets. After the model detects the small, medium, and large feature maps, it will output the corresponding sets of prediction boxes (position, confidence, category) at the corresponding scales respectively. Finally, they are merged and filtered through non-maximum suppression (NMS) to form the final object detection result.
[0031] S5, divide the processed surface defect dataset into a training set and a test set, and train the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file; Specifically, step S5 includes steps S51 to S52: S51, divide the processed surface defect dataset based on a stratified sampling strategy to obtain a training set, a validation set, and a test set; S52, store the training set, the validation set, and the test set into corresponding folders respectively; It should be noted that the surface defect dataset is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2. During the division, a stratified sampling strategy is adopted to ensure that the proportion of various defect samples in each subset is the same. When dividing, the random seed setting of sklearn's train_test_split is added to ensure the reproducibility of each division. At the same time, random.shuffle is used to randomly shuffle the data to avoid the problem of partial order in the sample set. After the division, the images and annotation files of different subsets are stored in the corresponding folders (images / train, images / val, images / test, and labels / train, etc.). For the distribution of labels in the training set, validation set, and test set, please refer to Table 2 for details: Table 2
[0032] It can be understood that the experiment is carried out according to the relevant experimental environment and experimental configuration. The input image size of the training set images of the steel defect dataset is set to 640*640, and then input into the updated feature extraction model for training to obtain a model for steel defect detection and a weight file.
[0033] S6. Based on the weight file and the steel defect detection model, the test set is detected to obtain the detection results; It can be understood that after obtaining the detection results, the detection results need to be evaluated. In this embodiment, the detection results are evaluated based on the object detection model and evaluation metrics. The evaluation metrics include accuracy, recall rate, and mean average precision (mAP); the calculation expression of the accuracy is: ; In the formula, represents the accuracy, represents the number of samples correctly classified as positive by the steel defect detection model, represents the number of samples misclassified as positive but actually negative by the steel defect detection model; The calculation expression of the recall rate is: ; In the formula, represents the recall rate, represents the number of samples misclassified as negative but actually positive by the steel defect detection model; The calculation expression of the mean average precision (mAP) is: ; In the formula, represents the mean average precision (mAP), represents the total number of categories, represents the average precision of the i-th category, represents the average precision; It should be noted that the average precision is the area under the precision-recall curve. AP comprehensively reflects the overall performance of the model by integrating precision and recall at different thresholds. The calculation expression of the average precision is: ; In the formula, represents the exact value, represents the recall rate. By integrating the precision with respect to the recall rate from 0 to 1, the area under the precision-recall curve is calculated to obtain AP. This reflects the overall performance of the model at different thresholds.
[0034] The mean average precision value is a comprehensive indicator for measuring the performance of multi-class classification or detection models. It calculates the average precision of each category and then takes the average of the APs of all categories. In this embodiment, mAP@0.5 and mAP@0.5:0.95 are selected as evaluation indicators to verify the effectiveness of the detection results; In addition, FPS (Frames Per Second) is an important indicator for measuring the performance of video or graphics processing, representing the number of frames processed or displayed per second, and is used to evaluate the fluency, real-time processing ability, and performance of the model system, etc. The calculation expression of the number of frames processed or displayed per second is: ; In the formula, FPS represents the number of frames processed or displayed per second, and Processing time per frame represents the processing time per frame; Furthermore, an ablation experiment is conducted on the steel defect detection model. For the specific experimental results, please refer to Table 3: Table 3
[0035] As can be seen from Table 3, in Experimental Group 1, after the SPD-Conv module was introduced alone, the number of model parameters decreased slightly, but both mAP@0.5 and FPS decreased, indicating that although SPD-Conv reduced the model complexity, it did not significantly improve the overall performance when applied alone. In Experimental Group 2, after the C2f_EMA module was introduced, the detection accuracy increased significantly, but the frame rate decreased, indicating that this module performed excellently in enhancing the feature extraction ability but increased a certain amount of computational overhead. In Experimental Group 3, only the CCFM module was introduced, and the detection accuracy was also improved. However, due to the increase in the amount of calculation, the FPS decreased. In Experimental Group 4, the SPD-Conv and C2f_EMA modules were jointly introduced. While slightly reducing the number of model parameters, the detection accuracy and frame rate were significantly improved. The mAP@0.5 was increased by 2.1% compared with the original model, indicating that there was a good synergistic effect between the two in performance optimization. In Experimental Group 5, the SPD-Conv and CCFM modules were combined. Although the number of parameters decreased, the improvement in detection accuracy was limited, indicating that this combination had certain advantages in model lightweighting but had little effect on performance improvement. In Experimental Group 6, after the C2f_EMA and CCFM modules were jointly used, both the detection accuracy and FPS increased, but the overall effect was still inferior to the best combination. Finally, in Experimental Group 7, the SPD-Conv, C2f_EMA, and CCFM modules were comprehensively introduced, and it performed optimally in all indicators, achieving the highest mAP@0.5 (78.6%), the lowest number of parameters (1.68M), and a relatively high frame rate (270.2). Although the FPS decreased compared with the original model, SPD-Conv improved the feature extraction efficiency, C2f_EMA optimized the smoothness of the prediction results, and CCFM enhanced the detection ability of multi-scale targets. The synergistic effect of the three effectively balanced the detection accuracy, model complexity, and inference speed.
[0036] In summary, for the steel defect detection method in the embodiments of the present invention, 1. Replace the convolution in the YOLOv8 feature extraction backbone network with the SPD-Conv (Sparse Partial Depthwise Convolution) module. This module can focus on the key defect areas in the image by adaptively adjusting the weights of the convolution kernels, thus significantly improving the detection accuracy of low-resolution and small-target defects, especially having advantages in detecting small-scale defects. 2. Design the C2f_EMA (C2f with Efficient Multi-scale Attention) module. This module enhances the expression ability of multi-scale features. By introducing the multi-scale attention mechanism, it effectively improves the detection effect of the model in complex backgrounds, enabling it to better handle defects of different scales and improving the detection accuracy and robustness. 3. Introduce the lightweight cross-scale feature fusion module CCFM (Cross-Scale Feature Fusion Module). SCCI-YOLO adds the CCFM module to the Neck network, which can fully integrate feature information of different scales, improve the adaptability of the model in multi-scale defect detection, and especially significantly improve the detection of small target defects. 4. Replace the loss function with the Inner-IoU loss function. Aiming at the deficiencies of the traditional IoU loss function, SCCI-YOLO introduces the Inner-IoU loss function. This loss function can more accurately measure the matching degree between target boxes, thereby improving the regression accuracy, accelerating the model convergence speed, and optimizing the target localization ability. 5. Balance of high precision and high efficiency. Through the above optimization modules, SCCI-YOLO can maintain a high inference speed while improving the detection accuracy, and can meet the requirements of industrial production lines for real-time detection, especially performing well in complex background and small target defect detection.
[0037] Embodiment 2 The present invention also proposes a steel defect detection system. Please refer to Figure 6 , which shows the steel defect detection system in the second embodiment of the present invention. The system includes: A collection module 10, which is used to collect the steel surface defect data set and sequentially organize and transform the surface defect data set. A processing module 20, which is used to process the organized and transformed surface defect data set based on the ContextGuided convolution and the PConv convolution in sequence to obtain a processed surface defect data set. An introduction module 30, which is used to introduce the SPD-Conv module into the YOLOv8 feature extraction backbone network and embed the EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model. A fusion module 40, which is used to fuse the Neck part in the feature extraction model, add a lightweight cross-scale feature fusion module, and introduce the Inner-IoU loss function into the feature extraction model to obtain an updated feature extraction model. A training module 50, which is used to divide the processed surface defect data set into a training set and a test set, and train the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file. A detection module 60, which is used to detect the test set based on the weight file and the steel defect detection model to obtain a detection result.
[0038] In some alternative embodiments, the collection module 10 includes: An arrangement unit for arranging the surface defect data set to label the defects in each image in the surface defect data set; A conversion unit for converting the format of the arranged surface defect data set.
[0039] In some alternative embodiments, the processing module 20 includes: An extraction unit for performing local feature extraction on the arranged and converted surface defect data set based on the ContextGuided convolution to obtain local features; A fusion unit for extracting context information of the local features based on the PConv convolution, jointly fusing features to fuse the co-occurrence relationship between the local features and the surrounding environment, and dynamically adjusting the co-occurrence relationship through global semantic information.
[0040] In some alternative embodiments, the introduction module 30 includes: A replacement unit for replacing the CBS module in the YOLOv8 feature extraction backbone network and introducing the SPD-Conv neural network building block; A design unit for designing a C2f_EMA module with an attention mechanism and introducing the C2f_EMA module into the YOLOv8 feature extraction backbone network to replace the C2f module in the YOLOv8 feature extraction backbone network; An embedding unit for embedding the C2f_EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network.
[0041] In some alternative embodiments, the fusion module 40 includes: An addition fusion unit for adding a lightweight cross-scale feature fusion module to the Neck part and fusing the Neck part with the added lightweight cross-scale feature fusion module into the feature extraction model; A connection unit for adding convolution operations before and after each upsampling in the feature extraction model and connecting the upsampling with the splicing operation.
[0042] In some alternative embodiments, the training module 50 includes: A division unit for dividing the processed surface defect data set based on a hierarchical sampling strategy to obtain a training set, a validation set, and a test set; A storage unit for storing the training set, the validation set, and the test set in corresponding folders respectively.
[0043] In some alternative embodiments, the detection module 60 includes: An evaluation unit for evaluating the detection result based on a target detection model and evaluation metrics, where the evaluation metrics include accuracy, recall rate, and mean average precision; The calculation expression of the accuracy is: ; In the formula, represents the accuracy, represents the number of samples correctly classified as positive by the steel defect detection model, represents the number of samples misclassified as positive but actually negative by the steel defect detection model; The calculation expression of the recall rate is: ; In the formula, represents the recall rate, represents the number of samples misclassified as negative but actually positive by the steel defect detection model; The calculation expression of the mean average precision is: ; In the formula, represents the mean average precision, represents the total number of categories, represents the average precision of the i-th category, represents the average precision.
[0044] The functions or operation steps implemented when the above modules and units are executed are substantially the same as those in the above method embodiments, and will not be elaborated here.
[0045] The steel defect detection system provided by the embodiments of the present invention has the same implementation principle and technical effects as those in the foregoing method embodiments. For a brief description, for the parts not mentioned in the system embodiments, reference may be made to the corresponding content in the foregoing method embodiments.
[0046] Embodiment III The third embodiment of the present invention further proposes an electronic device. Please refer to Figure 7 , which shows a schematic hardware structure diagram of the electronic device in the third embodiment of the present invention.
[0047] The electronic device may include a processor 71 and a memory 72 storing computer program instructions.
[0048] Specifically, the above-mentioned processor 71 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the present application.
[0049] Among them, the memory 72 may include a mass memory for data or instructions. By way of example and not limitation, the memory 72 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 72 may include removable or non-removable (or fixed) media. Where appropriate, the memory 72 may be internal or external to the data processing device. In a particular embodiment, the memory 72 is non-volatile memory. In a particular embodiment, the memory 72 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended date out dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0050] The memory 72 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 71.
[0051] The processor 71 reads and executes the computer program instructions stored in the memory 72 to implement the steel defect detection method in the first embodiment above.
[0052] In some of the embodiments, the electronic device may further include a communication interface 73 and a bus 70. Among them, as Figure 7 shown, the processor 71, the memory 72, and the communication interface 73 are connected through the bus 70 and complete communication with each other.
[0053] The communication interface 73 is used to implement communication between the various modules, devices, units, and / or devices in the present application. The communication interface 73 can also implement data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0054] Bus 70 includes hardware, software, or both, and couples components of the device together. Bus 70 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 70 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In a suitable case, Bus 70 may include one or more buses. Although the present application describes and illustrates specific buses, the present application contemplates any suitable bus or interconnect.
[0055] The electronic device can obtain a steel defect detection system and execute the steel defect detection method of Embodiment 1.
[0056] In addition, in combination with the steel defect detection method in Embodiment 1 above, the present application can be implemented by providing a storage medium. Computer program instructions are stored on the storage medium; when the computer program instructions are executed by a processor, the steel defect detection method of Embodiment 1 above is implemented.
[0057] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0058] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.< / height> < / width>
Claims
1. A method for detecting steel defects, characterized in that, The method includes: Collect a dataset of steel surface defects, and sequentially organize and transform the surface defect dataset; Process the organized and transformed surface defect dataset based on ContextGuided convolution and PConv convolution in sequence to obtain a processed surface defect dataset; Introduce an SPD-Conv module into the YOLOv8 feature extraction backbone network, and embed an EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model; Fuse the Neck part in the feature extraction model, add a lightweight cross-scale feature fusion module, and introduce an Inner-IoU loss function into the feature extraction model to obtain an updated feature extraction model; Divide the processed surface defect dataset into a training set and a test set, and train the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file; Detect the test set based on the weight file and the steel defect detection model to obtain a detection result.
2. The steel defect detection method according to claim 1, characterized in that, The step of sequentially organizing and transforming the surface defect dataset includes: Organize the surface defect dataset to label the defects in each image in the surface defect dataset; Convert the format of the organized surface defect dataset.
3. The steel defect detection method according to claim 1, characterized in that, The step of processing the organized and transformed surface defect dataset based on ContextGuided convolution and PConv convolution in sequence includes: Extract local features from the organized and transformed surface defect dataset based on the ContextGuided convolution to obtain local features; Extract the context information of the local features based on the PConv convolution, jointly fuse features to fuse the co-occurrence relationship between the local features and the surrounding environment, and dynamically adjust the co-occurrence relationship through global semantic information.
4. The steel defect detection method according to claim 1, characterized in that, The step of introducing an SPD-Conv module into the YOLOv8 feature extraction backbone network and embedding an EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network includes: Replace the CBS module in the YOLOv8 feature extraction backbone network and introduce an SPD-Conv neural network building block; Design a C2f_EMA module with an attention mechanism and introduce the C2f_EMA module into the YOLOv8 feature extraction backbone network to replace the C2f module in the YOLOv8 feature extraction backbone network; Embed the C2f_EMA module between the Split division layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network.
5. The steel defect detection method according to claim 1, characterized in that, The step of fusing the Neck part in the feature extraction model and adding a lightweight cross-scale feature fusion module includes: Add a lightweight cross-scale feature fusion module to the Neck part, and fuse the Neck part with the added lightweight cross-scale feature fusion module into the feature extraction model; Add convolution operations before and after each upsampling in the feature extraction model, and connect the upsampling with the splicing operation.
6. The steel defect detection method according to claim 1, wherein, The step of dividing the processed surface defect dataset into a training set and a test set includes: Divide the processed surface defect dataset based on a stratified sampling strategy to obtain a training set, a validation set, and a test set; Store the training set, the validation set, and the test set in corresponding folders respectively.
7. The steel defect detection method according to claim 1, characterized in that, After the step of detecting the test set based on the weight file and the steel defect detection model to obtain a detection result, the method further includes: Evaluate the detection result based on an object detection model and evaluation metrics, where the evaluation metrics include accuracy, recall, and mean average precision; The calculation expression of the accuracy is: ; In the formula, represents the accuracy rate, represents the number of samples correctly classified as positive by the steel defect detection model, represents the number of samples misclassified as positive but actually negative by the steel defect detection model; The calculation expression of the recall is: ; In the formula, represents the recall rate, represents the number of samples that are misclassified as negative by the steel defect detection model but are actually positive. The calculation expression of the mean average precision is: ; In the formula, represents the average accuracy value, represents the total number of categories, represents the average precision of the i-th category, represents the average precision.
8. A steel defect detection system, characterized in that, The system includes: A collection module for collecting a steel surface defect dataset and sequentially sorting and converting the surface defect dataset; A processing module for processing the sorted and converted surface defect dataset based on ContextGuided convolution and PConv convolution in sequence to obtain a processed surface defect dataset; An introduction module for introducing an SPD-Conv module into the YOLOv8 feature extraction backbone network and embedding an EMA module between the Split layer and the Bottleneck neck layer in the YOLOv8 feature extraction backbone network to construct a feature extraction model; A fusion module for fusing the Neck part in the feature extraction model, adding a lightweight cross-scale feature fusion module, and introducing an Inner-IoU loss function into the feature extraction model to obtain an updated feature extraction model; A training module for dividing the processed surface defect dataset into a training set and a test set, and training the updated feature extraction model based on the training set to obtain a steel defect detection model and a weight file; A detection module for detecting the test set based on the weight file and the steel defect detection model to obtain a detection result.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steel defect detection method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steel defect detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Compressed convolutional neural network-oriented parallel convolution operation method and apparatus
CN106951395A
Steel plate surface defect detection method based on improved YOLOv8
CN118279272A
Improved YOLOv8-based steel surface defect target detection method
CN118823299A
PCB defect detection method based on small target enhanced feature pyramid
CN119785175A
Steel surface defect detection method and device based on YOLO11n improvement
CN120088240A