Steel plate surface defect detection method based on improved YOLOv8
By improving the backbone network and feature fusion structure of the YOLOv8 network, combining multi-scale feature fusion and loss function optimization, the problem of insufficient detection accuracy of small objects is solved, and efficient and accurate detection of steel plate surface defects is achieved.
Patent Information
- Application Number
- CN202510450535.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing YOLOv8 lacks detection accuracy and accuracy when dealing with small target objects in complex backgrounds, resulting in frequent missed detection and missed detection, limiting its wide application in the industrial field.
Using the improved YOLOv8 network, multi-scale feature fusion is carried out by building the backbone network Backbone, feature fusion network Neck and detection head head, combining FPN, PAN structure, C2f-ContextFusion module and SPPF-DSLSKA module, and designing improved loss functions, including binary cross entropy loss and bounding box loss, to improve the accuracy of small object detection.
It significantly improves the detection accuracy and accuracy of small targets, reduces missed and missed detection phenomena, improves the model's detection ability of targets of different sizes, and ensures the reliability and efficiency of detection.
Smart Images

Figure CN120374543A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of stainless steel plate defect recognition, and particularly relates to a method for detecting surface defects of steel plates based on improved YOLOv8. Background Art
[0002] Due to its excellent corrosion resistance, high strength, and good mechanical properties, medium and thick stainless steel plates have been widely used in industrial fields such as construction, automotive, shipbuilding, and machinery manufacturing. Such materials are not only used for structural components that bear large loads but also can be used stably for a long time in harsh environments, with extremely high practical value. However, during the production and processing of medium and thick stainless steel plates, various defects may appear on their surfaces, such as scratches, pits, cracks, scale, impurities, etc. These surface defects not only affect the appearance quality of the products but may also cause problems such as stress concentration and local fatigue. In severe cases, they may even weaken the mechanical properties and service life of the materials, thereby reducing their corrosion resistance and strength, and ultimately affecting the overall quality and reliability of the products.
[0003] To ensure the high quality of products, it is particularly important to identify and detect these surface defects in a timely and accurate manner. Traditional surface defect detection methods mainly rely on manual visual inspection. Although manual inspection has a certain degree of flexibility, it has high labor intensity, low efficiency, and the detection results are easily affected by factors such as the fatigue level, experience level, and subjective judgment of the inspectors, making it difficult to guarantee the accuracy and consistency of the detection results. With the rapid development of industrial automation and intelligence, automated detection technology has gradually become the mainstream means to replace manual inspection. Especially driven by deep learning technology, defect detection technology based on image processing has been widely used.
[0004] Currently, object detection models in deep learning have shown significant advantages in the field of industrial surface defect detection. Among them, the YOLO series of single-stage object detection networks occupy an important position in this field due to their high efficiency and real-time performance. Compared with traditional multi-stage detection methods, the YOLO network can input the entire image into the neural network at once and quickly complete defect detection in an end-to-end manner, with extremely high detection speed and good accuracy. With the continuous iteration and optimization of the YOLO series of networks, their detection accuracy and performance have also been significantly improved, and they are widely used in the automated defect detection of industrial images, greatly improving the detection efficiency and accuracy.
[0005] In the prior art, as the latest version of the YOLO series, YOLOv8 adopts a more lightweight and efficient backbone network, which can further improve the detection speed while maintaining high accuracy. In terms of data preprocessing, YOLOv8 continues to use the mosaic augmentation technique and introduces more modern data augmentation methods to improve the generalization ability of the model. Similar to YOLOv5, YOLOv8 adopts an enhanced CSP structure, combined with multi-scale feature extraction, to further improve the model's detection ability for small targets and complex scenes. In addition, the bottleneck network of YOLOv8 adopts a new feature fusion strategy, which can better fuse feature maps from different levels, making the model capture fine-grained information more accurately. YOLOv8 still generates multi-scale feature maps and predicts class, object confidence, and bounding box coordinate information through a multi-task head. Its loss function and optimization algorithm have also been improved. For example, a new IoU variant is introduced to further improve the accuracy of the predicted bounding box and the optimization efficiency. The entire model uses SiLU (Sigmoid-weighted Linear Unit) as the main activation function and combines a more efficient attention mechanism in some layers to further enhance the selective feature extraction ability. As an end-to-end model, YOLOv8 has extremely strong real-time performance and accuracy and is widely applicable to various object detection tasks.
[0006] However, although YOLOv8 performs excellently in the field of object detection, existing detection networks still have certain limitations in dealing with small target objects in complex backgrounds. Due to the lack of feature information of small targets, it is difficult for existing models to fully extract their features, resulting in insufficient detection accuracy and precision. In addition, as the network depth increases, the feature information of small targets may gradually weaken or even disappear, leading to frequent missed detections and false detections of small target defects. This situation not only reduces the reliability of detection but also significantly affects production efficiency and cost, limiting its wide application in the industrial field.
[0007] Therefore, there is a need for a steel plate surface defect detection method based on improved YOLOv8 that can effectively improve the detection accuracy and precision of small targets and reduce the missed detections and false detections of small target defects. Summary of the Invention
[0008] The purpose of the present invention is to provide a steel plate surface defect detection method based on improved YOLOv8, which can effectively improve the detection accuracy and precision of small targets and reduce the missed detections and false detections of small target defects.
[0009] To achieve the above purpose, the technical solution adopted by the present invention is:
[0010] A steel plate surface defect detection method based on improved YOLOv8, comprising the following steps:
[0011] Step S1: Data acquisition and division: Obtain the image data of the steel plate surface defects and divide them into a training set and a test set;
[0012] Step S2: Data augmentation and preprocessing: Perform one or more of random scaling, cropping, flipping, rotation, color jittering, brightness and contrast adjustment on the image data of the steel plate surface defects in the training set to obtain a sample training set;
[0013] Step S3: Build a defect detection model based on the improved YOLOv8 network. After initializing the epoch, learning rate and model weights, input the data of the samples to be detected in the sample training set into the defect detection model to obtain the prediction results;
[0014] Step S4: Design a loss function and decode the prediction results;
[0015] Step S5: Repeatedly input the data of the samples to be detected in the sample training set into the defect detection model for training until the number of training times reaches the preset epoch, output the final defect detection model, and then input the data of the samples to be detected in the test set into the final defect detection model to obtain the detection results of the steel plate surface defects.
[0016] A further improvement of the technical solution of the present invention lies in: Step S3 includes the following steps:
[0017] Step S301: Build a backbone network Backbone, and input the data of the samples to be detected in the sample training set into the backbone network Backbone for feature extraction, so as to send multi-level feature maps into the feature fusion network Neck for feature fusion;
[0018] Step S302: Build a feature fusion network Neck. The feature fusion network Neck adopts the FPN plus PAN structure to effectively fuse shallow and deep features and send them into the detection head Head to improve the detection ability for small targets;
[0019] Step S303: Build a detection head Head. The detection head Head includes a regression branch, a confidence branch and a classification branch, which are used to predict the coordinate information of the target box, the confidence of foreground and background, and the category information of the target respectively. The prediction results output by each branch include the coordinate information of the prediction box, whether there is an object and the object type. Stack the output results of the three branches to obtain the final prediction results, including the category, position and confidence of the target.
[0020] A further improvement of the technical solution of the present invention lies in: in step S301, the backbone network Backbone includes a plurality of sequentially connected convolutional modules CBS, an improved C2f-ContextFusion module, and an SPPF-DSLSKA module; according to the propagation direction, Backbone extracts features through the convolutional module CBS and the improved C2f-ContextFusion module, and Backbone uses the SPPF-DSLSKA module to capture multi-scale context information in the image to improve the model's detection ability for targets of different sizes.
[0021] A further improvement of the technical solution of the present invention lies in: the convolutional module CBS includes a combination module of a convolutional layer with a 3*3 convolution kernel and a stride of 2, batch normalization, and the activation function SiLU, which is used for initial feature extraction.
[0022] A further improvement of the technical solution of the present invention lies in: the Bottleneck structure in the improved C2f-ContextFusion module is optimized into a CF structure. The CF module in the C2f-ContextFusion module includes: a 3×3 convolutional path and a dilated convolutional path connected in parallel along the forward propagation direction. After the two paths, a batch normalization unit and an activation function unit are respectively set. The activation function unit includes a ReLU activation function and a Sigmoid activation function, which are used to generate a weight vector. The feature map generates spatial feature weights through a spatial attention mechanism. The weight vector is used to guide feature fusion. The concatenated feature map is weighted by element-wise multiplication and output an effective feature map through a residual connection.
[0023] A further improvement of the technical solution of the present invention lies in: the SPPF module introduces a local adaptive convolutional attention mechanism DSLSKA module on the basis of the spatial pyramid pooling structure; in the SPPF module, the input feature map is first preprocessed by a convolutional block, which contains a 1x1 convolutional kernel, and its output channel number is the same as that of the input feature map; subsequently, the feature map is passed to three MaxPool2d layers. These pooling layers downsample the feature map in a serial manner and extract the main features. Each pooling operation is performed on different regions of the feature map to capture information of different scales; through the Concat splicing operation, the output feature maps of the three MaxPool2d layers are spliced in the channel dimension.
[0024] A further improvement of the technical solution of the present invention lies in: The DSLSKA module includes a local feature extractor, an overall feature comparator, and multi-channel interaction; among them, the DW-DSConv module used in the local feature extractor splits the depth convolution into two cascaded one-dimensional convolutions, and the convolution kernel is 2d-1, where d is the dilation rate; the DW-D-DSConv module used in the overall feature comparator is a dilated convolution with a stride of k / d. The dilated convolution introduces holes in the convolution kernel to expand its effective receptive field and reduces the number of parameters while keeping the size of the convolution layer unchanged; the multi-channel interaction uses 1x1 convolution to operate on all channels of the pixels to increase the interaction between channels.
[0025] A further improvement of the technical solution of the present invention lies in: In step S302, the process of feature fusion is as follows: First, the high-level enhanced feature map input by the backbone network Backbone is upsampled through FPN and fused with the low-level feature map by element-wise addition to construct a top-down feature pyramid. Second, the upsampled high-level feature map is decomposed into three branches with different dilation rates to capture multi-scale context information. Then, a Concat splicing operation is performed on each branch and sent to the C2f-ContextFusion module for processing. Then, the PAN structure is used to increase the bottom-up path aggregation to further fuse the features, and finally, the feature maps of the three branches are output to the detection head Head.
[0026] A further improvement of the technical solution of the present invention lies in: Step S4 includes the following steps:
[0027] Step S401: Design a classification loss function: The classification loss function is the Binary CrossEntropy with Logits Loss, and its calculation formula is:
[0028] Represent the cost matrix as L cls = -∑[y·log(σ(x))+(1 - y)·log(1 - σ(x))],
[0029] In the formula, y represents the true label, x represents the unactivated predicted value output by the model, and σ(x) represents the Sigmoid function;
[0030] Step S402: Design a bounding box loss function: The bounding box loss function includes the center point coordinate loss of the bounding box, the width and height loss of the bounding box, and the confidence loss of the bounding box.
[0031] A further improvement of the technical solution of the present invention lies in: In step S402, the center point coordinate loss of the bounding box is: Its function is to predict the center point coordinates (x i , yi ) The mean square error between the coordinates (x t , y t ) of the center point of the true bounding box;
[0032] The width and height loss of the bounding box is: Its function is to calculate the width and height (w i , h i ) of the predicted bounding box and the width and height (w t , h t ) of the true bounding box, and take the square root of the width and height to reduce the weight of large bounding boxes;
[0033] The confidence loss of the bounding box is:
[0034]
[0035] Its function is to calculate the binary cross-entropy loss between the confidence C i of the predicted bounding box for positive and negative samples respectively and the confidence C t of the true bounding box, so that the confidence of the predicted bounding box is close to that of the true bounding box.
[0036] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:
[0037] The steel plate surface defect detection method based on the improved YOLOv8 of the present invention can effectively improve the detection accuracy and precision of small targets, and reduce the missed detection and false detection of small target defects.
[0038] In the YOLOv8s network adopted by the present invention, the feature fusion module efficiently fuses multi-scale features through the path aggregation network, significantly improving the detection ability of the model for targets of different sizes. Specifically, the feature fusion network Neck not only utilizes the top-down path from high-level to low-level to capture high-level features with rich semantic information, but also introduces the bottom-up path, enabling better transmission and retention of spatial details in low-level features. Through the aggregation of these two-way paths, the YOLOv8s network can effectively fuse features at different scales, ensuring that the model can maintain high-precision recognition ability for small targets while detecting large targets. Description of the Drawings
[0039] Figure 1 is the flowchart of the steel plate surface defect detection method of the present invention;
[0040] Figure 2 is the overall structural schematic diagram of the defect detection model in the present invention;
[0041] Figure 3 is the compositional structural schematic diagram of the CF module adopted by the present invention;
[0042] Figure 4 It is a schematic diagram of the composition structure of the SPPF-DSLSKA module adopted by the present invention;
[0043] Figure 5 It is a schematic diagram of the composition structure of the DSLSKA adaptive convolutional attention adopted by the present invention;
[0044] Figure 6 It is a convolutional schematic diagram of different ways of existing DSLSKA attention;
[0045] Figure 7 It is the original model structure diagram based on YOLOv8 provided by the present invention. Specific embodiments
[0046] The present invention will be further described in detail below in conjunction with embodiments:
[0047] As Figure 1 shown, the present invention provides a method for detecting steel plate surface defects based on improved YOLOv8, including the following steps:
[0048] Step S1: Data collection and division: Obtain steel plate surface defect picture data and divide it into a training set and a test set; This data comes from the actual production line of a steel factory and is collected using a Baumer high-resolution industrial camera (model VCXG.2-51C, 5 million pixels). The specific method is to continuously take pictures while the steel plate is conveyed by a roller. To improve the image quality and eliminate the influence of light, preprocessing such as cropping and normalization is performed on the collected images, which provides a high-quality data basis for the training of the subsequent defect detection model;
[0049] Step S2: Data augmentation and preprocessing: Perform one or more of random scaling, cropping, flipping, rotating, color jittering, brightness and contrast adjustment on the steel plate surface defect picture data in the training set to obtain a sample training set;
[0050] Step S3: As Figure 2 shown, construct a defect detection model based on the improved YOLOv8 network. After initializing the epoch, learning rate, and model weights, input the data of the samples to be detected in the sample training set into the defect detection model to obtain the prediction results, as Figure 7As shown, the original YOLOv8 network structure consists of three parts: the backbone network Backbone, the feature fusion network Neck, and the detection head Head. Among them, the Backbone uses the CSPDarknet53 structure to extract multi-level features, the Neck realizes the feature pyramid fusion through the FPN+PAN structure, and the Head predicts the classification and regression results respectively through the decoupled detection head. The present invention has made the following improvements on this basis, specifically including the following steps:
[0051] Step S301: Construct the backbone network Backbone, and input the sample data to be detected in the sample training set into the backbone network Backbone for feature extraction, so as to send the multi-level feature maps into the feature fusion network Neck for feature fusion;
[0052] The backbone network Backbone of LC-YOLO consists of standard YOLOv8s modules. Among them, the SPPF module of YOLOv8s is added with the large separable kernel DSLSKA attention mechanism. The backbone network Backbone includes a number of convolutional modules CBS, an improved C2f-ContextFusion module, and an SPPF-DSLSKA module connected in series in sequence; according to the propagation direction, the Backbone extracts features through the convolutional module CBS and the improved C2f-ContextFusion module, and the Backbone uses the SPPF-DSLSKA module to capture the multi-scale context information in the image to improve the model's detection ability for targets of different sizes;
[0053] Among them, the convolutional module CBS includes a combination module of a convolutional layer with a convolution kernel of 3*3 and a stride of 2, batch normalization, and the activation function SiLU, which is used for initial feature extraction;
[0054] The Bottleneck structure in the improved C2f-ContextFusion module is optimized to the CF structure to enhance the feature extraction ability. Specifically, the CF module in the C2f-ContextFusion module includes: a 3×3 convolutional path and a dilated convolutional path connected in parallel along the forward propagation direction. After the two paths, a batch normalization unit and an activation function unit are respectively set. The activation function unit includes a ReLU activation function and a Sigmoid activation function, which are used to generate a weight vector. The feature map generates the spatial feature weight through the spatial attention mechanism. The weight vector is used to guide the feature fusion. The spliced feature map is weighted by element-wise multiplication and output the effective feature map through the residual connection;
[0055] More specifically, as Figure 3As shown in the figure, first, the CF module halves the number of input channels of the input feature map through a 1×1 convolution. Subsequently, a copy is made and divided into two paths. One path extracts features through a conventional 3×3 convolution, focusing on the details and local information of the input image. The other path uses dilated convolution to capture context information over a larger range. Dilated convolution expands the receptive field by inserting holes between the convolution kernels, thereby obtaining more context information while keeping the computational cost unchanged.
[0056] Then, the CF module concatenates the local features and the global features to form a richer feature map, and then adopts a spatial attention mechanism for spatial weight allocation. The concatenated feature map is processed by batch normalization and the ReLU activation function to enhance the stability and non-linear ability of feature representation.
[0057] Finally, the CF module uses a global pooling layer to extract global features. Global pooling performs average pooling on the entire feature map to obtain a global feature vector. The global feature vector passes through two fully connected layers, using the ReLU and Sigmoid activation functions respectively, to generate a weight vector. This weight vector is used to guide feature fusion and weights the concatenated feature map by element-wise multiplication.
[0058] The SPPF-DSLSKA module introduces a local adaptive convolutional attention mechanism DSLSKA module on the basis of the spatial pyramid pooling structure to enhance the extraction of key information in multi-scale feature maps, effectively capture targets of different sizes in the image, and improve the detection accuracy.
[0059] As Figure 4 shown in the figure, in the SPPF-DSLSKA module, the input feature map is first preprocessed by a convolutional block. This convolutional block contains a 1×1 convolutional kernel, and its output channels are the same as those of the input feature map. Subsequently, the feature map is passed to three MaxPool2d layers. These pooling layers downsample the feature map in a serial manner, effectively reducing the computational complexity of the network and extracting the main features. Each pooling operation is performed on different regions of the feature map to capture information at different scales. Through the Concat concatenation operation, the output feature maps of the three MaxPool2d layers are concatenated together in the channel dimension.
[0060] By introducing the DSLSKA attention module, the receptive field is expanded with a large convolutional kernel and global information is effectively obtained, solving the problem of information loss caused by downsampling, thereby significantly improving the detection ability of the model. As Figure 5As shown, the DSLSKA module includes a local feature extractor, a global feature comparator, and multi-channel interaction. Among them, the DW-DSConv module used in the local feature extractor splits the depth convolution into two cascaded one-dimensional convolutions, with a convolution kernel of 2d-1, where d is the dilation rate. The essence of the DW-D-DSConv module used in the global feature comparator is dilated convolution with a stride of k / d. Dilated convolution introduces holes in the convolution kernel to expand its effective receptive field, reducing the number of parameters while keeping the size of the convolutional layer unchanged. The multi-channel interaction uses 1x1 convolution to operate on all channels of the pixels, increasing the interaction between channels.
[0061] As Figure 6 shown, the DSLSKA module further enhances the feature extraction ability through convolutional operations in the horizontal and vertical directions. Among them, (a) and (b) respectively show the decomposition of the one-dimensional convolution kernels in the horizontal and vertical directions. This decomposition method effectively reduces the computational complexity while retaining the integrity of the spatial information. (c) and (d) reflect the synergistic effect of the multi-scale convolution kernels. Through the convolution combinations in different directions, the model can capture local details and global context information more comprehensively. This design not only optimizes the expansion efficiency of the receptive field but also significantly improves the accuracy of small target detection through directional feature fusion, especially showing stronger robustness in complex backgrounds.
[0062] Step S302: Construct a feature fusion network Neck. The feature fusion network Neck adopts the FPN + PAN structure to effectively fuse shallow and deep features and send them to the detection head Head to improve the detection ability for small targets. The process of feature fusion is as follows: First, the high-level enhanced feature map input by the backbone network Backbone is upsampled through FPN and fused with the low-level feature map by element-wise addition to construct a top-down feature pyramid. Second, the upsampled high-level feature map is decomposed into three branches with different dilation rates to capture multi-scale context information. Then, a Concat splicing operation is performed on each branch and sent to the C2f-ContextFusion module for processing. Then, the PAN structure is used to add a bottom-up path aggregation to further fuse the features. Finally, the feature maps of the three branches are output to the detection head Head.
[0063] Step S303: Construct a detection head Head. The detection head Head includes a regression branch, a confidence branch, and a classification branch, which are used to predict the coordinate information of the target box, the confidence of the foreground and background, and the category information of the target respectively. The prediction results output by each branch include the coordinate information of the prediction box, whether there is an object (foreground or background), and the object category. The output results of the three branches are stacked to obtain the final prediction results, including the category, location, and confidence of the target.
[0064] Step S4: Design a loss function and decode the prediction results, which specifically includes the following steps:
[0065] Step S401: Design a classification loss function: The classification loss function is Binary CrossEntropy with Logits Loss, and its calculation formula is:
[0066] Represent the cost matrix as L cls = -∑[y·log(σ(x))+(1 - y)·log(1 - σ(x))],
[0067] In the formula, y represents the true label, x represents the unactivated predicted value output by the model, and σ(x) represents the Sigmoid function;
[0068] Step S402: Design a bounding box loss function: The bounding box loss function includes the center point coordinate loss of the bounding box, the width and height loss of the bounding box, and the confidence loss of the bounding box (for positive and negative samples);
[0069] Among them, the center point coordinate loss of the bounding box is: Its function is to predict the mean square error between the center point coordinates (x i , y i ) of the predicted bounding box and the center point coordinates (x t , y t ) of the true bounding box;
[0070] The width and height loss of the bounding box is: Its function is to calculate the mean square error between the width and height (w i , h i ) of the predicted bounding box and the width and height (w t , h t ) of the true bounding box and take the square root of the width and height to reduce the weight of large bounding boxes;
[0071] The confidence loss of the bounding box is:
[0072]
[0073] Its function is to calculate the binary cross entropy loss between the confidence C i of the predicted bounding box for positive and negative samples respectively and the confidence C t of the true bounding box, ensuring that the confidence of the predicted bounding box is as close as possible to the confidence of the true bounding box;
[0074] Step S5: Repeatedly input the sample data to be detected in the sample training set into the defect detection model for training until the number of training times reaches the preset epoch, output the final defect detection model, and then input the sample data to be detected in the test set into the final defect detection model to obtain the steel plate surface defect detection result.
[0075] It can be understood that the present invention is described by means of some embodiments. Those skilled in the art will know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A method for detecting defects on the surface of steel plates based on improved YOLOv8, characterized in that It includes the following steps: Step S1: Data collection and division: Obtain the image data of the steel plate surface defects and divide them into a training set and a test set; Step S2: Data augmentation and preprocessing: Perform one or more of random scaling, cropping, flipping, rotating, color jittering, brightness and contrast adjustment on the image data of the steel plate surface defects in the training set to obtain a sample training set; Step S3: Build a defect detection model based on the improved YOLOv8 network. After initializing the epoch, learning rate and model weights, input the sample data to be detected in the sample training set into the defect detection model to obtain the prediction results; Step S4: Design a loss function and decode the prediction results; Step S5: Repeatedly input the sample data to be detected in the sample training set into the defect detection model for training until the number of training times reaches the preset epoch, output the final defect detection model, and then input the sample data to be detected in the test set into the final defect detection model to obtain the steel plate surface defect detection results.
2. The method for detecting steel plate surface defects based on improved YOLOv8 according to claim 1, wherein: Step S3 includes the following steps: Step S301: Build a backbone network Backbone, and input the sample data to be detected in the sample training set into the backbone network Backbone for feature extraction, so as to send the multi-level feature maps into the feature fusion network Neck for feature fusion; Step S302: Build a feature fusion network Neck. The feature fusion network Neck adopts the FPN + PAN structure to effectively fuse the shallow and deep features and send them to the detection head Head to improve the detection ability for small targets; Step S303: Build a detection head Head. The detection head Head includes a regression branch, a confidence branch and a classification branch, which are used to predict the coordinate information of the target box, the confidence of the foreground and background, and the category information of the target respectively. The prediction results output by each branch include the coordinate information of the prediction box, whether there is an object and the object type. Stack the output results of the three branches to obtain the final prediction results, including the category, location and confidence of the target.
3. The method for detecting steel plate surface defects based on improved YOLOv8 according to claim 2, wherein: In step S301, the backbone network Backbone includes a number of convolutional modules CBS, an improved C2f-ContextFusion module and an SPPF-DSLSKA module connected in series in sequence; according to the propagation direction, Backbone extracts features through the convolutional module CBS and the improved C2f-ContextFusion module, and Backbone uses the SPPF-DSLSKA module to capture the multi-scale context information in the image to improve the detection ability of the model for targets of different sizes.
4. The method for detecting steel plate surface defects based on improved YOLOv8 according to claim 3, characterized in that: The convolutional module CBS includes a combined module of a convolutional layer with a convolution kernel of 3*3 and a stride of 2, batch normalization and an activation function SiLU, which is used for initial feature extraction.
5. A method for detecting steel plate surface defects based on improved YOLOv8 according to claim 4, characterized in that: The Bottleneck structure in the improved C2f-ContextFusion module is optimized to the CF structure. The CF module in the C2f-ContextFusion module includes: a 3×3 convolution path and a dilated convolution path in parallel along the forward propagation direction. After the two paths, a batch normalization unit and an activation function unit are respectively set. The activation function unit includes a ReLU activation function and a Sigmoid activation function, which are used to generate a weight vector. The feature map generates spatial feature weights through a spatial attention mechanism. The weight vector is used to guide feature fusion. The concatenated feature map is weighted by element-wise multiplication and output the effective feature map through a residual connection.
6. The method for detecting steel plate surface defects based on improved YOLOv8 according to claim 5, wherein: The SPPF module introduces a local adaptive convolutional attention mechanism DSLSKA module on the basis of the spatial pyramid pooling structure; in the SPPF-DSLSKA module, the input feature map is first preprocessed by a convolutional block, which contains a convolutional kernel with a size of 1x1, and its output channel number is the same as that of the input feature map; subsequently, the feature map is passed to three MaxPool2d layers, and these pooling layers downsample the feature map in a serial manner and extract the main features. Each pooling operation is performed on different regions of the feature map to capture information at different scales; Through the Concat splicing operation, the output feature maps of the three MaxPool2d layers are spliced in the channel dimension.
7. A method for detecting defects on the surface of steel plates based on improved YOLOv8 according to claim 6, characterized in that: The DSLSKA module includes a local feature extractor, an overall feature comparator, and multi-channel interaction; among them, the DW-DSConv module used by the local feature extractor splits the depth convolution into two cascaded one-dimensional convolutions, and the convolutional kernel is 2d-1, where d is the dilation rate; the DW-D-DSConv module used by the overall feature comparator is a dilated convolution with a stride of k / d. The dilated convolution introduces holes in the convolutional kernel to expand its effective receptive field and reduces the number of parameters while keeping the size of the convolutional layer unchanged; the multi-channel interaction uses a 1x1 convolution to operate on all channels of the pixels to increase the interaction between channels.
8. A method for detecting steel plate surface defects based on improved YOLOv8 according to claim 7, characterized in that: In step S302, the process of feature fusion is as follows: First, the high-level enhanced feature map input by the backbone network Backbone is upsampled by FPN and fused with the low-level feature map by element-wise addition to construct a top-down feature pyramid. Second, the upsampled high-level feature map is decomposed into three branches with different dilation rates to capture multi-scale context information. Then, a Concat splicing operation is performed on each branch and sent to the C2f-ContextFusion module for processing. Then, the PAN structure is used to increase the bottom-up path aggregation to further fuse the features. Finally, the feature maps of the three branches are output to the detection head Head.
9. A method for detecting steel plate surface defects based on improved YOLOv8 according to claim 8, characterized in that: The step S4 includes the following steps: Step S401: Design a classification loss function: The classification loss function is the Binary Cross Entropy with Logits Loss, and its calculation formula is as follows: Represent the cost matrix as L cls = -∑[y·log(σ(x))+(1 - y)·log(1 - σ(x))], In the formula, y represents the true label, x represents the unactivated predicted value output by the model, and σ(x) represents the Sigmoid function; Step S402: Design a bounding box loss function: The bounding box loss function includes the center point coordinate loss of the bounding box, the width and height loss of the bounding box, and the confidence loss of the bounding box.
10. A method for detecting steel plate surface defects based on improved YOLOv8 according to claim 9, characterized in that: In step S402, the loss of the center point coordinates of the bounding box is as follows: Its function is to predict the mean square error between the center point coordinates (x i , y i ) of the bounding box and the center point coordinates (x t , y t ) of the true bounding box; The width and height loss of the bounding box is: Its function is to calculate the mean squared error between the width and height (w i , h i ) of the predicted bounding box and the width and height (w t , h t ) of the ground truth bounding box and take the square root of the width and height to reduce the weight of large bounding boxes; The confidence loss of the bounding box is: Its function is to calculate the confidence C of the predicted bounding boxes for positive and negative samples respectively i and the confidence C of the ground truth bounding boxes t between the binary cross-entropy losses, making the confidence of the predicted bounding boxes close to that of the ground truth bounding boxes.
Citation Information
Cited By
Surface defect detection method and equipment based on multi-scale adaptive guidance and medium
CN120672758A
Surface defect detection method and device based on multi-scale adaptive guidance and medium
CN120672758B
Industrial product surface defect detection method based on feature coupling
CN120766047A
Flood discharge detection model training method and device, equipment and medium
CN121095696A
Surface defect detection method and equipment based on DINOv3 model and medium
CN121190466A