Improved YOLOv8n-based steel surface defect detection model construction method and system

By optimizing YOLOv8n's backbone network and feature fusion module, combining the EffectiveSE attention mechanism and CAFM module, and using the SlideLoss loss function, the problem of difficult to balance detection accuracy and efficiency in steel surface defect detection is solved, and efficient and accurate defect detection effect is achieved.

CN120070369APending Publication Date: 2025-05-30GUANGZHOU INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510142624.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has the problem of difficult to balance detection accuracy and efficiency in steel surface defect detection. Too many model parameters lead to large calculations and difficult to quickly deploy on edge terminal devices. At the same time, there are some shortcomings in detection accuracy and speed.

Method used

By optimizing the C2f architecture of the YOLOv8n backbone network as a C2f-ESEM module, the EffectiveSE attention mechanism is introduced, and the convolution and attention fusion module CAFM is introduced in the feature fusion stage, and SlideLoss is used as the loss function to improve detection accuracy and efficiency.

Benefits of technology

The accuracy and efficiency of defect detection are significantly improved, and the model shows stronger adaptability and robustness in complex and changing environments, and enables rapid deployment on edge end devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070369A_ABST
    Figure CN120070369A_ABST
Patent Text Reader

Abstract

The invention relates to a steel surface defect detection model construction method and system based on improved YOLOv8n, and the method comprises the steps: replacing a convolution assembly in an original C2f module with an MBConv module in a backbone network, introducing an EffectiveSE attention mechanism, and enhancing feature representation; a convolution and attention fusion module CAFM is introduced in the feature fusion stage, and effective fusion of multi-level features is achieved; the SlideLoss is adopted as a loss function, and the loss weight is dynamically adjusted through an intersection-to-parallel ratio IoU index, so that the problem of class imbalance is relieved. Compared with the prior art, the method has the advantages that the accuracy and efficiency of defect detection are improved, the cross-scale target capturing capability of the model is enhanced, the problem of class imbalance is effectively relieved, higher adaptability and robustness are shown, and the method has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly to a method and system for constructing a steel surface defect detection model based on improved YOLOv8n. Background Art

[0002] As the core pillar of China's economic growth, steel plays a crucial role in the manufacturing of various industrial products. However, in the production and processing process of steel, due to environmental conditions and processing technology limitations, various defects often occur on the steel surface. These defects not only damage the external beauty and durability of the products but also may pose potential safety risks in specific application fields. Therefore, steel surface defect detection has become an important link to ensure product quality and safety.

[0003] In recent years, with the rise of deep learning technology, object detection algorithms based on image recognition have shown great potential in steel surface defect detection due to their excellent anti-interference ability and high detection efficiency. Surface defect detection algorithms under the deep learning framework are mainly divided into two categories: two-stage detection algorithms and one-stage detection algorithms. Two-stage algorithms, such as the RCNN series (including Fast R-CNN, Faster R-CNN), achieve high-precision detection by first generating candidate regions and then performing feature extraction and target regression. However, due to their large parameter scale and slow processing speed, their deployment and application on edge devices are limited. In contrast, one-stage detection algorithms, such as SSD and the YOLO series, significantly improve the detection speed by integrating classification and localization tasks into a single regression framework. Although they have obvious advantages in terms of speed, the detection accuracy of these algorithms is still slightly insufficient. The training and testing of deep learning models require a large amount of computing resources, while the computing power of terminal devices is limited, and excessive parameters are not conducive to the deployment of the model on edge terminal devices. In practical application scenarios, the real-time performance of defect detection has also become an important indicator for evaluating network performance, which brings certain difficulties and challenges to the practical application of steel surface defect detection.

[0004] To overcome the drawbacks of traditional steel surface defect detection technologies, various deep learning-based detection methods have been adopted in the prior art. For example, there is a method that uses EfficientNet as the backbone feature extraction network to enhance the feature extraction ability of the feature network for steel surface defects, and improves the quality of defect segmentation and the accuracy of detection by designing a weighted bidirectional recursive feature pyramid and a MASK detection head incorporating an attention module. Another method improves the detection efficiency and accuracy by using the improved YOLO algorithm, replacing the module in the original backbone network with the Res2Net module, designing a new feature fusion module, and adopting the structure of a decoupled head to replace the original detection head. There is also a method that realizes real-time detection of steel surface defects through multi-scale detection and the improved Yolo-v5 algorithm. In addition, there is a method that improves based on the YOLOv5 algorithm, introduces the BiFormer attention mechanism and improves the CSP module to enhance the detection ability for tiny targets.

[0005] Although the prior art has made certain progress in steel surface defect detection, there are still some problems. On the one hand, although some methods have improved the detection accuracy, due to the excessive number of model parameters, the computational load is large, making it difficult to be quickly deployed on edge terminal devices. On the other hand, although some methods have improved the detection speed, there are still deficiencies in the detection accuracy. In addition, for the problem of the uneven scales of steel surface defects, the existing methods may not be fine enough in processing, resulting in limited recognition ability for certain defects. Summary of the Invention

[0006] In view of this, it is necessary to provide a method and system for constructing a steel surface defect detection model based on the improved YOLOv8n to solve the above-mentioned defects of the prior art.

[0007] To solve the above problems, on the one hand, an embodiment of the present invention provides a method for constructing a steel surface defect detection model based on the improved YOLOv8n, including:

[0008] On the basis of the YOLOv8n backbone network, replace the convolutional component in the original C2f module with the MBConv module, and introduce the EffectiveSE attention mechanism;

[0009] In the feature fusion stage of YOLOv8n, introduce the convolutional and attention fusion module CAFM; the convolutional and attention fusion module is used to fuse features at different levels by constructing a local branch and a global branch and using the feature fusion strategy CAFMFusion.

[0010] Use SlideLoss as the loss function of YOLOv8n; the SlideLoss loss function is used to optimize the loss calculation in the deep learning object detection task by incorporating the Slide weighting mechanism and differentiating the weighting of samples according to the Intersection over Union (IoU) metric.

[0011] Preferably, based on the YOLOv8n backbone network, the convolutional component in the original C2f module is replaced with an MBConv module, and the ESE attention mechanism is introduced, including:

[0012] Introduce the MBConv module, which includes a standard convolutional layer, a depthwise separable convolution, a squeeze-and-excitation module, and a Dropout component;

[0013] Replace the SE attention mechanism in the MBConv module with the EffectiveSE attention mechanism to construct the ESE-MBConv module; the EffectiveSE attention mechanism includes a global average pooling layer, a fully connected layer, and a Sigmoid activation function; the global average pooling layer is used to reduce the feature map of each channel to a single scalar value to achieve information condensation; the fully connected layer is used to learn and assign the weight coefficients of each channel; the Sigmoid activation function is used to generate the feature weights reflecting the channel importance;

[0014] Improve the Bottleneck layer through the ESE-MBConv module to obtain the feature extraction component C2f-ESEM.

[0015] Preferably, the mathematical expression form of the EffectiveSE attention mechanism is:

[0016]

[0017] A eSE (X) = σ(W C (F gap (X)))

[0018]

[0019] In the formula, X represents the input feature map, H and W are the height and width of the feature map respectively, X ij is the value of the feature map at the position (i, j), F gap represents the result after global average pooling, W C represents the result after the fully connected layer, σ represents the activation function, A eSE represents the channel attention weight, X refine represents the refined feature map.

[0020] Preferably, the working process of the feature extraction component C2f-ESEM includes:

[0021] Channel expansion is performed through 1×1 convolution, and the Splitting operation is executed to divide the expanded channels into two groups. One group of data streams is processed through n layers of ESE-MBConv modules, and the other group of data is concatenated with the output of the ESE-MBConv module in the channel dimension. Finally, the number of channels is adjusted through 1×1 convolution.

[0022] Preferably, the convolution and attention fusion module CAFM includes a local branch and a global branch;

[0023] In the local branch, the input feature map first adjusts the number of channels through 1×1 convolution, then performs a channel shuffle operation to enhance cross-channel interaction and information integration, and then uses 3×3×3 convolution for feature extraction;

[0024] In the global branch, tensors for query, key, and value are generated through 1×1 convolution and 3×3 depth convolution. An attention map is calculated through the interaction of these tensors, and then feature extraction is performed through 1×1 convolution and added to the input feature to obtain the output of the global branch.

[0025] Preferably, the CAFM module uses the feature fusion strategy CAFMFusion to calculate the weights of the input low-level features and high-level features and then perform weighted merging to obtain the fused features. The expression is:

[0026] F fuse =Conv 1×1 (F low ·w+(1-w)+F low +F high )

[0027] In the formula, F fuse represents the fused feature representation; Conv 1×1 is a 1×1 convolution operation, F low represents the low-level feature, F high represents the high-level feature, w is the weight corresponding to the low-level feature, and 1-w is the weight corresponding to the high-level feature.

[0028] Preferably, the SlideLoss loss function combines the BCEWithLogitsLoss loss function and the Slide weighting function; the expression of the SlideLoss loss function is:

[0029]

[0030] In the formula, SlideLoss is the loss function based on the Slide weighting function, I nis the expression of the loss function, f(x) is the Slide weighting function, and x n is the predicted value, y n is the actual value, w n is the weight coefficient, σ(x n ) is the Sigmoid function, and the Sigmoid function is used to map the predicted value x n to the interval (0, 1).

[0031] In a second aspect, an embodiment of the present invention provides a system for constructing a steel surface defect detection model based on improved ECS-YOLOv8n, including:

[0032] A backbone network improvement module, which is used to replace the convolutional component in the original C2f module with an MBConv module on the basis of the YOLOv8n backbone network, and introduce an EffectiveSE attention mechanism;

[0033] A feature fusion module, which is used to introduce a convolutional and attention fusion module CAFM during the feature fusion stage of YOLOv8n; the convolutional and attention fusion module is used to fuse features at different levels by constructing a local branch and a global branch and using a feature fusion strategy CAFMFusion;

[0034] A loss function optimization module, which is used to use SlideLoss as the loss function of YOLOv8n; the SlideLoss loss function is used to optimize the loss calculation in the deep learning object detection task by incorporating a Slide weighting mechanism and differentially weighting samples according to the intersection over union (IoU) metric.

[0035] In a third aspect, the present invention also provides an electronic device, including a memory and a processor, wherein,

[0036] The memory is used to store programs;

[0037] The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the method for constructing a steel surface defect detection model based on improved YOLOv8n according to the embodiment of the first aspect of the present invention.

[0038] In a fourth aspect, the present invention also provides a computer-readable storage medium for storing computer-readable programs or instructions, and when the programs or instructions are executed by a processor, they can implement the steps in the method for constructing a steel surface defect detection model based on improved YOLOv8n according to the embodiment of the first aspect of the present invention.

[0039] The method and system for constructing a steel surface defect detection model based on improved YOLOv8n provided by the present invention have the following beneficial effects compared with the prior art:

[0040] 1) The present invention optimizes the C2f architecture of the YOLOv8n backbone network into a C2f-ESEM module, introduces the EffectiveSE attention mechanism, enhances the pertinence of feature representation, enables the model to focus more on key features, and thus improves the accuracy of defect detection.

[0041] 2) The present invention introduces a CAFM module in the feature fusion stage of the YOLOv8 framework. The CAFM module realizes the effective fusion of multi-level features by constructing top-down and bottom-up bidirectional paths, and uses the attention mechanism to weight the importance of different features, greatly enhancing the model's cross-scale target capture ability and improving its adaptability to complex application scenarios.

[0042] 3) SlideLoss is used as the loss function of YOLOv8n. SlideLoss effectively alleviates the problem of class imbalance by accurately calculating the overlapping area between the predicted bounding box and the ground truth bounding box and dynamically adjusting the loss weight accordingly, significantly improving the accuracy of target localization, and enabling the model to show stronger adaptability and robustness in complex and changing environments. Description of the Drawings

[0043] Figure 1 It is a flowchart of the method for constructing a steel surface defect detection model based on the improved YOLOv8n provided by the present invention;

[0044] Figure 2 It is a schematic structural diagram of the MBConv module provided by the present invention;

[0045] Figure 3(a) is a schematic structural diagram of the SE module provided by the present invention;

[0046] Figure 3(b) is a schematic structural diagram of the EffectiveSE (abbreviated as ESE) module provided by the present invention;

[0047] Figure 4 It is a schematic structural diagram of the ESE-MBConv module provided by the present invention;

[0048] Figure 5 It is a schematic structural diagram of the feature extraction component C2f-ESEM provided by the present invention;

[0049] Figure 6 It is a schematic structural diagram of the CAFM module provided by the present invention;

[0050] Figure 7 It is a schematic structural diagram of the fusion mechanism CAFMFusion based on the convolutional and attention fusion module provided by the present invention;

[0051] Figure 8 Schematic diagram of the slide weighting function;

[0052] Figure 9 Network architecture diagram of ECS-Yolov8n provided by the present invention;

[0053] Figure 10 Structural block diagram of the electronic device provided by the present invention. Specific embodiments

[0054] The following will specifically describe the preferred embodiments of the present invention in conjunction with the accompanying drawings, in which the accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.

[0055] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0056] Currently, although the prior art has made some progress in the detection of steel surface defects, there are still some problems. On the one hand, although some methods have improved the detection accuracy, due to too many model parameters, the computational complexity is large, making it difficult to be quickly deployed on edge terminal devices. On the other hand, although some methods have improved the detection speed, there are still deficiencies in the detection accuracy. In addition, for the problem of different scales of steel surface defects, the existing methods may not be fine enough in processing, resulting in limited recognition ability for some defects.

[0057] In view of this, the present invention provides a method for constructing a steel surface defect detection model based on improved YOLOv8n, aiming to solve the problem that it is difficult to balance detection accuracy and efficiency in the prior art. By optimizing the C2f architecture of the backbone network, integrating the attention mechanism CAFM module, and adopting the SlideLoss loss function, the present invention significantly improves the accuracy and efficiency of defect detection while keeping the model lightweight. The following will be described and introduced through multiple embodiments.

[0058] Figure 1 Flowchart of the method for constructing a steel surface defect detection model based on improved YOLOv8n provided by the present invention, referring to Figure 1 , the method for constructing a steel surface defect detection model based on improved YOLOv8n provided by the present invention at least includes the following steps:

[0059] Step S1: Based on the YOLOv8n backbone network, replace the convolutional components in the original C2f module with the MBConv module and introduce the EffectiveSE attention mechanism.

[0060] Among them, YOLOv8n is a lightweight version of YOLOv8, where "n" represents "nano". YOLOv8n aims to improve the inference speed of the model and reduce resource consumption while maintaining high detection accuracy by reducing the number of network layers, using more efficient convolutional operations or feature extraction modules, etc. When discussing feature fusion and loss function optimization in the embodiments of the present invention, although it may be described with YOLOv8 as the framework, these improvements are equally applicable to YOLOv8n.

[0061] In a preferred embodiment of the present invention, step S1 may specifically include the following steps S11 to S13:

[0062] S11: Introduce the MBConv module, which includes a standard convolutional layer, a depthwise separable convolution (Depwise Conv), a squeeze-and-excitation (SE) module, and a Dropout component.

[0063] The structural design of the MBConv module is as Figure 2 shown. Referring to Figure 2 , the working process of the MBConv module is as follows:

[0064] First, perform a 1x1 convolution (Conv 1x1 s1) on the input feature map to expand the number of channels. This step can increase the expressive ability of the feature map. Then, perform batch normalization (BatchNormalization, BN) on the feature map after expanding the channels, which helps to stabilize the training process and accelerate the convergence of the model; at the same time, apply the Swish activation function for non-linear transformation, and the Swish activation function helps the model capture non-linear relationships in the input data. Then, use the depthwise separable convolution to process the feature map. The depthwise separable convolution decomposes the standard convolution into a depthwise convolution (Depthwise Conv) and a pointwise convolution (Pointwise Conv). Both are batch-normalized through the BN layer, and a Dropout mechanism is embedded in the pointwise convolution part, which accelerates the convergence process of the network through the strategy of randomly discarding parameters. Referring to Figure 2 , the core part of the MBConv architecture also integrates the SE (Squeeze-and-Excitation) attention mechanism, aiming to improve the sensitivity and accuracy of feature extraction.

[0065] The MBConv module effectively extracts and enhances the expressive power of feature maps through a series of operations, while combining attention mechanisms and regularization techniques to improve the performance and generalization ability of the model.

[0066] In S12, the EffectiveSE attention mechanism is used to replace the SE attention mechanism in the MBConv module to construct the ESE-MBConv module.

[0067] Specifically, as a convolutional neural network architecture designed for image classification tasks, EffectiveSE is a further optimization and upgrade of Squeeze-and-Excitation Networks (abbreviated as SENet). EffectiveSE draws on the core concept of SENet and obtains the "channel attention" mechanism through training to dynamically adjust the weights of each channel, thereby enhancing the overall efficiency of the network. The network architecture of EffectiveSE consists of several key modules. Among them, the "channel attention module" plays a core role, and this module is subdivided into two main components: Squeeze and Excitation. The Squeeze component performs global average pooling operations, which reduces the feature maps of each channel to a single scalar value to achieve information condensation. The Excitation component adopts a multi-layer perceptron (MLP) structure, which is responsible for learning and assigning the weight coefficients of each channel. These weight coefficients are then multiplied by the original feature maps to generate weighted feature maps.

[0068] Figure 3(a) is a schematic structural diagram of the SE module provided by the present invention; Figure 3(b) is a schematic structural diagram of the EffectiveSE (abbreviated as ESE) module provided by the present invention. Referring to Figure 3(b), the EffectiveSE attention mechanism includes a global average pooling layer (Global pooling), a fully connected layer (Fully Connected Layer, abbreviated as FC), a Sigmoid activation function, and a residual connection (Residual); the global average pooling layer is used to reduce the feature maps of each channel to a single scalar value to achieve information condensation; the fully connected layer is used to learn and assign the weight coefficients of each channel; the Sigmoid activation function is used to generate feature weights reflecting the importance of channels. The residual connection (Residual) is used to alleviate the problem of gradient disappearance in the training of deep neural networks and accelerate the convergence of the model.

[0069] The working process of the EffectiveSE (ESE) attention mechanism when processing the input feature map is as follows: First, a global average pooling operation is performed along the height h and width w dimensions to retain the correlation information between channels. The feature map after global average pooling is then processed through a fully connected layer, which maps the features of each channel to a new space to generate a set of recalibrated weights. Then, the output of the fully connected layer is further processed through a Sigmoid activation function (denoted as σ) to generate feature weights reflecting the channel importance, aiming to enhance and enrich the expression diversity of small object features.

[0070] It can be understood that the traditional SE attention mechanism usually includes dimension reduction and dimension increase operations, which may cause loss of channel information during transmission. The EffectiveSE attention mechanism adopted in the embodiments of the present invention effectively avoids this problem through the innovative design of only deploying one fully connected layer (FC). The mathematical expression form of the EffectiveSE attention mechanism is:

[0071]

[0072] A eSE (X) = σ(W C (F gap (X)))

[0073]

[0074] In the formula, X represents the input feature map, H and W are the height and width of the feature map respectively, X ij is the value of the feature map at the position (i, j), F gap represents the result after global average pooling, W C represents the result after passing through the fully connected layer, σ represents the activation function, A eSE represents the channel attention weight, and X refine represents the refined feature map.

[0075] By replacing the SE attention mechanism in the MBConv module with the EffectiveSE attention mechanism, the constructed ESE-MBConv module is as Figure 4 shown.

[0076] S13. Improve the Bottleneck layer through the ESE-MBConv module to obtain the feature extraction component C2f-ESEM.

[0077] In order to improve the feature extraction efficiency of the backbone network and minimize the parameter scale, while maintaining the original network architecture unchanged, the present invention introduces ESE-MBConv to improve the Bottleneck layer, and then constructs a new feature extraction component, named C2f-ESEM. For the structure of the feature extraction component C2f-ESEM, see Figure 5 .

[0078] Refer to Figure 5 . The working process of the feature extraction component C2f-ESEM is as follows: First, perform channel expansion through 1×1 convolution to increase the number of channels of the feature map and provide more feature information for subsequent processing. Then, perform the Splitting (channel splitting) operation to divide the expanded channels into two groups. One group of data streams is processed through n layers of ESE-MBConv modules, and the other group of data is concatenated with the output of the ESE-MBConv module in the channel dimension to combine the original features and the processed features, thereby retaining and fusing feature information at different levels. Finally, perform a fine-tuning of the merged number of channels through an additional 1×1 convolutional layer to obtain the output tensor, realizing the integration and optimization of the fused features, and ensuring that the output features have appropriate number of channels and feature expression capabilities. In the whole process, through strategies such as channel splitting, ESE-MBConv processing, and channel concatenation, flexible configuration and optimization of the number of channels are achieved.

[0079] Step S2, in the feature fusion stage of YOLOv8n, introduce a convolutional and attention fusion module CAFM; the convolutional and attention fusion module is used to fuse features at different levels by constructing a local branch and a global branch and using the feature fusion strategy CAFMFusion.

[0080] By analyzing the network architecture of YOLOv8, it can be observed that although its feature fusion mechanism is designed to maintain both shallow and deep information simultaneously, after passing through multiple layers, the position information of the shallow layer gradually decays and its influence weakens. The feature fusion strategy helps feature retention and effective gradient backpropagation by promoting information flow from shallow to deep. Traditional fusion methods, such as element-wise addition and adaptive weighted mixing, although they can adaptively adjust the fusion weights, face the challenge of inconsistent receptive fields. Specifically, the information contained in shallow and deep features is significantly different, and a single pixel point in the latter often corresponds to a large area in the former. Therefore, simple addition, concatenation, or mixing operations are difficult to overcome the receptive field mismatch before fusion, resulting in insufficient information exchange between feature layers and ultimately affecting the detection performance.

[0081] To address this problem, the present invention introduces a new feature fusion strategy - CAFMFusion, whose core is based on the convolutional and attention fusion module (CAFM).Figure 6 The structural layout of CAFM is shown in detail. The Convolution and Attention Fusion Module CAFM includes a Local Branch and a Global Branch.

[0082] In the local branch, the input feature map first adjusts the number of channels through a 1×1 convolution, and then performs a Channel Shuffle operation to enhance cross-channel interaction and information integration. In this process, the input tensor is first grouped by channels, and depthwise separable convolutions are applied within each group to activate the mixing between channels. Then, the outputs of each group are concatenated along the channel dimension to form a new tensor. Finally, a 3×3×3 convolution is used for feature extraction. Through the above operations, the local branch effectively enhances the cross-channel interaction and information integration capabilities of features, providing a rich and discriminative feature representation for the CAFM module. These operations act together on the input feature map, enabling the model to better capture and utilize the key information in the features. The output of the local branch can be expressed by the following formula:

[0083] F Conv =Conv 3×3×3 (CS(Conv 1×1 (Y)))

[0084] In the formula, F Conv is the output of the local branch, Conv 3×3×3 represents the 3×3×3 convolution, Conv 1×1 represents the 1×1 convolution, CS represents the channel shuffle operation, and Y is the input feature.

[0085] In the global branch, tensors for query, key, and value, namely query (Q), key (K), and value (V), are generated through 1×1 convolution and 3×3 depth convolution, obtaining three tensors with shapes of . Then, Q is reshaped into and K is reshaped into The attention map is calculated through the interaction of these tensors. In this embodiment, the attention map is calculated through the interaction of and Then, after feature extraction using a 1×1 convolution and adding it to the input feature, the output F of the global branch is obtained. The expression for the output F att of the global branch is: att In the formula,

[0086]

[0087] where represents passing through and The attention map obtained by the interaction calculation, Softmax represents the Softmax function, and α is a learnable scaling parameter used to control the magnitude of the matrix multiplication of Q and K before applying the Softmax function.

[0088] The output calculated by the CAFM module is calculated as follows:

[0089] F out = F att + F Conv

[0090] In the formula, F out represents the output of the CAFM module.

[0091] Figure 7 FIG. is a schematic structural diagram of the fusion mechanism CAFMFusion based on the convolution and attention fusion module provided by the present invention. It can be understood that due to its inherent locality and limited receptive field, the standard convolution operation is difficult to effectively capture global feature information. To overcome this limitation, the present invention introduces a convolution and attention mechanism fusion module (CAFM) aiming to achieve a comprehensive modeling of global and local features, thereby enhancing the efficiency and accuracy of the model in the surface defect small target detection task. This module ensures the robustness and accuracy improvement of the model when facing complex scenes and multi-scale targets.

[0092] Specifically, CAFM uses the learned spatial weights to dynamically adjust the feature representation to achieve the adaptive fusion of low-level features and corresponding high-level features. Figure 7 FIG. details the proposed fusion strategy framework, the core of which is to adopt the CAFM mechanism to calculate the spatial weight coefficients of feature modulation. In this process, low-level features and high-level features are input into CAFM. After calculating the weights, the features are integrated by means of weighted merging. This mechanism enables the model to dynamically evaluate the importance of features, strengthens the feature signals crucial for target detection, and effectively suppresses the interference of irrelevant backgrounds. In addition, by introducing the skip connection strategy and integrating the original input features into it, this method not only alleviates the gradient disappearance phenomenon but also promotes the optimization of the learning process. Finally, the fused features are further mapped through a 1×1 convolutional layer to extract the final feature representation, and the formula is as follows:

[0093] F fuse = Conv 1×1 (F low · w+(1 - w)+ F low + F high )

[0094] In the formula, F fuse represents the fused feature representation; Conv 1×1is a 1×1 convolution operation, F low represents low-level features, F high represents high-level features, w is the weight corresponding to low-level features, and 1 - w is the weight corresponding to high-level features.

[0095] Step S3, using SlideLoss as the loss function of YOLOv8n; the SlideLoss loss function is used to optimize the loss calculation in the deep learning object detection task by incorporating a Slide weighting mechanism and differentiating the weighting of samples according to the intersection over union (IoU) metric.

[0096] The present invention uses SlideLoss as the loss function of YOLOv8n to optimize the deep learning object detection task. SlideLoss improves the traditional loss function framework by incorporating a weighting mechanism called Slide to address the problem of uneven distribution of easy and hard samples during training. Specifically, SlideLoss differentiates the weighting of samples according to the intersection over union (IoU) metric. At the same time, the design of the anchor points incorporates the information of the effective receptive field. The difficulty of a sample is defined by comparing the IoU value between the predicted bounding box and the ground truth.

[0097] The SlideLoss loss function combines the BCEWithLogitsLoss loss function and the Slide weighting function. Figure 8 is the schematic diagram of the Slide weighting function. The core of the Slide weighting function is that it uses the mean value of all bounding box IoU values as the threshold for differentiation (see Figure 8 ). IoU is an index to measure the overlapping degree between the predicted bounding box and the ground truth, and its value ranges from 0 to 1. By calculating the mean value of all bounding box IoU values, a threshold is obtained to distinguish positive and negative samples. Samples below this threshold are regarded as negative samples, and vice versa. It should be noted that SlideLoss has the ability to adaptively learn the threshold parameter. By assigning higher weights near the threshold, it means that samples with IoU values close to the threshold (i.e., difficult-to-classify samples) will account for a larger proportion in the loss calculation. In this way, the model is guided to pay more attention to those difficult-to-classify samples, thereby improving the recognition ability of these samples. By emphasizing the importance of difficult-to-classify samples, SlideLoss effectively improves the model's recognition ability for complex scenes, which further promotes the improvement of object detection accuracy.

[0098] Emphasize the samples at the boundary through the weighting function Slide, so as to focus more attention on the difficult-to-classify misclassified examples. The specific calculation formula of the Slide weighting function is as follows:

[0099]

[0100] In the formula, f(x) is the Slide weighting function, x is the predicted value, and μ is the central position of the sliding window, which is used to determine the threshold of difficult-to-classify samples.

[0101] The classification loss function used by the baseline model is the BCEWithLogitsLoss loss function, which combines BCELoss and Sigmoid. The specific calculation formula is as follows:

[0102]

[0103] In the formula, BCE represents the BCELoss loss function, and BCE consists of a series of loss terms I n and each loss term corresponds to the prediction error of a sample.

[0104] The SlideLoss loss function combines the BCEWithLogitsLoss loss function and the Slide weighting function. The expression of the SlideLoss loss function is:

[0105]

[0106] In the formula, SlideLoss is the loss function based on the Slide weighting function, I n is the expression of the loss function, x n is the predicted value, y n is the actual value, w n is the weight coefficient, σ(x n ) is the Sigmoid function, and the Sigmoid function is used to map the predicted value x n to the interval of (0, 1).

[0107] Based on YOLOv8n, after optimizing the backbone network, feature fusion link, and loss function through steps S1 to S3, the steel surface defect detection model ECS-YOLOv8n based on the improved YOLOv8n is obtained. The optimized ECS-Yolov8n network architecture is as Figure 9 shown.

[0108] In summary, the present invention first optimizes the backbone network of YOLOv8n. Specifically, the convolutional components in the original C2f module are first replaced with the MBConv architecture. This process includes channel expansion through 1×1 convolution, application of depthwise separable convolution (i.e., DWConv), integration of the SE module, and channel compression using 1×1 convolution again. This series of operations significantly reduces the model parameter scale and computational complexity, while enhancing the feature extraction and representation capabilities, thereby improving the accuracy of defect detection. In the feature fusion link, the present invention innovatively incorporates the CAFM module, which realizes the effective fusion of multi-level features by constructing top-down and bottom-up bidirectional paths. In addition, by using the attention mechanism to weight the importance of different features, this strategy greatly enhances the model's cross-scale target capture ability and improves the adaptability to complex application scenarios. Finally, in the selection of the loss function, the present invention replaces the YOLOv8 loss function with SlideLoss. Considering the diverse shapes and sizes of steel surface defects, SlideLoss effectively alleviates the class imbalance problem and significantly improves the accuracy of target localization by accurately calculating the overlapping area between the predicted box and the ground truth box and dynamically adjusting the loss weight according to this. The optimized ECS-Yolov8n network architecture is as Figure 9 shown. Based on YOLOv8n, the ECS-Yolov8n model optimizes the backbone network, feature fusion link, and loss function, significantly enhancing the model's feature extraction ability, cross-scale target capture ability, and adaptability to complex application scenarios.

[0109] In a preferred embodiment of the present invention, after obtaining the steel surface defect detection model ECS-YOLOv8n based on the improved YOLOv8n, the method further includes:

[0110] Step S4, training the steel surface defect detection model ECS-YOLOv8n.

[0111] First, prepare the dataset: The publicly available strip steel surface defect dataset NEU-DET is used. This dataset contains six types of defect target images: pitted surface (PS), inclusion (In), crazing (Cr), scratches (Sc), rolled-in scale (RS), and patches (Pa). There are 300 images for each defect category, with a size of 200×200, for a total of 1800 data sample images. To effectively train the model, a random division of 7:1:2 is adopted for the training set, test set, and validation set.

[0112] Next, set up the experimental environment: Set the relevant parameters for model training. Use Python 3.10.14, the torch 2.5.2+cu121 deep learning framework, an Intel Core i9-14900k processor, an RTX 4090 GPU, and a 24GB graphics card. During the training phase, the input image size is 640×640, the number of training epochs is 300, the batch size is 16, use SGD as the optimizer, the initial learning rate is 0.01, the momentum factor is 0.937, and the weight decay coefficient is 0.0005.

[0113] Then, determine the evaluation metrics: To evaluate the defect detection model, this application selects accuracy (Precision), recall (Recall), and mean average precision (mAP, denoted as PmAP) as the evaluation metrics. The formulas are as follows:

[0114]

[0115] In the formulas, TP, FP, and FN represent the number of true positives correctly detected, false positives, and false negatives, respectively. AP represents the area enclosed by the P-R curve and the coordinate axis. FPS represents the number of images detected per unit time; Frameum represents the total number of detected images; ElapsedTime represents the total time required for detection.

[0116] The training process is based on the YOLOv8n baseline model and incorporates three specifically designed improvements from the above embodiments to form the ECS-YOLOv8n model. First, replace the C2f module in the original backbone network with the enhanced C2f-ESEM module; second, replace the loss function in YOLOv8 with the optimized SlideLoss; finally, introduce the CAFM module in the feature fusion stage. During the training process, the model continuously iterates and learns to optimize the parameter configuration to minimize the loss function and improve the detection performance.

[0117] Through experiments, it can be seen that each improvement strategy of the present invention has achieved an average precision improvement to varying degrees, and the number of parameters and computational complexity have been effectively reduced. After replacing the C2f module in the original backbone network with the C2f-ESEM module, the mAP% increases by 0.53%, the number of parameters increases slightly, but the computational complexity decreases by 14.81%, and the FPS increases by 5.2%, and the recall rate remains basically unchanged. By optimizing the SlideLoss function, the mAP% is increased by 2.67% compared to the original model, the number of parameters remains unchanged, the computational complexity is further reduced to 6.9G, and the FPS is increased by 5.63%. At the same time, using the C2f-ESEM module, the SlideLoss function, and introducing the CAFM module, although the number of parameters increases slightly, it does not affect the overall lightweighting, the computational complexity is reduced by about 25.93% compared to the original model, and the average precision mean is increased by 4%.

[0118] Taking all improvements into consideration, compared with the original YOLOv8n model, the final ECS-YOLOv8n model has a 4% increase in mAP%, a 0.1M reduction in the number of parameters, and the computational cost is reduced from 8.1G to 6G. After the training process with the above configuration, the ECS-YOLOv8n model performs excellently on the NEU-DET dataset. Compared with the original YOLOv8n model, while maintaining the lightweight characteristics, ECS-YOLOv8n significantly improves the detection accuracy, reduces the computational cost, and improves the running efficiency. These clear results prove the effectiveness of the training process and the rationality of the improvement measures, providing a solid foundation for deployment on resource-constrained devices.

[0119] The method and system for constructing a steel surface defect detection model based on improved YOLOv8n provided by the present invention have the following beneficial effects compared with the prior art:

[0120] 1) The present invention optimizes the C2f architecture of the YOLOv8n backbone network into a C2f-ESEM module, introduces the EffectiveSE attention mechanism, enhances the pertinence of feature representation, enables the model to focus more on key features, and thus improves the accuracy of defect detection.

[0121] 2) The present invention introduces the CAFM module in the feature fusion stage of the YOLOv8 framework. The CAFM module realizes the effective fusion of multi-level features by constructing top-down and bottom-up bidirectional paths, and uses the attention mechanism to weight the importance of different features, greatly enhancing the model's cross-scale target capture ability and improving the adaptability to complex application scenarios.

[0122] 3) SlideLoss is used as the loss function of YOLOv8n. SlideLoss effectively alleviates the class imbalance problem by accurately calculating the overlapping area between the predicted box and the ground truth box and dynamically adjusting the loss weight according to this, significantly improving the accuracy of target localization, and enabling the model to show stronger adaptability and robustness in complex and changeable environments.

[0123] In a preferred embodiment of the present invention, the present invention also provides a system for constructing a steel surface defect detection model based on improved YOLOv8n, and the system includes:

[0124] The backbone network improvement module is used to replace the convolutional components in the original C2f module with the MBConv module on the basis of the YOLOv8n backbone network and introduce the EffectiveSE attention mechanism;

[0125] A feature fusion module is used to introduce a convolutional and attention fusion module (CAFM) during the feature fusion stage of YOLOv8n. The convolutional and attention fusion module is used to construct a local branch and a global branch and fuse features at different levels using the feature fusion strategy CAFMFusion.

[0126] A loss function optimization module is used to adopt SlideLoss as the loss function of YOLOv8n. The SlideLoss function is used to optimize the loss calculation in the deep learning object detection task by incorporating a Slide weighting mechanism to differentially weight samples based on the intersection over union (IoU) metric.

[0127] The steel surface defect detection system based on the improved ECS-YOLOv8n provided by the present invention executes the method for constructing a steel surface defect detection model based on the improved YOLOv8n provided in the foregoing embodiments through the above-mentioned modules. The method for constructing a steel surface defect detection model based on the improved YOLOv8n has been described in detail in the foregoing embodiments, and will not be elaborated herein.

[0128] Figure 10 The block diagram of the electronic device provided by the present invention is shown in Figure 10 As shown, the present invention also provides an electronic device. The electronic device 1000 can be a computing device such as a mobile terminal, a desktop computer, a notebook, a palm computer, and a server. The electronic device 1000 includes a processor 1001 and a memory 1002. Among them, a steel surface defect detection model construction program 1003 is stored on the memory 1002.

[0129] The memory 1002 can be an internal storage unit of the computer device in some embodiments, such as the hard disk or memory of the computer device. The memory 1002 can also be an external storage device of the computer device in other embodiments, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 1002 can also include both the internal storage unit and the external storage device of the computer device. The memory 1002 is used to store application software installed on the computer device and various types of data, such as program codes installed on the computer device. The memory 1002 can also be used to temporarily store data that has been output or will be output. In one embodiment, when the steel surface defect detection model construction program 1003 is executed by the processor 1001, the following steps are implemented:

[0130] Based on the YOLOv8n backbone network, the MBConv module is adopted to replace the convolutional components in the original C2f module, and the EffectiveSE attention mechanism is introduced;

[0131] In the feature fusion stage of YOLOv8n, a convolutional and attention fusion module CAFM is introduced; the convolutional and attention fusion module is used to fuse features of different levels by constructing a local branch and a global branch and using the feature fusion strategy CAFMFusion;

[0132] SlideLoss is adopted as the loss function of YOLOv8n; the SlideLoss loss function is used to perform differential weighting on samples according to the intersection over union (IoU) index by incorporating the Slide weighting mechanism, so as to optimize the loss calculation in the deep learning object detection task.

[0133] In some embodiments, the processor 1001 can be a central processing unit (CPU), a microprocessor or other data processing chips, which are used to run the program code stored in the memory 1002 or process data, such as executing the program for constructing the steel surface defect detection model, etc.

[0134] This embodiment also provides a computer-readable storage medium, on which a program for constructing a steel surface defect detection model is stored. When the program for constructing the steel surface defect detection model is executed by a processor, the following steps are implemented:

[0135] Based on the YOLOv8n backbone network, the MBConv module is adopted to replace the convolutional components in the original C2f module, and the EffectiveSE attention mechanism is introduced;

[0136] In the feature fusion stage of YOLOv8n, a convolutional and attention fusion module CAFM is introduced; the convolutional and attention fusion module is used to fuse features of different levels by constructing a local branch and a global branch and using the feature fusion strategy CAFMFusion;

[0137] SlideLoss is adopted as the loss function of YOLOv8n; the SlideLoss loss function is used to perform differential weighting on samples according to the intersection over union (IoU) index by incorporating the Slide weighting mechanism, so as to optimize the loss calculation in the deep learning object detection task.

[0138] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a steel surface defect detection model based on improved YOLOv8n, characterized in that: include: Based on the YOLOv8n backbone network, the MBConv module is used to replace the convolution component in the original C2f module, and the EffectiveSE attention mechanism is introduced; In the feature fusion stage of YOLOv8n, a convolution and attention fusion module CAFM is introduced; the convolution and attention fusion module is used to fuse features of different levels by constructing local branches and global branches using the feature fusion strategy CAFMFusion; SlideLoss is used as the loss function of YOLOv8n; the SlideLoss loss function is used to optimize the loss calculation in deep learning target detection tasks by incorporating the Slide weighting mechanism and differentially weighting the samples according to the intersection-over-union (IoU) metric.

2. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 1, characterized in that: Based on the YOLOv8n backbone network, the MBConv module is used to replace the convolution component in the original C2f module, and the ESE attention mechanism is introduced, including: Introducing the MBConv module, which includes a standard convolutional layer, a depth-separable convolution, a squeeze-excitation module, and a Dropout component; The EffectiveSE attention mechanism is used to replace the SE attention mechanism in the MBConv module to construct the ESE-MBConv module; the EffectiveSE attention mechanism includes a global average pooling layer, a fully connected layer and a Sigmoid activation function; the global average pooling layer is used to reduce the feature map of each channel to a single scalar value to condense the information; the fully connected layer is used to learn and assign weight coefficients to each channel; the Sigmoid activation function is used to generate feature weights that reflect the importance of the channel; The Bottleneck layer is improved through the ESE-MBConv module to obtain the feature extraction component C2f-ESEM.

3. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 1, characterized in that: The mathematical expression of the EffectiveSE attention mechanism is: A eSE (X)=σ(W C (F gap (X))) Where X represents the input feature map, H and W are the height and width of the feature map respectively, and X ij is the value of the feature map at position (i, j), F gap represents global average pooling, WC represents fully connected layers, σ represents activation function, A eSE represents the channel attention weight, X refine Represents the refined feature map.

4. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 2, characterized in that: The workflow of the feature extraction component C2f-ESEM includes: Channel expansion is performed through 1×1 convolution, and the Splitting operation is performed to divide the expanded channels into two groups. One group of data streams is processed by the n-layer ESE-MBConv module, and the other group of data is spliced ​​with the output of the ESE-MBConv module in the channel dimension. Finally, the number of channels is adjusted through 1×1 convolution.

5. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 1, characterized in that: The convolution and attention fusion module CAFM includes a local branch and a global branch; In the local branch, the input feature map first adjusts the number of channels through 1×1 convolution, then performs channel shuffling to enhance cross-channel interaction and information integration, and then uses 3×3×3 convolution for feature extraction; In the global branch, 1×1 convolution and 3×3 deep convolution are used to generate query, key, and value tensors. The attention map is calculated through the interaction of these tensors, and then 1×1 convolution is used for feature extraction and added to the input features to obtain the output of the global branch.

6. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 1, characterized in that: The CAFM module adopts the feature fusion strategy CAFMFusion to calculate the weights of the input low-level features and high-level features, and then perform weighted merging to obtain the fused features, which are expressed as follows: F fuse =Conv 1×1 (F low ·w+(1-w)+F low +F high ) In the formula, F fuse Represents the fused feature representation; Conv 1×1 is a 1×1 convolution operation, F low represents low-level features, F high represents high-level features, w is the weight corresponding to the low-level features, and 1-w is the weight corresponding to the high-level features.

7. The method for constructing a steel surface defect detection model based on improved YOLOv8n according to claim 1, characterized in that: The SlideLoss loss function combines the BCEWithLogitsLoss loss function and the Slide weighting function; the expression of the SlideLoss loss function is: Where SlideLoss is the loss function based on the Slide weighted function, I n is the expression of the loss function, f(x) is the Slide weighting function, x n is the predicted value, y n is the actual value, w n is the weight coefficient, σ(x n ) is the Sigmoid function, which is used to convert the predicted value x n Mapped to the interval (0,1).

8. A steel surface defect detection model construction system applied to the steel surface defect detection model construction method based on improved YOLOv8n as described in any one of claims 1 to 7, characterized in that: include: The backbone network improvement module is used to replace the convolution component in the original C2f module with the MBConv module based on the YOLOv8n backbone network, and introduce the EffectiveSE attention mechanism; A feature fusion module is used to introduce a convolution and attention fusion module CAFM in the feature fusion stage of YOLOv8n; the convolution and attention fusion module is used to fuse features of different levels by constructing local branches and global branches using a feature fusion strategy CAFMFusion; The loss function optimization module is used to use SlideLoss as the loss function of YOLOv8n; the SlideLoss loss function is used to differentially weight samples according to the intersection-over-union (IoU) metric by incorporating the Slide weighting mechanism to optimize the loss calculation in deep learning target detection tasks.

9. An electronic device, It is characterized in that comprising a memory and a processor, wherein: The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the method for constructing a steel surface defect detection model based on improved YOLOv8n as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the method for constructing a steel surface defect detection model based on improved YOLOv8n as described in any one of claims 1 to 7.