YOLOv8 network improvement method and product surface defect detection method and system
By introducing VanillaNet, EG-C2f module, dynamic snake convolution module and Wise-IoU loss function in the YOLOv8 network, the problems of high computing resource requirements, poor detection of slender defects and limitations of loss function in industrial production are solved, and higher detection accuracy and faster inference speed are achieved, which is suitable for complex defect detection tasks.
Patent Information
- Application Number
- CN202510122394.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
When the existing YOLOv8 model is used for product surface defect detection in industrial production, the high demand for computing resources, poor detection of slender defects, and limitations in the loss function, resulting in difficulty in meeting the actual industrial needs of the detection accuracy and real-time.
By introducing VanillaNet, EG-C2f module, dynamic snake convolution module, and Wise-IoU loss function into the YOLOv8 network, the network structure is improved to improve detection accuracy and computing efficiency.
It significantly improves the detection accuracy of complex geometric defects, reduces the computing resource requirements, better meets the real-time requirements of industrial applications, and is suitable for strip surfaces and other complex defect detection tasks.
Smart Images

Figure CN120047410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial production, involves the detection of product surface defects, particularly relates to the YOLOv8 network, and specifically is an improved method of the YOLOv8 network and a product surface defect detection method and system. Background Technique
[0002] With the continuous development of manufacturing automation, the detection of strip steel surface defects has occupied a crucial position in industrial production. Strip steel is an important material widely used in metallurgy and manufacturing. Due to its large surface area and complex production environment, it is extremely easy to generate defects such as cracks, scratches, patches, oxidation, etc. during the production and processing process. The existence of these defects will not only affect the quality of the product, but also have a serious impact on the service life and safety of downstream products. Therefore, achieving efficient and accurate automated defect detection is a technical problem that industrial manufacturing urgently needs to solve.
[0003] Currently, the detection of strip steel surface defects mainly relies on manual inspection or traditional image processing algorithms. Manual inspection has low efficiency, and is limited by the experience and skills of inspectors, and is prone to false detection and missed detection. In addition, due to the limited ability of traditional image processing methods to extract complex morphological features, they cannot effectively process the diverse defects on the strip steel surface, especially when dealing with complex geometric shapes such as slender cracks and scratches, they perform poorly. Therefore, object detection algorithms based on deep learning have gradually been introduced into the field of surface defect detection. In the field of deep learning, the YOLO (You Only Look Once) series of algorithms have become a popular choice in the field of defect detection due to their efficient real-time object detection ability. The YOLO series of algorithms realize the simultaneous execution of the object localization and classification tasks through a single-stage object detection method. Especially, YOLOv8, as the latest version of the YOLO series, has significantly improved in detection accuracy and inference speed. However, there are still some problems in the actual application of the existing YOLOv8 model:
[0004] 1. Model complexity: The YOLOv8 network structure is complex, with high computational resource requirements, and it is difficult to meet the requirements of real-time performance and computational efficiency in industrial production, especially it performs poorly in resource-constrained embedded devices.
[0005] 2. Poor detection of slender defects: Defects with slender shapes such as cracks and scratches on the strip steel surface are difficult to be effectively captured in traditional convolution operations due to their special geometric features, resulting in a high missed detection rate of the model.
[0006] 3. Limitations of the loss function: The CIoU loss function adopted by YOLOv8 has limitations in dealing with bounding boxes with large shape differences. Especially when the aspect ratios of the predicted box and the ground truth box are similar, the discrimination degree of the loss value is insufficient, which affects the convergence speed and localization accuracy of the model. Summary of the Invention
[0007] The purpose of the present invention is to provide an improved method for the YOLOv8 network, a method and system for detecting product surface defects, so as to solve the problems pointed out in the above-mentioned background technology.
[0008] In the first aspect, the present invention provides an improved method for the YOLOv8 network, and the method includes: improving the original YOLOv8 network to obtain a new YOLOv8 network; wherein, the improvement of the original YOLOv8 network includes: using VanillaNet to replace the convolutional layer in the Backbone of the original YOLOv8 network; using the EG-C2f module to replace three C2f modules in the Neck of the original YOLOv8 network; the three C2f modules are respectively two C2f modules that are outputs and located at the front and rear positions in the Neck, and another C2f module other than the output; there are three C2f modules that are outputs in the Neck; using the dynamic snake-shaped convolutional module to replace the ordinary convolution that is output and located in the middle of the Neck of the original YOLOv8 network; using the Wise-IoU loss function to replace the CIoU loss function in the original YOLOv8 network; conducting experimental verification on the new YOLOv8 network on the NEU-DET dataset to obtain an improved YOLOv8 network, so as to detect product surface defects based on the improved YOLOv8 network.
[0009] In the present invention, an improved YOLOv8 network is proposed. By introducing VanillaNet, EG-C2f module, dynamic snake-shaped convolutional module, and Wise-IoU loss function, the performance of the YOLOv8 network in product surface defect detection is significantly improved, and the problems that the existing YOLOv8 network has high computational resource requirements, poor detection of slender defects, and limitations in the loss function when detecting product surface defects, resulting in the detection accuracy and real-time performance of the YOLOv8 network being difficult to meet the industrial actual needs, are solved.
[0010] In one implementation manner of the first aspect, the EG-C2f module combines GhostConv and ECA attention mechanism; wherein, GhostConv generates partial feature maps and generates the remaining feature maps through linear transformation to eliminate redundant features; the ECA attention mechanism adaptively adjusts the interaction weights between channels through convolutional operations.
[0011] In this implementation manner, GhostConv significantly reduces the number of network parameters and computational volume, while the ECA attention mechanism improves the network's recognition ability for complex defects, especially suitable for irregular defects on the strip steel surface.
[0012] In an implementation of the first aspect, the dynamic snake-shaped convolution module adaptively adjusts the receptive field of the convolution kernel to ensure that the convolution kernel can adjust its position following the geometric shape of the surface defects of the product.
[0013] In this implementation, the dynamic snake-shaped convolution module can capture slender defects, improving the performance of the network in detecting complex defects such as cracks and scratches.
[0014] In an implementation of the first aspect, the NEU-DET dataset is a standard dataset for the task of product surface defect detection.
[0015] In an implementation of the first aspect, the NEU-DET dataset includes at least any one or two or more combined types of surface defects: cracks, inclusions, patches, scratches, roll marks, and pitting.
[0016] In an implementation of the first aspect, the experimental verification of the novel YOLOv8 network on the NEU-DET dataset includes: dividing the NEU-DET dataset into a training set, a validation set, and a test set; using the training set, the validation set, and the test set to train, validate, and test the novel YOLOv8 network respectively.
[0017] In an implementation of the first aspect, before the step of training the novel YOLOv8 network using the training set, the experimental verification of the novel YOLOv8 network on the NEU-DET dataset further includes: performing data augmentation on the training set to use the augmented training set to train the novel YOLOv8 network.
[0018] In a second aspect, the present invention provides a method for detecting product surface defects, the method including: detecting product surface defects based on the improved YOLOv8 network described above.
[0019] In an implementation of the second aspect, detecting product surface defects based on the improved YOLOv8 network includes: acquiring an image of product surface defects; inputting the image of product surface defects into the improved YOLOv8 network so that the improved YOLOv8 network outputs product surface defects.
[0020] In a third aspect, the present invention provides a system for detecting product surface defects, the system including: an image acquisition device and the improved YOLOv8 network described above; wherein, the image acquisition device is used to acquire an image of product surface defects and to input the image of product surface defects into the improved YOLOv8 network; the improved YOLOv8 network is used to output product surface defects based on the image of product surface defects.
[0021] As described above, the improved method of the YOLOv8 network, the product surface defect detection method and system of the present invention have the following beneficial effects:
[0022] (1) Compared with the prior art, the present invention provides an improved YOLOv8 network. By introducing the VanillaNet, EG-C2f module, dynamic snake-shaped convolution module, and Wise-IoU loss function on the basis of the existing YOLOv8 network, the improved YOLOv8 network not only improves the detection accuracy of complex geometric shape defects, but also greatly reduces the demand for computing resources, and can better meet the real-time requirements of industrial applications.
[0023] (2) The improved YOLOv8 network provided by the present invention is not only applicable to the detection of strip steel surface defects, but can also be widely applied to the surface defect detection tasks of other products with complex geometric shapes and multi-scale features. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It shows a flowchart of the improved method of the YOLOv8 network described in the embodiment of the present invention.
[0025] Figure 2 It shows a schematic structural diagram of the YOLOv8 network in the prior art.
[0026] Figure 3 It shows a schematic structural diagram of the EG-C2f module described in the embodiment of the present invention.
[0027] Figure 4 It shows a schematic structural diagram of the GhostConv described in the embodiment of the present invention.
[0028] Figure 5 It shows a schematic structural diagram of the ECA attention mechanism described in the embodiment of the present invention.
[0029] Figure 6 It shows a schematic diagram of coordinate calculation of the dynamic snake-shaped convolution module described in the embodiment of the present invention.
[0030] Figure 7 It shows a schematic diagram of the receptive field change of the dynamic snake-shaped convolution module described in the embodiment of the present invention.
[0031] Figure 8 It shows a schematic structural diagram of the improved YOLOv8 network described in the embodiment of the present invention.
[0032] Figure 9 It shows a schematic diagram of the defect types of the NEU-DET dataset described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0033] The following specific examples illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0034] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0035] Refer to Figures 1 to 9 The following embodiments of the present invention provide an improved method for the YOLOv8 network, a method and system for detecting surface defects of products. Compared with the prior art, the present invention provides an improved YOLOv8 network. By introducing the VanillaNet, EG-C2f module, dynamic snake-shaped convolution module, and Wise-IoU loss function on the basis of the existing YOLOv8 network, the improved YOLOv8 network not only improves the detection accuracy of defects with complex geometric shapes, but also greatly reduces the demand for computing resources, and can better meet the real-time requirements of industrial applications; the improved YOLOv8 network provided by the present invention is not only applicable to the detection of strip steel surface defects, but also can be widely applied to the surface defect detection tasks of other products with complex geometric shapes and multi-scale features.
[0036] As Figure 2 shown, the existing YOLOv8 network (i.e., the "original YOLOv8 network" in the present invention) mainly consists of the following three major parts:
[0037] 1. Backbone: It uses a series of convolutional and deconvolutional layers to extract features, and also uses residual connections and bottleneck structures to reduce the size of the network and improve performance. This part uses the C2f module as the basic building unit. Compared with the C3 module of YOLOv5, the C2f module has fewer parameters and better feature extraction ability.
[0038] 2. Neck: It uses multi-scale feature fusion technology to fuse feature maps from different stages of the Backbone to enhance the feature representation ability. Specifically, the Neck part of YOLOv8 includes an SPPF module, a PAA module, and two PAN modules.
[0039] 3. Head: It is responsible for the final object detection and classification tasks, including a detection head and a classification head. The detection head contains a series of convolutional layers and deconvolutional layers for generating detection results; the classification head uses global average pooling to classify each feature map.
[0040] Next, the technical solutions in the embodiments of the present invention will be described in detail with reference to the accompanying drawings in the embodiments of the present invention.
[0041] As Figure 1 shown, in one embodiment, the present invention provides an improved method for the YOLOv8 network, and the method includes:
[0042] Step S1. Improve the original YOLOv8 network to obtain a new YOLOv8 network.
[0043] As Figure 1 shown, in this embodiment, the improvement of the original YOLOv8 network includes:
[0044] Step S101. Replace the convolutional layers in the Backbone of the original YOLOv8 network with VanillaNet.
[0045] VanillaNet is a new type of neural network architecture that emphasizes simple design and avoids complex operations such as depth, shortcuts, and self-attention. By enhancing non-linearity through deep training strategies and sequential activation functions, it achieves performance comparable to complex deep networks in computer vision tasks and is suitable for environments with limited resources.
[0046] It should be noted that by replacing the convolutional layers in the Backbone of the original YOLOv8 network with VanillaNet, VanillaNet is used as the backbone feature extraction network for the improved YOLOv8 network, reducing the complexity of the network and improving the computational efficiency.
[0047] In one embodiment, VanillaNet uses a convolutional operation with a stride of 4 to extract features from the input image (corresponding to the subsequent "product surface defect image"), reducing the size of the feature map, and applying a max-pooling operation after each convolutional layer to reduce the computational overhead. This can improve the inference speed and reduce the computational resource requirements while ensuring that the network extracts sufficient features.
[0048] It should be noted that the design of VanillaNet significantly reduces the computational complexity and parameter amount of the network by reducing the number of network layers and convolution operations, and can meet the needs of strip surface defect detection; by reducing the size of the feature map and combining the maximum pooling layer operation, it effectively reduces the amount of calculation and improves the inference speed, effectively improving the real-time and efficiency of the improved YOLOv8 network in strip surface defect detection.
[0049] Step S102: Use the EG-C2f module to replace the three C2f modules in the Neck of the original YOLOv8 network.
[0050] like Figure 2 As shown, in this embodiment, the three C2f modules are two C2f modules located at the front and rear positions as outputs in the Neck (corresponding to Figure 2 , a C2f module in the front position and a C2f module in the back position in accordance with the order of connection of the modules in the Neck, and another C2f module other than the output; there are three C2f modules in the Neck as output.
[0051] It should be noted that by introducing the EG-C2f module, the computational scale of the network is further reduced, the feature extraction process is optimized, and the model complexity is further reduced.
[0052] like Figures 3 to 5 As shown, in one embodiment, the EG-C2f module combines GhostConv and ECA attention mechanisms, eliminates redundant feature maps, reduces the number of network parameters, and improves the ability to extract complex defect features on the strip surface, ensuring the lightweight design of the network model. Through this improvement, the network can maintain high-precision detection performance in a resource-constrained environment; wherein, the GhostConv generates partial feature maps and generates the remaining feature maps through linear transformation to eliminate redundant features; the ECA attention mechanism adaptively adjusts the interaction weights between channels through convolution operations, so that the improved YOLOv8 network performs well in the detection of complex defects on the strip surface.
[0053] GhostConv is a "Ghost" feature map that mines the required information from the original features through a series of linear transformations with very little computational effort. It uses conventional technical means in the field, so it will not be described in detail here.
[0054] The ECA attention mechanism is a method of channel attention. This algorithm makes certain improvements on the basis of the SE attention mechanism. First, the ECA authors believe that although the full connection dimensionality reduction in SE can reduce the complexity of the model, it destroys the direct correspondence between channels and their weights. After dimensionality reduction and then dimensionality increase, the correspondence between weights and channels is indirect, and it also uses conventional technical means in the field, so it will not be elaborated in detail here.
[0055] The SE (full English name: Squeeze and Excitation) attention mechanism is a method of determining weights in the channel attention mode. It achieves the purpose of prioritizing the primary and secondary by allocating weights among different channels.
[0056] It should be noted that the EG-C2f module combines GhostConv and the ECA attention mechanism. GhostConv generates some feature maps and generates the remaining feature maps through linear transformation, eliminating redundant features and significantly reducing the number of network parameters and computational volume. The ECA attention mechanism adaptively adjusts the interaction weights between channels through convolution operations, thereby improving the model's ability to identify complex defects, especially suitable for detecting irregular defects on the strip steel surface.
[0057] Step S103: Replace the ordinary convolution in the C2f module that is used as the output and is in the middle position in the Neck of the original YOLOv8 network with a dynamic snake convolution module.
[0058] As Figure 2 shown, in this step S103, a dynamic snake convolution module is used to replace Figure 2 a kind of C2f module that is used for Output output in the Neck of the original YOLOv8 network and is in the middle position according to the order of connection of each module in the Neck.
[0059] The dynamic snake convolution module, full English name: Dynamic Snake Convolution, English abbreviation: DSC, is a convolution module designed for slender and weak local structural features and complex and variable global morphological features; in the present invention, in order to deal with defects such as cracks and scratches on the strip steel surface with slender and complex shapes, a dynamic snake convolution module is introduced, enhancing the network's detection ability for slender and curved defects, enabling the convolution kernel to be adaptively adjusted to accurately capture complex features such as cracks and scratches on the strip steel surface.
[0060] As Figure 6 and Figure 7 shown, in an embodiment, the dynamic snake convolution module adaptively adjusts the receptive field of the convolution kernel to ensure that the convolution kernel can follow the geometric shape of the product surface defect to adjust its position, especially suitable for the detection task of slender and curved defects.
[0061] It should be noted that traditional convolutional kernels often have deficiencies in detecting defects with slender shapes such as cracks and scratches, and are unable to effectively capture these complex geometric shapes. To overcome this problem, the present invention adopts a dynamic snake-shaped convolutional module, which can capture slender defects by adaptively adjusting the receptive field and position of the convolutional kernel. The convolutional kernel adjusts its position in the X-axis and Y-axis directions through offsets to ensure that it can perform convolutional operations following the geometric structure of the defect, thereby improving the performance of the model in detecting complex defects such as cracks and scratches.
[0062] The adjustment of the convolutional kernel is based on the offset of the receptive field, which can better adapt to the geometric characteristics of slender structures such as cracks and scratches.
[0063] Given a standard two-dimensional convolution with coordinates represented as K (used to define the spatial positions to be sampled during the convolution operation), its center coordinates are represented as K i =(x i , y i ), with a dilation rate set to 1, the 3×3 convolutional kernel can be represented as:
[0064] K = {(x - 1, y - 1), (x - 1, y), Λ, (x + 1, y + 1),};
[0065] To make the convolutional kernel better adapt to the complex geometric characteristics of steel surface defects, a deformation offset Δ is introduced. However, if the network model randomly learns the deformation offset all at once, the receptive field may deviate from the target. To prevent the excessive deviation of the receptive field from the target, the dynamic snake-shaped convolution adopts an iterative strategy to ensure that only one target is processed at a time and sequentially matches the observable positions, thereby maintaining the continuity of attention.
[0066] Straighten the standard convolutional kernel on the x-axis and y-axis. The specific position of each grid in K is represented as: K i±c =(x i±c , y i±c ), where c = {0, 1, 2, 3, 4} represents the horizontal distance between each grid of the convolutional kernel and the central position grid, used to describe the offsets of the convolutional kernel on the x-axis and y-axis in the following. The position calculation of each grid depends on the position of the previous grid, which is an accumulative process. K i+1 relative to the central position K i increases by an offset Δ = {δ | δ ∈ [-1, 1]}. Therefore, to ensure the matching of the convolutional kernel with the linear morphological structure, an accumulation operation needs to be performed on the offset.
[0067] The change in the x-axis direction is:
[0068]
[0069] The change in the y-axis direction is as follows:
[0070]
[0071] where i and j respectively represent variables along the x-axis and y-axis, used to describe the dynamic positions of the convolutional kernel on the two coordinate axes; K i±c represents the offset on the x-axis, indicating that the convolutional kernel will move in the x-axis direction according to the offset c. The specific explanation is:
[0072] (x i+c , y i+c ) and (x i-c , y i-c ) respectively represent moving c units to the right or left in the x-axis direction, and the y coordinate remains unchanged. Here, y i+c and y i-c are cumulative offsets, representing the change in the y-axis direction, and the adjustment is achieved through the cumulative offset ; K j±c represents the offset on the y-axis, indicating that the convolutional kernel will move in the y-axis direction according to the offset c. The specific explanation is: (x j+c , y j+c ) and (x j-c , y j-c ) respectively represent moving c units up or down in the y-axis direction, and the x coordinate remains unchanged. Here, x i+c and x i-c are cumulative offsets, representing the change in the x-axis direction, and the adjustment is achieved through the cumulative offset ; Δx and Δy respectively represent the change step sizes of the convolutional kernel in the horizontal and vertical coordinate axes directions; x i and y i respectively represent the abscissa in the horizontal direction and the ordinate in the vertical direction of the center of the convolutional kernel under the condition of the change in the x-axis direction, that is, the exact position of the center point of the convolutional kernel in the two-dimensional image space; x j and y j respectively represent the specific coordinates of the convolutional kernel in the horizontal and vertical coordinate axes directions under the condition of the change in the y-axis direction; x i+c and y i+c respectively represent the positions of the convolutional kernel after offsetting c units in the horizontal and vertical coordinate axes directions under the condition of the change in the x-axis direction; x i-c and y i-c respectively represent the positions of the convolutional kernel after offsetting c units in the horizontal and vertical coordinate axes directions under the condition of the change in the x-axis direction; x j+c and y j+crespectively represent the positions of the convolution kernel after shifting c units in the horizontal and vertical coordinate axes when the y-axis changes; x j-c and y j-c respectively represent the positions of the convolution kernel after shifting c units downward in the horizontal and vertical coordinate axes when the y-axis changes.
[0073] Since the offset Δ is usually a decimal number while coordinates usually exist in integer form, the bilinear interpolation method is used to process it, expressed as:
[0074] M = ∑ M' B(M', M) · M';
[0075] B(M, M ′ ) = b(M x , M x ′ ) · b(M y , M y ′ );
[0076] Among them, M represents the fractional parts of the changes K i±c in the x-axis direction and K j±c in the y-axis direction, including M x and M y . Among them, M x represents the fractional part of the change K i±c in the x-axis direction above, M y represents the fractional part of the change K j±c in the y-axis direction above, and both M x and M y are results obtained through interpolation calculations; M' lists all the spatial positions of integers and is the bilinear interpolation kernel; B can be decomposed into two one-dimensional kernels; b is a 1D linear interpolation function responsible for calculating the weight between the target position and the grid position: M x ' represents the integer coordinate after bilinear interpolation processing in the x-axis direction; M' y represents the integer coordinate after bilinear interpolation processing in the y-axis direction.
[0077] It should be noted that bilinear interpolation is a mathematical method for estimating the value of a specific point in a two-dimensional space, which is an extension of linear interpolation in a two-dimensional space. Its idea is to determine the value of this point by performing interpolation calculations between four given neighboring points.
[0078] Due to the two-dimensional changes in the x-axis and y-axis, the dynamic snake-shaped convolution kernel covers a receptive field range of 9×9 during the deformation process. This adaptation to the dynamic structure is more suitable for detecting irregular defect types on the steel surface.
[0079] Step S104: Replace the CIoU loss function in the original YOLOv8 network with the Wise-IoU loss function.
[0080] It should be noted that, in order to improve the accuracy of the YOLOv8 network in bounding box localization, the present invention uses the Wise-IoU loss function to replace the original CIoU loss function. The Wise-IoU loss function can reduce the negative impact of low-quality samples on the training of the network model and improve the localization accuracy of the network model for defect regions by optimizing the bounding box regression process. The Wise-IoU loss function further enhances the robustness of the network model by dynamically adjusting the weights of samples. Especially in the detection of complex defects on the strip steel surface, it significantly improves the accuracy of bounding box prediction. Specifically, Wise-IoU adaptively adjusts the gradients of low-quality samples through a dynamic non-monotonic focusing mechanism, reduces the negative impact of low-quality samples on model training, and improves the localization accuracy of the network model for complex-shaped defects by optimizing the regression method of bounding boxes. This improvement is particularly effective when dealing with defects with large shape differences.
[0081] The present invention uses the Wise-IoU function to replace the original CIoU function as the loss function of the strip steel surface defect detection model. Wise-IoUv1 constructs distance attention based on distance metrics, resulting in a two-layer attention mechanism. Its calculation formula is as follows:
[0082] L WIoUv1 = R WIoU L IoU ;
[0083]
[0084] where x gt and y gt respectively represent the abscissa and ordinate of the center point of the ground truth box; W g and H g respectively represent the width and height in the smallest predicted box; L IoU represents the overlap between the predicted box and the ground truth box; R WIoU is the penalty term; L WIoUv1 represents the loss value of the Wise-IoUv1 function; x' and y' respectively represent the abscissa and ordinate of the center point of the predicted box.
[0085] Wise-IoUv3 is a loss function with a dynamic non-monotonic focusing mechanism built on top of Wise-IoUv1. The dynamic non-monotonic focusing mechanism introduces an outlier degree β to replace IoU. The magnitude of the outlier degree β represents the quality of the prediction boxes of its sample data. Gradients are assigned according to the quality of the prediction boxes. For prediction boxes with a small outlier degree, a small gradient gain is assigned, which can make the bounding box regression target focus on prediction boxes of ordinary quality. For prediction boxes with a large outlier degree, a small gradient gain is also assigned, which can effectively avoid large harmful gradients generated by low-quality prediction boxes. Wise-IoUv3 also constructs a non-monotonic focusing coefficient using the outlier degree β, and its calculation formula is as follows:
[0086] L WIoUv3 = rL WIoUv1 ;
[0087]
[0088] where r is the dynamic gradient gain coefficient for adjusting the loss weight; σ and α are hyperparameters, where α controls the change speed of the gradient gain; β is the outlier degree of the prediction box. When β = α, r = 1. When the outlier degree of the prediction box satisfies β = Q (Q is a fixed value), the prediction box will obtain the highest gradient gain; represents the IoU loss function; represents the moving average of the IoU loss function; L WIoUv3 represents the loss value of the Wise-IoUv3 function.
[0089] Since is dynamic, the quality division standard for prediction boxes is also dynamic. This enables Wise-IoUv3 to allocate the most appropriate gradient gain to different prediction boxes according to the real-time situation at each moment, thereby improving the performance and generalization ability of the model.
[0090] It should be noted that the execution order of steps S101 to S104 is not a condition for limiting the present invention. In practical applications, they can be executed in sequence or simultaneously. Among them, when executed in sequence, the order of who comes first and who comes later is also uncertain ( Figure 1 only gives an example).
[0091] Step S2: Conduct experimental verification on the novel YOLOv8 network on the NEU-DET dataset to obtain an improved YOLOv8 network for detecting product surface defects based on the improved YOLOv8 network.
[0092] It should be noted that through a large number of experimental verifications on the NEU-DET dataset, the improved YOLOv8 network of the present invention (such as Figure 8As shown in the figure, when detecting strip surface defects, it shows higher detection accuracy and faster inference speed. Experimental data shows that the improved YOLOv8 network has an average detection accuracy (mean Average Precision, abbreviated as mAP) increased by 2.3 percentage points, and the model complexity is reduced by 52.4%, significantly reducing the computational resource requirements. Especially in the detection of slender defects such as cracks and scratches, the improved model shows higher detection accuracy and lower missed detection rate.
[0093] In one embodiment, the NEU-DET dataset is a standard dataset for product surface defect detection tasks.
[0094] In one embodiment, the product surface defects are imaged by an image acquisition device to obtain the NEU-DET dataset.
[0095] In one embodiment, the image acquisition device uses an industrial camera.
[0096] As Figure 9 shown, in one embodiment, the NEU-DET dataset includes at least, but is not limited to, any one or two or more combinations of the following surface defects: cracks, inclusions, patches, scratches, roll marks, and pitting.
[0097] In one embodiment, the NEU-DET dataset includes the following six types of surface defects: cracks, inclusions, patches, scratches, roll marks, and pitting. For each type of surface defect, it contains 300 high-resolution images, a total of 1800 images. The image resolution is 200×200, and the image quality is high, which can truly reflect the defect types on the strip surface.
[0098] In one embodiment, the experimental verification of the new YOLOv8 network on the NEU-DET dataset includes: dividing the NEU-DET dataset into a training set, a validation set, and a test set; using the training set, the validation set, and the test set to train, validate, and test the new YOLOv8 network respectively.
[0099] In one embodiment, the NEU-DET dataset is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1 to ensure the diversity of training data and the independence of validation and test data.
[0100] In one embodiment, before dividing the NEU-DET dataset into a training set, a validation set, and a test set, the image data in the NEU-DET dataset is screened and labeled.
[0101] The following further explains the experimental verification of the new YOLOv8 network through specific embodiments.
[0102] Experimental Environment and Details: This experiment was conducted in an environment with the Ubuntu 22.04 operating system. The GPU selected was NVIDIA GeForce GTX 1080ti (with 12GB of video memory). The deep learning framework was Pytorch 1.9.2, and CUDA11.4 and cudnn11.4 were used for acceleration. The training settings were the same as those of YOLOv8. The input image size was set to 640×640, and the batch size was 32. The optimizer used Stochastic Gradient Descent (SGD) with a momentum of 0.9, and the learning rate was 0.001. The network was trained for 500 rounds. During training, the Mosaic data augmentation strategy was used to accelerate model convergence, and Mosaic augmentation was turned off in the last 10 rounds.
[0103] Experimental Evaluation Metrics: Various evaluation metrics commonly used in object detection tasks were used to evaluate the performance of the experimental model in strip surface defect detection, including precision (P), recall (R), mAP (mean average precision) 0.5, mAP0.5:0.95, the number of model parameters, and the model size. The relevant formulas are as follows:
[0104]
[0105] Among them, TP represents the number of correctly predicted ones; FP represents the number of misjudging other categories as the true category; FN represents the number of misjudging the true category as other categories; AP represents the average correct rate; N represents the total number of iterative trainings; AP i represents the average correct rate of the i-th iterative training.
[0106] Ablation Experiment: The baseline model was YOLOv8n. First, VanillaNet was introduced as the backbone network on its basis. The experimental results showed that the number of parameters decreased by 1.3×10^6, the computational cost decreased by 3.1 GFLOPs, and the detection accuracy remained basically stable. Then, the EG-C2f module was added, further reducing the number of parameters and computational cost. The mAP@0.5 increased by 1 percentage point, and the computational cost decreased by 0.9 GFLOPs. Subsequently, the Wise-IoU loss function was used to replace the original loss function to reduce the harmful gradients generated by low-quality samples, and the mAP@0.5 increased by 1.1 percentage points again. Finally, the dynamic snake convolution module (i.e., C2f_DyS in Table 1) was introduced to improve the network model's perception ability of slender defect structures. Eventually, the mAP@0.5 increased by 2.2 percentage points compared with the baseline model, and the mAP@0.5:0.95 increased by 1.1 percentage points.
[0107] Table 1 Results of Ablation Experiment
[0108]
[0109] In summary, the improved method of the YOLOv8 network proposed by the present invention effectively improves the detection accuracy and efficiency of the network model by introducing the VanillaNet, EG-C2f module, dynamic snake-shaped convolution module, and Wise-IoU loss function. The experimental results show that the mAP of the improved algorithm on the NEU-DET dataset is significantly improved compared with the baseline model, while the number of parameters and computational complexity are greatly reduced, and the inference speed is significantly improved. Through the verification of ablation experiments and comparative experiments, the improved network model proposed by the present invention has stronger detection ability for slender and complex-shaped defects and is more suitable for real-time detection requirements in industrial production.
[0110] In one embodiment, before the step of training the novel YOLOv8 network using the training set, the experimental verification of the novel YOLOv8 network on the NEU-DET dataset further includes: performing data augmentation on the training set to train the novel YOLOv8 network using the augmented training set.
[0111] In one embodiment, the data augmentation includes at least, but is not limited to, the following data processing methods: random flipping and / or cropping to improve the generalization ability of the network model.
[0112] It should be noted that the present invention can effectively solve the limitations of the existing YOLOv8 network in industrial applications and make it more suitable for complex defect detection tasks such as strip steel surfaces. The improved method not only improves the detection accuracy but also significantly reduces the demand for computing resources, enabling real-time detection in resource-constrained environments. The present invention is not only applicable to strip steel surface defect detection but can also be widely applied to other surface defect detection tasks with complex geometric shapes and multi-scale features.
[0113] The protection scope of the improved method of the YOLOv8 network described in the embodiments of the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principles of the present invention is included in the protection scope of the present invention.
[0114] In one embodiment, the present invention also provides a method for detecting product surface defects, the method including: detecting product surface defects based on the above improved YOLOv8 network.
[0115] In one embodiment, detecting product surface defects based on the improved YOLOv8 network includes: obtaining a product surface defect image; inputting the product surface defect image into the improved YOLOv8 network so that the improved YOLOv8 network outputs product surface defects.
[0116] In one embodiment, the present invention further provides a product surface defect detection system, which includes: an image acquisition device and the improved YOLOv8 network described above.
[0117] Specifically, the image acquisition device is used to acquire product surface defect images and input the product surface defect images into the improved YOLOv8 network; the improved YOLOv8 network is used to output product surface defects based on the product surface defect images.
[0118] It should be noted that the working principle of the product surface defect detection method and system can refer to the introduction of the improvement method of the YOLOv8 network above, so it will not be elaborated in detail here.
[0119] In several embodiments provided by the present invention, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical or other forms.
[0120] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present invention. For example, in each embodiment of the present invention, the functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.
[0121] Those of ordinary skill in the art should also further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0122] The descriptions of the processes or structures corresponding to the above-mentioned respective drawings each have their own focuses. For the parts not elaborated in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0123] The above embodiments merely illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for improving a YOLOv8 network, characterized in that: The method comprises: The original YOLOv8 network is improved to obtain a new YOLOv8 network; wherein the improvement of the original YOLOv8 network includes: Use VanillaNet to replace the convolutional layer in the Backbone of the original YOLOv8 network; The EG-C2f module is used to replace the three C2f modules in the Neck of the original YOLOv8 network; the three The C2f modules are two C2f modules located at the front and rear positions as outputs in the Neck, and another C2f module other than the output. There are three C2f modules in the Neck as outputs. A dynamic snake convolution module is used to replace the ordinary convolution in the C2f module located in the middle of the Neck of the original YOLOv8 network as the output; Use the Wise-IoU loss function to replace the CIoU loss function in the original YOLOv8 network; The new YOLOv8 network is experimentally verified on the NEU-DET dataset to obtain an improved YOLOv8 network, so as to detect product surface defects based on the improved YOLOv8 network.
2. The improved method of the YOLOv8 network according to claim 1, characterized in that: The EG-C2f module combines GhostConv and ECA attention mechanisms; The GhostConv generates a partial feature map and generates a remaining feature map through linear transformation to eliminate redundant features; The ECA attention mechanism adaptively adjusts the interaction weights between channels through convolution operations.
3. The improved method of the YOLOv8 network according to claim 1, characterized in that: The dynamic serpentine convolution module adaptively adjusts the receptive field of the convolution kernel to ensure that the convolution kernel can adjust its position according to the geometric shape of the product surface defects.
4. The improved method of the YOLOv8 network according to claim 1, characterized in that: The NEU-DET dataset is a standard dataset for product surface defect detection tasks.
5. The improved method of the YOLOv8 network according to claim 1, characterized in that: The NEU-DET dataset includes at least one of the following types of surface defects or a combination of two or more types: cracks, inclusions, plaques, scratches, roller marks, and pitting.
6. The improved method of the YOLOv8 network according to claim 1, characterized in that: The experimental verification of the new YOLOv8 network on the NEU-DET dataset includes: The NEU-DET dataset is divided into a training set, a validation set and a test set; The novel YOLOv8 network is trained, verified and tested using the training set, the validation set and the test set, respectively.
7. The improved method of the YOLOv8 network according to claim 6, characterized in that: Before the step of using the training set to train the new YOLOv8 network, the experimental verification of the new YOLOv8 network on the NEU-DET data set also includes: performing data enhancement on the training set to train the new YOLOv8 network using the data enhanced training set.
8. A method for detecting surface defects of a product, characterized in that: The method comprises: detecting product surface defects based on the improved YOLOv8 network described in any one of claims 1 to 7.
9. The product surface defect detection method according to claim 8, characterized in that: The surface defects of products detected based on the improved YOLOv8 network include: Acquire product surface defect images; The product surface defect image is input into the improved YOLOv8 network so that the improved YOLOv8 network outputs the product surface defects.
10. A product surface defect detection system, characterized in that: The system comprises: an image acquisition device and an improved YOLOv8 network as described in any one of claims 1 to 7; wherein, The image acquisition device is used to acquire product surface defect images, and is used to input the product surface defect images into the improved YOLOv8 network; The improved YOLOv8 network is used to output product surface defects based on the product surface defect image.