Ship industry steel surface defect detection algorithm based on improved YOLOv8

By improving the YOLOv8 model, using the EfficientViT-M2 network to convolution with the SPPELAN module, Lion optimizer and STSConv, the problem of traditional manual detection is solved, and efficient and accurate detection of surface defects of ship steel is achieved.

CN120387990APending Publication Date: 2025-07-29CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510464440.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional manual detection of surface defects of ship steel is inefficient and unstable, making it difficult to meet the needs of large-scale production, and is especially difficult to identify small and complex defects.

Method used

Using the improved YOLOv8 model, the backbone is formed through the EfficientViT-M2 network and the SPPELAN module, combined with the Lion optimizer and the STSConv convolution module, the STS Bottleneck and thin-neck network structure are designed to improve the model robustness and detection accuracy.

Benefits of technology

It realizes efficient and accurate detection of surface defects of ship steel, improves detection accuracy, and meets the production needs of the ship industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387990A_ABST
    Figure CN120387990A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of defect detection, and particularly relates to a ship industry steel surface defect detection algorithm based on improved YOLOv8. According to the method, the ship industry steel surface defect image is obtained for preprocessing operation, the improved YOLOv8 model is trained, and the steel surface defect category and position are obtained. The specific algorithm improvement comprises the following steps: on the basis of an original algorithm, firstly, designing and using an EfficentViT-M2 network and an SPPELAN module to form an EfficentViT-SPPELAN trunk to replace an original YOLOv8 model trunk, so that more context information can be captured on different scales, the accuracy of the model is improved, and meanwhile, the model applies a Lion optimization algorithm to improve the training efficiency and performance of the model; secondly, proposing to design an STSConv convolution module to replace a standard convolution module in an original model neck network; and finally, designing an STS Bottleneck module by using STSConv convolution, and meanwhile, designing an STS module by adopting a thin-neck network structure to replace a C2f module in an original model neck network structure, so that the complexity of calculation and the network structure is reduced. Based on the improved YOLOv8n target detection algorithm, the steel surface defect type and position can be identified and detected, the detection accuracy is improved, and the surface defect detection requirement of the ship industry is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of defect detection, and specifically to a surface defect detection algorithm for shipbuilding industrial steel based on improved YOLOv8. Background Art

[0002] In the pursuit of high-efficiency and high-quality production in the shipbuilding industry, the problem of surface defects in steel has always been a challenge affecting the product quality and safety. Surface defects in steel can not only lead to a decrease in the structural strength of ships, affecting navigation safety, but also trigger a series of production accidents. Therefore, accurate and efficient detection of surface defects in shipbuilding steel has become an urgent problem to be solved in the shipbuilding industry.

[0003] Traditional methods for detecting surface defects in shipbuilding steel mainly rely on manual visual sampling inspection, but this method has many deficiencies. For example, the inspection efficiency is low, making it difficult to meet the needs of large-scale production; it is easily affected by factors such as workers' experience and fatigue, resulting in unstable detection results; it is difficult to accurately identify small and complex defects. Moreover, the detection efficiency and accuracy are difficult to meet production requirements. Therefore, we propose a surface defect detection algorithm for shipbuilding industrial steel based on improved YOLOv8. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] In view of the deficiencies of the prior art, the present invention provides a surface defect detection algorithm for shipbuilding industrial steel based on improved YOLOv8, which solves the problems raised in the above background art.

[0006] (2) Technical Solutions

[0007] The present invention specifically adopts the following technical solutions to achieve the above objectives:

[0008] A surface defect detection algorithm for shipbuilding industrial steel based on improved YOLOv8, comprising:

[0009] Obtain images of surface defects in shipbuilding industrial steel, and use the NEU-DET dataset, which includes 6 types of defect pictures, a total of 1800 defect images. Each picture has been labeled, with 300 defect images of each type;

[0010] Furthermore, perform image preprocessing, randomly divide the dataset into a training set, a validation set, and a test set according to a ratio of 8:1:1. The training set includes 1400 defect pictures, with 240 defect pictures of each type; the validation set and the test set each include 180 defect sample pictures, with 30 defect pictures of each type; expand the dataset image size from 200*200 to 640*640 and input it into the improved YOLOv8 model for training.

[0011] Furthermore, it is proposed to design an EfficientViT-SPPELAN backbone by using the EfficientViT-M2 network and the SPPELAN module to replace the original YOLOv8 model backbone. This backbone network is mainly composed of EfficientViTBlock, sampling layers, and the SPPELAN module. The convolution in it uses a convolutional layer with convolutional operations, BN layers, and ReLU functions. The SPPELAN module uses pyramidal pooling with two convolutions of 1×1 and 5×5 to obtain sub-feature maps of different sizes and combines them to generate larger feature maps. SPP is a spatial pyramidal pooling method that can effectively capture information at different scales, thereby improving the robustness of the model. ELAN is an efficient layer aggregation network that improves the expressive power of the model by effectively fusing features from different layers. By combining SPP and ELAN, the introduction of the SPPELAN module maintains high accuracy. The pyramidal pooling operation enables the algorithm to capture context information at different scales while minimizing information loss and computation, further improving the robustness and generalization ability of the model. Due to the lightweight nature of ELAN, the SPPELAN also helps reduce the computational cost of the model and improve the inference speed.

[0012] Furthermore, the improved model applies the Lion optimizer. Lion introduces the concept of momentum and accelerates gradient updates by accumulating a portion of the historical gradient. During gradient calculation, the learning rate is dynamically adjusted according to the gradient information of each parameter. Then, the loss of the model is calculated through forward propagation, and the gradients of the loss function with respect to each parameter are calculated through backpropagation. This adaptive adjustment makes the model training more stable. The Lion optimizer also introduces a new type of regularization technique. This regularization technique can effectively control the complexity of the model, reduce the risk of overfitting, and also helps improve the generalization ability of the model.

[0013] Furthermore, it is proposed to design an STSConv convolutional module to replace the standard convolutional module in the neck network of the original model. An STSConv module that combines DSConv and Conv is designed to maintain detection accuracy while ensuring detection speed. STSConv uses shuffle to evenly mix the information generated by the standard convolution into each piece of information generated by the depthwise separable convolution, realizing the uniform exchange of feature information between different channels. The specific process is as follows: The input is passed through a normal convolution, and its output is processed by two depthwise convolutions. The results of the outputs of the two depthwise convolution modules are concatenated, and finally, a Shuffle operation is performed to concatenate the channels corresponding to each previous convolution result together. In the STSConv convolutional module, when inputting, a 1×1 convolution is used to convolve the feature map obtained in the previous step and halve the number of channels. Subsequently, it passes through two DSConv convolutions with a convolution kernel of 5. The outputs of the two DSConv convolution modules are combined. Finally, STSConv uses shuffle to evenly mix the information generated by the standard convolution into each piece of information generated by the depthwise separable convolution, realizing the uniform exchange of feature information between different channels, and maintaining detection accuracy while ensuring detection speed.

[0014] Furthermore, an STS Bottleneck module is designed using STSConv convolution. In the STS Bottleneck, a 1×1 convolution is used to convolve the feature map obtained in the previous step to obtain a new feature map representing the weight information of each channel. The input passes through a standard convolution and two STSConv convolutional modules respectively, and the new feature maps output by the STSconv and the 1×1 convolution are weighted to obtain the output result, realizing the information interaction between different channels and reducing the loss of feature information.

[0015] Furthermore, after designing the STS Bottleneck module using STSConv convolution, an STS module is designed using a thin neck network structure. The STS network module is designed based on the STS Bottleneck using a one-time aggregation method. The output features after passing through the standard convolution are respectively input into the standard convolution and the STS Bottleneck module, and the number of channels in each branch is halved. In the STS Bottleneck, a 1×1 convolution is used to convolve the feature map obtained in the previous step to obtain a new feature map representing the weight information of each channel. The new feature map is weighted, and finally, it is multiplied by the original feature map channel by channel to obtain a weighted feature map. This structure can assign higher weights to effective feature channels, suppress irrelevant background features, and reduce the impact of background noise on target detection, thereby enhancing the ability of the network model to judge the location and size of defects. (III) Beneficial effects

[0016] Compared with the prior art, the present invention provides a ship industrial steel surface defect detection algorithm based on improved YOLOv8, having the following beneficial effects:

[0017] In the present invention, by proposing to use the EfficientViT-M2 network and the SPPELAN module to form the EfficientViT-SPPELAN backbone to replace the original YOLOv8 model backbone, it is possible to capture feature information of different sizes, further improving the robustness and generalization ability of the model. Due to the lightweight characteristics of ELAN, SPPELAN also helps to reduce the computational amount of the model and improve the inference speed. Introducing the application of the Lion optimizer makes the model training more stable, with a smaller overfitting risk, and at the same time improves the generalization ability of the model. It is proposed to design an STSConv convolutional module to replace the standard convolutional module in the original model neck network, which can reduce the computational and network structure complexity, and at the same time achieve uniform exchange of feature information of different channels. At the same time, use the STSConv convolution to design the STS Bottleneck module, and adopt the thin neck network structure to combine into the STS module to replace the C2f module in the original model neck network structure, reducing the computational and network structure complexity. Finally, by training the improved YOLOv8 model, the accuracy of ship industrial steel surface defect detection is higher than that of the original YOLOv8 model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of a ship industrial steel surface defect detection algorithm based on improved YOLOv8 of the present invention;

[0019] Figure 2 is a schematic diagram of the improved YOLOv8 backbone structure of the present invention;

[0020] Figure 3 is a schematic diagram of the structure of the SPPELAN module used in the present invention;

[0021] Figure 4 is a schematic diagram of the structure of the STSConv convolution used in the present invention;

[0022] Figure 5 is a schematic diagram of the structure of the STS Bottleneck used in the present invention;

[0023] Figure 6 is a schematic diagram of the structure of the STS module used in the present invention;

[0024] Figure 7 is a simplified diagram of the network structure of the improved YOLOv8 model of the present invention;

[0025] Figure 8 is a visualization schematic diagram of the defect detection results of the models before and after improvement of the present invention. Detailed implementation manners

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] Embodiment

[0028] With the rapid development of technologies such as computer vision and deep learning, new ideas and methods are provided for the detection of steel surface defects. These technologies can achieve accurate identification and positioning of defects through automatic analysis and processing of steel surface images. Compared with traditional manual detection methods, the detection system based on computer vision and deep learning has higher detection accuracy, faster detection speed and stronger stability.

[0029] As Figure 1 shown, a steel surface defect detection algorithm for the shipbuilding industry based on improved YOLOv8 proposed in an embodiment of the present invention includes the following steps:

[0030] Step 1: Obtain steel surface defect images of the shipbuilding industry and preprocess the pictures by methods such as size transformation, classification and annotation.

[0031] Step 2: Divide the processed pictures into a training set and a validation set and input them into the improved YOLOv8 steel surface defect detection algorithm model for training.

[0032] Step 3: Output steel surface defect pictures after training, which contain steel defect types and defect location information.

[0033] Among them, as Figure 2 shown, in some embodiments, the process of feature extraction of the backbone network of the improved YOLOv8 model is specifically as follows:

[0034] When the processed ship steel defect pictures are input into the improved YOLOv8 backbone, first, an overlapping embedding layer is introduced at the input position to enhance the model's ability in low-vision learning. The entire framework uses three-scale stages, and each stage adopts a different number of EfficientViT modules. A downsampling layer composed of a linear layer and an inverted residual module is introduced between different stages to reduce the loss of feature information. During input, the overlapping embedding layer is used for 16-fold downsampling. After passing through the EfficientViT Block modules and sampling modules at different stages, different numbers of sandwich structures are used in the three stages respectively. By continuously adjusting the parameters of each stage, the most suitable combination is finally selected, that is, L = 1, M = 2, N = 3, and this combination can reduce the loss of information. Finally, region extraction and feature extraction are performed through the Spatial Pyramid Pooling Enhanced Local Attention Network (SPPELAN). To improve the model inference efficiency, Batch Normalization and ReLU activation functions are uniformly used. Among them, the EfficientViT Block consists of 2N depthwise separable convolutions (DWConv), FFN, and a cascaded attention module (CGA), which improve the model detection efficiency from three aspects: memory, calculation, and parameters. The MHSA layer is embedded between the FFN feed-forward network layers, and multiple FFN feed-forward network layers are applied. This structure not only reduces the time overhead caused by memory limitation in MHSA but also uses multiple FFN feed-forward network layers to achieve communication between different channels.

[0035] As Figure 3 shown, in some embodiments, the specific operation of using the SPPELAN module in the improved YOLOv8 model backbone network is as follows:

[0036] The input features first pass through a standard convolution. The outputs of the standard convolution are respectively input into Concat and Maxpool2d (maximum pooling layer). The outputs of the maximum pooling layer are respectively input into Concat and the next Maxpool2d. The output of the last Maxpool2d is input into Concat. Finally, the output features of this module are obtained through a standard convolution.

[0037] As Figure 4 shown, in some embodiments, the training process of the STSConv convolution module in the improved YOLOv8 model network is as follows:

[0038] During input, a 1×1 convolution is used to convolve the feature map obtained in the previous step, and the number of channels is halved. Subsequently, two DSConv convolutions with a convolution kernel of 5 are performed. The outputs of the two DSConv convolution modules are combined. Finally, STSConv uses shuffle to evenly mix the information generated by the standard convolution into each piece of information generated by the depthwise separable convolution, achieving uniform exchange of feature information between different channels, while maintaining detection accuracy while ensuring the detection speed.

[0039] As Figure 5 shown, in some embodiments, the training process of the STS Bottleneck structure of a shipbuilding industrial steel surface defect detection algorithm based on improved YOLOv8 is as follows:

[0040] The input passes through a standard convolution and two STSConv convolution modules respectively. The 1×1 standard convolution convolves the feature map obtained in the previous step, and a weighting operation is performed on the new feature maps output by STSConv and the 1×1 convolution to obtain a new feature map representing the weight information of each channel.

[0041] As Figure 6 shown, in some embodiments, the training process of the STS network structure of a shipbuilding industrial steel surface defect detection algorithm based on improved YOLOv8 is as follows:

[0042] The STS network module is designed using a one-time aggregation method based on the STS Bottleneck. The output features after passing through the standard convolution are respectively input into the standard convolution and the STS Bottleneck module. The number of channels in each branch is halved, and finally, it is multiplied with the original feature map channel by channel to obtain a weighted feature map, assigning higher weights to effective feature channels, suppressing irrelevant background features, and reducing the impact of background noise on target detection, thereby enhancing the ability of the network model to judge the location and size of defects.

[0043] As Figure 7 shown, in some embodiments, the training process of a structural schematic diagram of a shipbuilding industrial steel surface defect detection algorithm based on improved YOLOv8 is as follows:

[0044] The improved YOLOv8 model mainly consists of an input, a backbone, a neck, a head, and an output. The input part first preprocesses the obtained shipbuilding industrial steel surface defect detection pictures, expands them to a size of 640*640, and then inputs them into the backbone network; after performing the above Figure 2 shown steps to extract features from the input defect images in the backbone network, they are input into the neck network; in the neck network, the extracted features are sampled, combined, and Figure 4 and Figure 6The module shown performs feature fusion to obtain feature information of different scales; then, the fused features are detected through the model detection head; finally, the output results include information such as the type and location of steel surface defects. During the entire training process, the Lion optimizer is used to optimize the improved algorithm, and in the entire model network, the standard convolution applies a convolutional layer with convolutional operations, a BN layer, and a SiLU function.

[0045] As Figure 8 shown, in some embodiments, the visualization results of the detection of steel surface defects in the shipbuilding industry before and after the improvement of the YOLOv8 model:

[0046] Since the foreground and background of the steel surface defect dataset in the shipbuilding industry are similar and the detection is difficult, the original YOLOv8 model is not accurate enough in detecting and locating steel defects in the shipbuilding industry, and there are problems such as missed detections and repeated detections, and it cannot effectively detect steel surface defects; the improved model is more accurate in locating steel surface defects in the shipbuilding industry, and the detection accuracy is significantly improved, which can solve the problems encountered in the above background technology and better meet the requirements of the shipbuilding industry for steel defect detection.

[0047] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An algorithm for detecting surface defects of shipbuilding industrial steel based on improved YOLOv8, characterized in that, Including: (1) Obtain the surface defect images of shipbuilding industry steel. (2) Image preprocessing. (3) Design to use the EfficientViT-M2 network and the SPPELAN module to form the EfficientViT-SPPELAN backbone to replace the original YOLOv8 model backbone, and introduce the Lion optimizer. (4) Propose to design an STSConv convolution module to replace the standard convolution module in the original model's neck network, and perform feature fusion on the extracted features to obtain features of different sizes. (5) Use the STSConv convolution to design the STS Bottleneck module, and adopt the thin neck network structure to design the STSConv and STS Bottleneck as the STS module to replace the C2f module in the original model's neck network structure, reducing the computational and network structure complexity.

2. The surface defect detection algorithm for shipbuilding industrial steel based on the improved YOLOv8 according to claim 1, characterized in that: For method step (1), the public dataset NEU-DEF of Northeastern University is used to obtain the surface defect images of shipbuilding industry steel. This dataset contains six types of defect images commonly found in industry, namely scratches (Sc), crazing (Cr), rolled-in scale (RS), pitted surface (PS), patches (Pa), and inclusion (In). There are 300 images for each type of defect, a total of 1800 defect images. Each dataset provides the annotation information of the defects. NEU-DEF contains a complete range of defect types, and the various types of defects are similar, which poses a certain challenge to the recognition and localization of defect types.

3. An algorithm for detecting surface defects of shipbuilding industrial steel based on improved YOLOv8 according to claim 1, characterized in that: For method step (2), the 1800 defect images in the dataset are randomly divided into a training set, a validation set, and a test set according to the ratio of 8:1:

1. The training set contains 1440 defect sample images, with 240 images for each type of defect; the validation set and the test set each contain 180 defect sample images, with 30 images for each type of defect. Finally, the 200*200 steel surface defect images are expanded to a size of 640*640 and then input into the improved model.

4. An algorithm for detecting surface defects of shipbuilding industrial steel based on improved YOLOv8 according to claim 1, characterized in that: Method step (3) proposes to design an EfficientViT-SPPELAN backbone by combining the EfficientViT-M2 network and the SPPELAN module to replace the original YOLOv8 model backbone. This backbone network is mainly composed of EfficientViTBlock, sampling layers, and the SPPELAN module. The convolution in it uses a convolutional layer with convolutional operations, BN layers, and ReLU functions. The SPPELAN module uses two convolutions of 1×1 and 5×5 in the pyramid pooling to obtain sub-feature maps of different sizes, and combines them to generate a larger feature map. In the backbone, SPPELAN is a combination of SPP and ELAN. The purpose of this module is to combine the advantages of both to improve the detection effect of the model. ELAN is a lightweight network structure that effectively improves the model's feature extraction ability through local aggregation and global integration. SPP is a spatial pyramid pooling method that can effectively capture information at different scales, thereby improving the robustness of the model. The introduction of SPPELAN maintains high accuracy. The pyramid pooling operation enables the algorithm to capture context information at different scales while minimizing information loss and computation, further improving the robustness and generalization ability of the model. Due to the lightweight characteristics of ELAN, SPPELAN also helps to reduce the computational cost of the model and improve the inference speed.

5. An algorithm for detecting surface defects of shipbuilding industrial steel based on improved YOLOv8 according to claim 1, characterized in that: Method step (3) introduces the use of the Lion optimizer. Lion introduces the concept of momentum and accelerates gradient updates by accumulating a part of the historical gradient. In gradient calculation, the learning rate is dynamically adjusted according to the gradient information of each parameter. For parameters with larger gradients, the Lion optimizer tends to use a smaller learning rate to avoid oscillations in the model during training; for parameters with smaller gradients, a larger learning rate is used to accelerate convergence. Lion treats each component equally through the sign operation, enabling the model to fully utilize the role of each component. Its update process is as follows: where, is the gradient of the loss function, and sign is the sign function, i.e., 1 for positive numbers and -1 for negative numbers.

6. The surface defect detection algorithm for shipbuilding industrial steel based on the improved YOLOv8 according to claim 1, wherein: Method step (4) proposes to design an STSConv convolutional module to replace the standard convolutional module in the original model's neck network. An STSConv module that combines DSConv and Conv is designed to maintain detection accuracy while ensuring detection speed. STSConv uses shuffle to evenly mix the information generated by the standard convolution into each piece of information generated by the depthwise separable convolution, achieving uniform exchange of feature information in different channels. The specific process is as follows: The input is passed through an ordinary convolution, and its output is processed by two depthwise convolutions. The results of the outputs of the two depthwise convolution modules are concatenated, and finally, a Shuffle operation is performed to concatenate the channels corresponding to each convolution result before. Its purpose is to make the result of the depthwise convolution more similar to the standard convolution.

7. An algorithm for detecting surface defects of shipbuilding industrial steel based on improved YOLOv8 according to claim 1, characterized in that: In method step (5), after using STSConv convolution to design the STS Bottleneck module, a neck network structure is adopted to combine it into an STS module to replace the C2f module in the neck network structure of the original model. The output features after standard convolution are respectively input into the standard convolution and the STS Bottleneck module, and the number of channels in each branch is halved. In the STS Bottleneck, a 1×1 convolution is used to convolve the feature map obtained in the previous step to obtain a new feature map representing the weight information of each channel. A weighting operation is performed on the new feature map, and finally, it is multiplied by the original feature map channel by channel to obtain a weighted feature map. This structure can assign higher weights to effective feature channels, suppress irrelevant background features, and reduce the impact of background noise on target detection, thereby enhancing the ability of the network model to judge the location and size of defects.

8. The STS Bottleneck module designed using the STSConv convolution according to claim 7, wherein: In the STS Bottleneck, a 1×1 convolution is used to convolve the feature map obtained in the previous step. A weighting operation is performed on the new feature maps output by STSconv and the 1×1 convolution to obtain a new feature map representing the weight information of each channel.

Citation Information

Cited By

  • Welding seam surface defect detection method based on DAB-YOLO algorithm

    CN120953260A

  • A welding seam surface defect detection method based on a DAB-YOLO algorithm

    CN120953260B