Metal surface defect detection method based on improved YOLOv8

By improving the YOLOv8 network, introducing ECA, CBAM and ShuffleNet V2 modules, and using the Inner-IoU loss function, the problems of low detection accuracy and large parameters in the existing metal surface defect detection methods are solved, achieving more efficient and accurate defect detection.

CN120031798APending Publication Date: 2025-05-23YICHUN JIANGLI LITHIUM BATTERY NEW ENERGY IND RES INST
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411991144.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing metal surface defect detection methods have problems such as low detection accuracy, large parameter quantity, high error detection rate and insufficient feature extraction capabilities, and it is difficult to achieve a good balance between detection accuracy and efficiency.

Method used

By improving the YOLOv8 network, the ECA attention mechanism, the CBAM attention module and the ShuffleNet V2 module were introduced, and a new metal surface defect detection model was constructed using the Inner-IoU loss function.

Benefits of technology

The model's detection accuracy of metal surface defects is improved, the parameter quantity and calculation overhead are reduced, the feature extraction ability is enhanced, and the false detection and missed detection rates are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031798A_ABST
    Figure CN120031798A_ABST
Patent Text Reader

Abstract

The invention discloses a metal surface defect detection method based on improved YOLOv8, and belongs to the technical field of metal surface defect detection.The metal surface defect detection method based on the improved YOLOv8 comprises the steps that a metal surface image is collected, defects in the image are marked, and a sample data set is constructed; the method comprises the following steps of: improving a YOLOv8 network to obtain an improved YOLOv8 network; training the improved YOLOv8 network on the basis of the sample data set; and realizing metal surface defect detection by utilizing the trained and improved YOLOv8 network. According to the scheme, the model parameter quantity is reduced, the model detection precision is improved, and the false detection probability is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of metal surface defect detection, and in particular to a metal surface defect detection method based on improved YOLOv8. Background Art

[0002] With the development of industry, metal materials are increasingly used in manufacturing, involving aerospace, transportation, machinery and other fields. However, in actual production and processing, factors such as material quality, processing equipment and environment will lead to various types of defects on the metal surface, including scratches, cracks and oxidation. These defects have a great impact on the product performance and service life of metal products, and cause great economic losses. In recent years, with the continuous improvement and updating of detection technology, the field of metal surface defect detection has received more and more attention. Metal surface defect detection can check the quality of metal, find existing defects and reduce production costs.

[0003] Traditional metal surface defect detection mainly relies on manual visual inspection, which has disadvantages such as low detection efficiency and strong subjectivity. There is no strict definition of metal surface defects. In order to solve these problems, scholars engaged in the metal industry have proposed some automated methods, such as optical inspection, magnetic particle inspection and eddy current inspection. Among them, magnetic particle inspection can only detect metals with magnetism, but cannot detect defects in non-magnetic metals, and it is easily disturbed by the external environment and causes false alarms.

[0004] With the continuous development of deep learning technology, metal defect detection technology based on deep learning has been widely used in actual processing and production. It has the advantages of fast speed and good versatility. It can realize large-scale defect detection and is gradually replacing traditional metal defect detection methods. Convolutional neural networks are highly praised for their powerful feature extraction capabilities. Many scholars apply deep learning technology to metal surface defect detection. The metal defect detection method based on deep learning is mainly divided into one stage and two stages. The main one-stage defect detection algorithms are YOLO and SSD (Single Shot MultiBox Detector). SSD is similar to YOLO in running speed and has the problem of insufficient feature extraction. Unlike the first stage, the second stage algorithm is to generate the required area first, and then predict and classify the corresponding area. Representative algorithms include R-CNN, Fast R-CNN and U-Net.

[0005] The You Only Look Once (YOLO) algorithm proposed by Redmon et al. directly returns the object and its class bounding box by dividing the grid, and has a high detection speed. Zhou et al. integrated the attention mechanism module into the YOLOv5s model to improve the detection performance, but the detection accuracy was reduced. Zhang et al. combined the lightweight convolution layer GSConv with YOLOv5s to improve the detection rate of strip surface defects at the cost of reducing detection accuracy. Although these studies have made certain contributions to the field of metal surface defect detection, these models have not yet achieved a good balance between detection accuracy and efficiency. In practical applications, the detection of metal surface defects will be affected by factors such as overexposure and uneven brightness. Therefore, in actual strip surface defect detection applications, the existing metal surface defect detection models still cannot overcome these challenges and achieve a balance between detection accuracy and detection speed.

[0006] In summary, the existing metal surface defect detection methods have the following defects:

[0007] 1. The model has a weak perception of subtle defects on the metal surface, low detection accuracy, and little overlap between the predicted frame and the true frame;

[0008] 2. The model has a large number of parameters, resulting in high operating costs;

[0009] 3. False detection and missed detection rate are high, and non-defective parts are often detected as defective, or defective parts are detected as non-defective;

[0010] 4. The ability to extract features is low and the ability to represent features is not strong. Summary of the invention

[0011] The present invention provides a metal surface defect detection method based on improved YOLOv8, so as to at least solve one of the above-mentioned technical problems existing in the prior art to a certain extent.

[0012] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0013] On the one hand, the present invention provides a metal surface defect detection method based on improved YOLOv8, and the metal surface defect detection method based on improved YOLOv8 includes:

[0014] Collect metal surface images and annotate the defects in the images to build a sample data set;

[0015] Improve the YOLOv8 network to obtain an improved YOLOv8 network;

[0016] Training the improved YOLOv8 network based on the sample data set;

[0017] The trained improved YOLOv8 network is used to detect metal surface defects.

[0018] Furthermore, the YOLOv8 network is improved to obtain an improved YOLOv8 network, including:

[0019] The ECA mechanism is introduced before SPPF in the Backbone of the YOLOv8 network.

[0020] Furthermore, the improving the YOLOv8 network to obtain an improved YOLOv8 network also includes:

[0021] The channel and spatial attention module CBAM is integrated into the SPPF of the YOLOv8 network.

[0022] Furthermore, the improving the YOLOv8 network to obtain an improved YOLOv8 network also includes:

[0023] Replace the C2f module in the Backbone of the YOLOv8 network with the ShuffleNet V2 module.

[0024] Furthermore, the improving the YOLOv8 network to obtain an improved YOLOv8 network also includes:

[0025] The Inner-IoU loss function is used as the loss function of the improved YOLOv8 network.

[0026] Furthermore, after training the improved YOLOv8 network based on the sample data set, the method further includes:

[0027] The improved YOLOv8 network is evaluated using preset indicators.

[0028] Furthermore, the preset indicator is a combination of one or more of accuracy, recall rate, average precision and model parameter quantity.

[0029] Furthermore, the improved SPPF includes: a first convolutional layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a first CBAM module, a second CBAM module, a third CBAM module and a fifth convolutional layer.

[0030] Furthermore, the convolution kernels of the first convolution layer and the fifth convolution layer are 1×1 convolution kernels; and the convolution kernels of the second convolution layer, the third convolution layer and the fourth convolution layer are 3×3 convolution kernels.

[0031] Furthermore, the improved SPPF processes the input data as follows:

[0032] The input data is input into the first convolution layer, the processing result of the first convolution layer is input into the first maximum pooling layer, the processing result of the first maximum pooling layer is input into the second maximum pooling layer and the second convolution layer, the processing result of the second maximum pooling layer is input into the third maximum pooling layer and the third convolution layer, and the processing result of the third maximum pooling layer is input into the fourth convolution layer; the processing result of the second convolution layer is input into the first CBAM module, the processing result of the third convolution layer is input into the second CBAM module, and the processing result of the fourth convolution layer is input into the third CBAM module; after the processing results of the first CBAM module, the second CBAM module and the third CBAM module are fused, they are input into the fifth convolution layer after the Concat operation on the channel dimension with the processing result of the first convolution layer.

[0033] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0034] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the above method.

[0035] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0036] 1. The parameters of the metal surface defect detection model proposed in the present invention are reduced compared with the existing model, and the calculation overhead is effectively controlled. When the parameters are equivalent, the detection accuracy of the model of the present invention is improved;

[0037] 2. The model of the present invention has more accurate anchor frames for defective areas and clearer boundaries;

[0038] 3. The model of the present invention detects different types of defects more accurately, reducing the probability of false detection;

[0039] 4. The metal surface defect detection model proposed in the present invention has enhanced feature extraction capability, improved characterization capability, and improved model perception of subtle defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0041] Figure 1 It is a schematic diagram of the execution flow of the metal surface defect detection method based on the improved YOLOv8 provided in an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of an improved network structure provided by an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the structure of an improved SPPF provided by an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the attention calculation process of CBAM provided by an embodiment of the present invention;

[0045] Figure 5 It is a schematic diagram of the ECA module structure provided by an embodiment of the present invention;

[0046] Figure 6 It is a schematic diagram of the basic structure and calculation process of ShuffleNet V2 provided by an embodiment of the present invention.

[0047] Figure 7 is a performance comparison chart of different models provided by an embodiment of the present invention;

[0048] Figure 8 It is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0050] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0051] First embodiment

[0052] This embodiment provides a metal surface defect detection method based on improved YOLOv8, which can be implemented by an electronic device. The execution process of the method is as follows: Figure 1 As shown, the following steps are included:

[0053] S1, collect metal surface images and annotate the defects in the images to build a sample data set;

[0054] S2, improve the YOLOv8 network to obtain an improved YOLOv8 network;

[0055] S3, training the improved YOLOv8 network based on the sample data set;

[0056] S4, using the trained improved YOLOv8 network to realize metal surface defect detection.

[0057] Specifically, in this embodiment, the scheme for improving the YOLOv8 network is as follows:

[0058] In order to improve the model's ability to extract complex defects on metal surfaces, the SPPF (spatial pyramid pooling layer) module in Backbone was first optimized, and the CBSPPF module was designed, and the CBAM (channel and spatial attention module) was integrated into the SPPF. Through the dual attention mechanism of CBAM, the model can extract features more accurately from the channel and spatial dimensions, and adaptively adjust the weights of the feature map to improve the ability to capture defects of different sizes, expand the model's receptive field, and improve the accuracy of small target detection. In order to further enhance the model's feature representation ability, the ECA (Efficient Channel Attention) mechanism was introduced in Backbone. ECA captures the dependencies between different channels through efficient channel interactions, strengthens the model's attention to important features, and realizes efficient channel attention allocation, significantly reducing the interference of redundant information on feature extraction. In addition, in order to achieve a lightweight design of the model without losing detection performance, this paper replaces the C2f module in Backbone with the ShuffleNet V2 module. ShuffleNet V2 is an efficient and lightweight network that reduces model parameters through deep separable convolution and channel mixing strategies and adopts a new information flow shuffling method. It is suitable for real-time defect detection tasks in industrial environments. Finally, the Inner-IoU loss function is used as the loss function of the improved network. Based on the above optimization, this embodiment proposes a metal surface defect detection model based on improved YOLOv8. The model integrates the efficient and lightweight design of ShuffleNet V2 in the backbone network, and significantly improves the performance of the model in complex metal surface defect detection by introducing the improved SPPF module and ECA attention mechanism. The improved network structure is as follows Figure 2 shown.

[0059] Below, various improvement methods are described in detail one by one.

[0060] 1.SPPF (Spatial Pyramid Pooling Layer)

[0061] SPPF is a commonly used module in convolutional neural networks. It aims to extract multi-scale features through pooling operations and enhance the model's ability to handle complex image tasks. However, the original SPPF module uniformly processes the feature maps of all input regions and lacks the distinction of the importance of different regions. When there are defects of multiple sizes in the image, it is difficult for the model to effectively handle these complexities, resulting in the receptive field being unable to fully adapt to multi-size defect detection tasks, and it is especially difficult to focus on key defect areas. To this end, we optimized the SPPF module and integrated CBAM (Convolutional Block Attention Module) with SPPF, enabling it to dynamically allocate more computing resources to key areas and ignore irrelevant or minor parts.

[0062] The improved SPPF structure is as follows: Figure 3 As shown in the figure, it integrates the CBAM attention mechanism, which can dynamically adjust the feature extraction strength of different regions, effectively overcoming the limitations of the original SPPF in multi-scale feature extraction, especially showing significant advantages in small target defect detection.

[0063] After integrating the CBAM attention mechanism, the model can achieve more refined attention to different features in the image through the dual mechanisms of channel attention and spatial attention. The channel attention mechanism can dynamically adjust the weights of each feature channel, so that the model prioritizes the channel information that is most important for metal defect detection and suppresses noise features. At the same time, the spatial attention mechanism enables the model to more accurately locate key metal surface defect areas in different spatial dimensions by capturing important spatial positions in the image. This combination not only significantly reduces the interference from background noise, but also enhances the model's learning ability when dealing with complex defects, allowing the model to show higher accuracy and generalization ability in a variety of defect detection tasks. Figure 4 In , we show the attention calculation process of CBAM, where channel attention generates different weights through global average pooling and maximum pooling, and spatial attention captures key spatial regions through convolution operations.

[0064] 2.ECA Attention Mechanism

[0065] In the metal defect detection task, the feature extraction capability of the model is crucial to the final detection results. As a lightweight and efficient target detection framework, YOLOv8 has shown strong detection capabilities in practical applications. However, when dealing with tiny and complex metal defect features, the standard backbone architecture of YOLOv8 may not be able to fully capture the importance differences between channels, and some key features may be ignored, resulting in a decrease in the accuracy of the detection results, especially in the identification of subtle defects. To solve this problem, we introduced the Efficient Channel Attention (ECA) attention mechanism before the SPPF module in the backbone of YOLOv8.

[0066] The ECA attention mechanism is a lightweight channel attention mechanism that can effectively assign weights to each channel through simple 1D convolution operations with almost no increase in model complexity. Compared with other attention mechanisms, ECA avoids the use of fully connected layers, thereby reducing computational overhead while retaining sensitivity to important features. This efficient design ensures that ECA can be easily integrated into existing detection frameworks, improving the model's ability to capture inter-channel relationships while maintaining lightweight. By introducing the ECA attention mechanism in the YOLOv8 backbone, the model can better distinguish between key features and noise, improve sensitivity to minor defects, and significantly improve overall detection accuracy and model robustness, enabling it to perform better when dealing with complex detection tasks.

[0067] Specifically, the ECA module automatically adjusts the model's focus on each channel by adding weights to each channel of the feature map. This approach ensures that the model can more accurately extract the most critical feature information when dealing with defects of various sizes and shapes, especially in the detection of subtle flaws in metal defect detection. The ECA module structure is as follows Figure 5 shown.

[0068] 3.ShuffleNet

[0069] In order to achieve lightweight model, this paper replaces all C2f modules in the YOLOv8 backbone network with ShuffleNet V2 modules. ShuffleNet V2 is a lightweight convolutional neural network architecture that uses channel rearrangement operations to reduce the amount of computation and parameters. It divides the input channels into multiple groups and rearranges the channels of each group to facilitate information exchange between components. ShuffleNet V2 effectively reduces the computational complexity and model size by combining depthwise separable convolution and channel mixing technology. Depthwise separable convolution is divided into two steps: depthwise convolution and pointwise convolution. This design greatly reduces the number of parameters required and reduces the computational cost, while channel mixing further improves the network's expressiveness and computational efficiency through rearrangement operations. Depthwise separable convolution decomposes the convolution operation into two consecutive steps: depthwise convolution and pointwise convolution. This decomposition greatly reduces the number of parameters involved and the computational cost. Through these optimized designs, ShuffleNet V2 can significantly reduce the model inference time while maintaining the high accuracy of the model. Especially on resource-constrained hardware (such as mobile devices or embedded systems), ShuffleNet V2 exhibits extremely high inference speed and low computational burden, making it an ideal choice for lightweight models.

[0070] like Figure 6 As shown on the left side of the figure, the basic structural module of ShuffleNet V2 includes two parallel branches. This structure enhances the parallel computing capability of the model by dividing the input feature map into two parts and performing different computing operations on each branch. In the left branch, a 3×3 depth convolution operation is first performed on the input feature map. This step is used to extract local spatial features. Subsequently, a 1×1 point-by-point convolution is performed on the processed feature map. The point-by-point convolution is used to integrate information from different channels and further reduce the amount of calculation. Next, the outputs of the two branches are merged through the Concat operation on the channel dimension to form a more representative feature map. Next, a channel shuffle operation is performed. By rearranging the channel order, the information is ensured to be evenly distributed among different channels, and redundant calculations are avoided, thereby improving the network's expression ability and feature extraction efficiency. The specific process of channel shuffling is shown as follows. Figure 6 The right half of the diagram is shown in Figure 1.

[0071] 4. Inner-IoU loss function

[0072] Inner-IoU is an indicator that measures the degree of overlap between the predicted box and the true box in the target detection task. In this project, the loss function of the YOLOv8 model was improved, and the original CIoU (Complete Intersection over Union) loss was replaced by Inner-IoU (Inner Intersection over Union) loss. CIoU is used to optimize the bounding box matching of the target detection model. However, when it deals with scenes with complex backgrounds or blurred boundaries, the matching may be inaccurate due to the influence of background noise. To solve this problem, this embodiment uses Inner-IoU to replace CIoU. Inner-IoU focuses on calculating the overlapping area between the inside of the predicted box and the true box, ignoring the background area outside the predicted box. This improvement effectively reduces the interference of the background and improves the accuracy of target detection in complex scenes. In model training, the Inner-IoU loss optimizes the positioning accuracy of the bounding box by only calculating the overlap of the internal area of ​​the prediction box, which is particularly suitable for detection tasks with complex backgrounds or unclear target boundaries. The formula of Inner-IoU is as follows:

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] union=(w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter (6)

[0079]

[0080] L inner -IoU=1-IoU inner (8)

[0081] Next, the effectiveness of the metal surface defect detection model constructed in this embodiment is verified.

[0082] 1. Comparative Experiments on Different Attention Mechanisms

[0083] This experiment compares the application of different attention mechanisms in the backbone of YOLOv8, using CA, SE, CBAM and ECA attention modules respectively, and tests the performance of these modules in steel plate surface defect detection. The results show that the CBSPPF-ECA-Shuffle scheme using ECA attention performs well in multiple indicators, especially in the two types of defects, Inclusion, Rolled-in Scale and Scrathes, reaching 86.2%, 77.4% and 85.1% mAP@0.5 respectively, which is significantly better than other schemes. At the same time, ECA attention also performs well in comprehensive detection performance indicators, with mAP@0.5 reaching 77.9 and R% being 78.0, both of which are the highest values. This shows that the ECA attention mechanism can more effectively enhance the feature extraction capability, especially when dealing with complex surface defects, it can better capture the local details and global information of the target, thereby improving the detection accuracy of the model. In contrast, although CA, SE and CBAM attention also perform well in some indicators, they fail to surpass the performance of ECA overall. Therefore, ECA attention performs best in this experiment and is suitable as the backbone improvement solution for the YOLOv8 model. The experimental results are shown in Table 1.

[0084] Table 1 Performance of different attention modules in steel plate surface defect detection

[0085]

[0086] 2. Model comparison experiment

[0087] In order to verify the effectiveness of the improved model, the model was compared with other models in the YOLO series, and P% (precision), R% (recall), mAP@0.5% (average precision) and the number of parameters of the model (Params / M) were examined. In terms of recall, our model reached 78.0%, which is significantly better than other YOLO series models, indicating that our model has better results in detecting positive samples, can identify more targets, and improves the practical application ability of the model. Secondly, our model achieved an accuracy of 73.6%, which is comparable to YOLOv8n's 74.5%, and significantly better than YOLOv7's 69.3%. In terms of mAP@0.5%, our model reached 77.9%, showing that our model has good detection accuracy at a higher intersection-over-union ratio. This further illustrates the superiority of our model in balancing precision and recall.

[0088] In terms of the number of parameters (Params / M), our model has only 3.0M parameters, which is significantly smaller than the number of parameters of models such as SSD and YOLOv8s. Compared with SSD's 21.5M, our model has significantly reduced parameters while still maintaining good detection performance, which shows that the lightweight design of the model has great advantages for practical deployment and application. Overall, our improved model has significant advantages in recall rate and model complexity, while maintaining a high level of accuracy and average precision. At the same time, the results of six defects of YOLOv8s, YOLOv8n and the improved model on the NEU-DET dataset are compared in detail, and the results are shown in Table 2.

[0089] Table 2 Results of different models on six defects of NEU-DET dataset

[0090]

[0091] 3. Ablation Experiment

[0092] In the ablation experiment, we compared the effects of different module combinations on target detection performance. Table 3 shows the detection effects of using CBSPPF, ECA, ShuffleNet, and Inner-IoU in YOLOv8. When all components are added, the detection performance is optimal, with a mAP@0.5 value of 77.9, an R% of 78.0, and a parameter amount (Params / M) of 3.0M. This shows that ECA can enhance feature extraction capabilities, while the use of Inner-IoU loss improves the matching accuracy between the predicted box and the target. In contrast, the model without ECA and Inner-IoU performs better in some indicators, but its overall performance is slightly inferior. In addition, the comparison of parameter amounts shows that although the parameter amounts of different combinations vary, they are all within a reasonable range and do not significantly affect the complexity of the model. Therefore, the ablation experiment results verify the effectiveness of the model.

[0093] Table 3 Detection effects of different module combinations in YOLOv8

[0094]

[0095]

[0096] 4. Experimental Visualization

[0097] In order to verify the effectiveness of the improved model, the model is compared with models such as YOLOv8s and YOLOv8n. Figure 7 As shown in the comparative experiment diagram, the improved model of the present invention is compared with the existing YOLOv5s and YOLOv8n models in terms of performance. Figure 7 It can be seen that the detection effect of this model is closer to Ground Truth, especially in detail recognition, bounding box positioning and the accuracy of the number of targets. In contrast, YOLOv8s and YOLOv8n have missed detection or false detection in some areas, while this model avoids these problems. These results verify the robustness and detection accuracy of the improved model in complex scenes, indicating that our optimization method has practical value in improving target detection performance.

[0098] In summary, this embodiment provides a metal surface defect detection method based on improved YOLOv8. By improving the YOLOv8 network, a new metal surface defect detection model is constructed, which solves the limitations of the original feature pyramid (SPPF) in the YOLO model in multi-scale feature extraction; for important information features, higher weights can be assigned to information-rich parts, and in the final segmentation map, the edges of the defect segmentation map are made clearer; and while maintaining high precision, the number of model parameters is effectively reduced.

[0099] Second embodiment

[0100] This embodiment provides an electronic device, such as Figure 8 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment. In addition, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0101] Next, combine Figure 8 The following is a detailed introduction to the various components of the electronic device:

[0102] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor may execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0103] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 8 The CPU0 and CPU1 shown in the figure are, of course, only exemplary.

[0104] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0105] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 8 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0106] The transceiver may include a receiver and a transmitter ( Figure 8 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 8 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0107] In addition, it should be noted that Figure 8 The structure of the electronic device shown in the figure does not constitute a limitation on the device, and the actual device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they are not repeated here.

[0108] Third embodiment

[0109] This embodiment provides a computer-readable storage medium, which stores at least one instruction, and the instruction is loaded and executed by a processor to implement the method of the first embodiment. The computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein may be loaded by a processor in a terminal to execute the method.

[0110] In addition, it should be noted that the present invention can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present invention can be in the form of a full or partial hardware embodiment, a full or partial software embodiment or an embodiment combining software and hardware. Moreover, when implemented using software, the embodiment of the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more available media sets. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state hard disk.

[0111] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0112] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0113] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. In addition, the term "and / or" is only an association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc or abc, where a, b, c can be single or plural.

[0114] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0115] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0116] In several embodiments provided by the present invention, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of functional modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0117] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0118] Finally, it should be noted that the above is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiment of the present invention has been described, for ordinary technicians in this technical field, once the basic creative concept of the present invention is known, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. Therefore, the attached claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A metal surface defect detection method based on improved YOLOv8, characterized in that: include: Collect metal surface images and annotate the defects in the images to build a sample data set; Improve the YOLOv8 network to obtain an improved YOLOv8 network; Training the improved YOLOv8 network based on the sample data set; The trained improved YOLOv8 network is used to detect metal surface defects.

2. The metal surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The improved YOLOv8 network is improved to obtain an improved YOLOv8 network, including: The ECA mechanism is introduced before SPPF in the Backbone of the YOLOv8 network.

3. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 2, characterized in that: The improved YOLOv8 network is improved to obtain an improved YOLOv8 network, further comprising: The channel and spatial attention module CBAM is integrated into the SPPF of the YOLOv8 network.

4. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 3, characterized in that: The improved YOLOv8 network is improved to obtain an improved YOLOv8 network, further comprising: Replace the C2f module in the Backbone of the YOLOv8 network with the ShuffleNet V2 module.

5. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 4, characterized in that: The improved YOLOv8 network is improved to obtain an improved YOLOv8 network, further comprising: The Inner-IoU loss function is used as the loss function of the improved YOLOv8 network.

6. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 1, characterized in that: After training the improved YOLOv8 network based on the sample data set, the method further includes: The improved YOLOv8 network is evaluated using preset indicators.

7. The metal surface defect detection method based on improved YOLOv8 according to claim 6, characterized in that: The preset indicator is a combination of one or more of accuracy, recall rate, average precision and model parameter quantity.

8. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 3, characterized in that: The improved SPPF includes: a first convolution layer, a first maximum pooling layer, a second maximum pooling layer, a third maximum pooling layer, a second convolution layer, a third convolution layer, a fourth convolution layer, a first CBAM module, a second CBAM module, a third CBAM module and a fifth convolution layer.

9. The metal surface defect detection method based on improved YOLOv8 as claimed in claim 8, characterized in that: The convolution kernels of the first convolution layer and the fifth convolution layer are 1×1 convolution kernels; the convolution kernels of the second convolution layer, the third convolution layer and the fourth convolution layer are 3×3 convolution kernels.

10. The metal surface defect detection method based on improved YOLOv8 according to claim 9, characterized in that: The improved SPPF process of input data includes: The input data is input into the first convolution layer, the processing result of the first convolution layer is input into the first maximum pooling layer, the processing result of the first maximum pooling layer is input into the second maximum pooling layer and the second convolution layer, the processing result of the second maximum pooling layer is input into the third maximum pooling layer and the third convolution layer, and the processing result of the third maximum pooling layer is input into the fourth convolution layer; the processing result of the second convolution layer is input into the first CBAM module, the processing result of the third convolution layer is input into the second CBAM module, and the processing result of the fourth convolution layer is input into the third CBAM module; after the processing results of the first CBAM module, the second CBAM module and the third CBAM module are fused, they are input into the fifth convolution layer after the Concat operation on the channel dimension with the processing result of the first convolution layer.

Citation Information

Cited By

  • Road crack lightweight segmentation and quantification method based on enhanced feature fusion

    CN122510579A

  • A method, system, device and storage medium for detecting surface defects in aluminum sheets.

    CN122675861A