A Thin Film Defect Detection Method and Device Based on the MSSA-YOLOv8 Lightweight Model

By constructing a lightweight YOLOv8 network of multi-scale striped attention module, the detection accuracy and speed problems in film defect detection are solved, and efficient film defect detection is achieved.

CN119850537BActive Publication Date: 2025-07-29HANGZHOU SHENGHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411897945.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-07-29
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing film defect detection methods have limited detection accuracy when they are characterized by low contrast, weak semantics, pinstriations, etc., and deep learning models such as YOLOv8 are difficult to deploy in real-time in industrial production with limited resources.

Method used

Using a lightweight YOLOv8 network structure based on multi-scale striped attention (MSSA), the network depth and parameter amount are reduced by building the MSSA-YOLOv8 lightweight model, combined with the characteristics of film defects, the detection accuracy and speed are improved.

Benefits of technology

High resolution and fast defect detection in film detection are achieved, with the model parameter volume of only 0.4M, which significantly improves the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850537B_ABST
    Figure CN119850537B_ABST
Patent Text Reader

Abstract

A film defect detection method and device based on the MSSA-YOLOv8 lightweight model according to the present invention specifically includes the following steps: S1. Obtain film image data containing film surface defect information, preprocess it to obtain an annotated data set, and divide the annotated data set into a training set and a validation set; S2. Construct an initial MSSA-YOLOv8 lightweight model through the YOLOv8 network architecture and a multi-scale stripe attention module; S3. Input the training set in step S1 into the initial MSSA-YOLOv8 lightweight model for model training, and monitor the accuracy through the validation set until the model converges, stop training, and obtain the final MSSA-YOLOv8 network model; S4. Input the film image data to be detected into the final MSSA-YOLOv8 network model to complete the detection of film defects. The film defect detection method based on the lightweight MSSA-YOLOv8 proposed by the present invention is aimed at the scenario of fast detection speed with high resolution for film detection defects, and according to the characteristics of the weak semantic defects of the single-channel image of the film, the backbone network is cleverly designed to reduce the depth and the number of parameters of the network model, and the number of model parameters is only 0.4M.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and image technology, and particularly relates to a film defect detection method and device based on a lightweight MSSA-YOLOv8 model. Background Art

[0002] Traditional methods for film defect detection mostly rely on manual visual inspection or methods based on traditional machine vision. However, these methods have the problem that it is difficult to stably extract defect features. Especially when film defects usually exhibit characteristics such as low contrast, weak semantics, and fine stripes, the detection accuracy of traditional methods is limited. Existing deep learning methods such as the YOLOv8 model have made significant progress in the field of object detection. However, due to the complexity of the model and the large number of parameters, it is difficult to be deployed in real time in industrial production with limited resources. To solve this problem, the present invention combines the characteristics of film defect detection and proposes a lightweight YOLOv8 network structure based on multi-scale stripe attention (MSSA) to improve the detection accuracy and speed. Summary of the Invention

[0003] The purpose of the present invention is to provide a film defect detection method and device based on a lightweight MSSA-YOLOv8 model to overcome the deficiencies in the prior art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] The present application discloses a film defect detection method based on a lightweight MSSA-YOLOv8 model, which specifically includes the following steps:

[0006] S1. Obtain film image data containing film surface defect information, preprocess it to obtain an annotation dataset, and divide the annotation dataset into a training set and a validation set;

[0007] S2. Construct an initial MSSA-YOLOv8 lightweight model through the YOLOv8 network architecture and a multi-scale stripe attention module;

[0008] S3. Input the training set in step S1 into the initial MSSA-YOLOv8 lightweight model for model training, and monitor the accuracy through the validation set until the model converges, stop training, and obtain the final MSSA-YOLOv8 network model;

[0009] S4. Input the film image data to be detected into the final MSSA-YOLOv8 network model to complete the detection of film defects.

[0010] Preferably, the preprocessing in step S1 includes:

[0011] S11. Crop and resize the thin-film image data to obtain image patches of 1024*1024; the image patches contain thin-film surface defect information.

[0012] S12. Mark the regions where the thin-film surface defect information is located in the image patches obtained in step S11, and assign category labels according to the types of the thin-film surface defect information to obtain an annotated data set; the types of the thin-film surface defect information include dirt, creases, and bright spots.

[0013] Preferably, the ratio of the training set to the validation set is 8:2.

[0014] Preferably, step S2 specifically includes the following sub-steps:

[0015] S21. Set the number of input channels of the first convolutional layer of the YOLOv8 network architecture to 1; the backbone network includes 5 stages: stage1, stage2, stage3, stage4, and stage5. The number of channels of stage1 is 8 and it consists of 1 CONV_BN_SILU module; the numbers of channels of stage2, stage3, stage4, and stage5 are 8, 16, 32, and 64 respectively, and they are composed of CONV_BN_SILU modules connected to C2F modules.

[0016] S22. Insert multi-scale stripe attention modules after stage2, stage3, stage4, and stage5.

[0017] S23. Input the output feature maps of stage2, stage3, stage4, and stage5 into the FAPAN module; complete the construction of the initial MSSA-YOLOv8 lightweight model.

[0018] Preferably, the operation of the multi-scale stripe attention module is as follows: for a two-dimensional tensor , perform convolutional operations with stripe widths of 1, 3, and 5. The calculation in the horizontal direction is as follows:

[0019] ;

[0020] The calculation in the horizontal direction is as follows:

[0021] ;

[0022] The output of MSSA is as follows:

[0023] ;

[0024] where is the convolutional kernel size of , and the stride is The two-dimensional convolution operation.

[0025] Preferably, the specific operations in step S3 include the following:

[0026] S31. Set the training parameters;

[0027] S32. Input the training set into the initial MSSA-YOLOv8 lightweight model for model training;

[0028] S33. During the model training process, monitor the accuracy of the model through the validation set to determine whether the training stop criterion is met; if it is met, go to step S34, otherwise, loop step S33;

[0029] S34. Stop training to obtain the final MSSA-YOLOv8 network model.

[0030] Preferably, the training parameters in step S31 include the input image size of 1024*1024, the batch size of 8, and the upper limit of the total training epoch of 500.

[0031] Preferably, the training stop criterion in step S33 is as follows: the accuracy of the model does not increase for 100 consecutive epochs or reaches the upper limit of the training epoch.

[0032] The present invention also discloses a film defect detection device based on the MSSA-YOLOv8 lightweight model, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned film defect detection method based on the MSSA-YOLOv8 lightweight model.

[0033] The present invention also discloses a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned film defect detection method based on the MSSA-YOLOv8 lightweight model.

[0034] Advantages of the present invention:

[0035] 1. The film defect detection method based on the lightweight MSSA-YOLOv8 proposed by the present invention is aimed at the scenario of film detection defects with high-resolution and fast detection speed. And according to the characteristics of the weak semantic defects of the single-channel image of the film, the backbone network is cleverly designed to reduce the depth and the number of parameters of the network model. The number of model parameters is only 0.4M.

[0036] 2. The present invention proposes a Multi Scale Strip Attention (MSSA) module. In the scenario of thin film defect detection, there are many thin strip-shaped defects, such as dirt, creases, bright spots, etc. Inspired by the StringPooling algorithm, the present invention designs a multi-scale strip attention module, which can enable the model to capture richer and key feature representations, and improve the detection effect of the model on thin film defects with only a small increase in the number of parameters.

[0037] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a schematic flowchart of a thin film defect detection method cabinet based on the MSSA-YOLOv8 lightweight model of the present invention;

[0039] Figure 2 is the architecture diagram of the MSSA-YOLOv8 model of the present invention;

[0040] Figure 3 is a comparison table of the number of parameters and mAP@0.5 indicators of different models;

[0041] Figure 4 is a schematic structural diagram of a thin film defect detection device based on the MSSA-YOLOv8 lightweight model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through the drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0043] Refer to Figure 1 , an embodiment of the present invention provides a thin film defect detection method based on the MSSA-YOLOv8 lightweight model, which specifically includes the following steps:

[0044] S1. Obtain thin film image data containing thin film surface defect information, preprocess it to obtain an annotated data set, and divide the annotated data set into a training set and a validation set;

[0045] Specifically, the preprocessing in step S1 includes:

[0046] S11. Crop and resize the thin film image data to obtain an image block of 1024*1024; the image block contains thin film surface defect information;

[0047] S12. Mark the area where the film surface defect information is located in the image block of step S11, and assign a category label according to the type of the film surface defect information to obtain an annotation dataset; the types of the film surface defect information include dirt, crease, and bright spot.

[0048] S2. Construct an initial MSSA-YOLOv8 lightweight model through the YOLOv8 network architecture and the multi-scale stripe attention module;

[0049] Specifically, step S2 includes:

[0050] S21. Set the number of input channels of the first convolutional layer of the YOLOv8 network architecture to 1; the backbone network includes 5 stages: stage1, stage2, stage3, stage4, and stage5. The number of channels of stage1 is 8 and it consists of 1 CONV_BN_SILU module; the numbers of channels of stage2, stage3, stage4, and stage5 are 8, 16, 32, and 64 respectively, and they are composed of CONV_BN_SILU modules connected to C2F modules;

[0051] S22. Insert a multi-scale stripe attention module after stage2, stage3, stage4, and stage5,

[0052] S23. Input the output feature maps of stage2, stage3, stage4, and stage5 into the FAPAN module; complete the construction of the initial MSSA-YOLOv8 lightweight model, as Figure 2 shown;

[0053] Among them, the operation of the multi-scale stripe attention module is as follows: for the two-dimensional tensor , perform convolutional operations with stripe widths of 1, 3, and 5. The calculation in the horizontal direction is as follows:

[0054] ;

[0055] The calculation in the horizontal direction is as follows:

[0056] ;

[0057] The output of MSSA is as follows:

[0058] ;

[0059] Among them is a two-dimensional convolutional operation with a convolutional kernel size of and a stride of .

[0060] S3. Input the training set in step S1 into the initial MSSA-YOLOv8 lightweight model for model training, and monitor the accuracy through the validation set until the model converges, then stop training to obtain the final MSSA-YOLOv8 network model;

[0061] Specifically, the following operations are included in step S3:

[0062] S31. Set the training parameters;

[0063] S32. Input the training set into the initial MSSA-YOLOv8 lightweight model for model training;

[0064] S33. During the model training process, monitor the accuracy of the model through the validation set to determine whether the stop training standard is reached; if so, enter step S34, otherwise, loop step S33;

[0065] S34. Stop training to obtain the final MSSA-YOLOv8 network model.

[0066] S4. Input the thin film image data to be detected into the final MSSA-YOLOv8 network model to complete the detection of thin film defects.

[0067] In a feasible embodiment, the ratio of the training set to the validation set is 8:2.

[0068] In a feasible embodiment, the training parameters in step S31 include the input image size of 1024*1024, the batchsize of 8, and the upper limit of the total training epochs of 500.

[0069] In a feasible embodiment, the stop training standard in step S33 is as follows: the accuracy of the model does not increase for 100 consecutive epochs or reaches the upper limit of the training epochs.

[0070] Embodiment:

[0071] 1. Thin film image data acquisition. Collect thin film image data on the industrial production line to ensure that it includes common defects such as dirt, creases, and bright spots. Ensure the image clarity to ensure that the details of small defects can be distinguished;

[0072] 2. Image cropping and size adjustment. Crop the original collected image into blocks of 1024x1024 to meet the input requirements of YOLOv8. Ensure that each image block contains the thin film surface defect information during cropping.

[0073] 3. Use the labeling tool (LabelMe) to label the defective areas and assign category labels to each type of defect (such as dirt, creases, bright spots, etc.). Convert the annotation format to a YOLO annotation file, and divide the annotated dataset into a training set (80%) and a validation set (20%) to ensure good performance of the model in both the training and validation phases.

[0074] 4. Build a lightweight backbone network and construct a lightweight model based on the YOLOv8 architecture. The backbone network is divided into five stages (stage1 to stage5), and the number of input channels of the first convolutional layer is set to 1 to adapt to single-channel image input. The characteristics of each stage are as follows:

[0075] Stage1: Consists of a CONV_BN_SILU module for initial feature extraction.

[0076] Stage2 to Stage5: Each stage includes a CONV_BN_SILU module followed by a C2F module, and the number of channels is set to 8, 8, 16, 32, and 64.

[0077] 5. Add a multi-scale stripe attention module (MSSA) and insert the MSSA module after stage2 to stage5. The MSSA module uses convolutional operations with stripe widths of 1, 3, and 5, calculates the feature information in the horizontal and vertical directions, combines the multi-scale stripe features, and enhances the feature capture ability of the model.

[0078] The multi-scale stripe attention module (Multi Scale Strip Attention, MSSA), the mathematical calculation of this module is defined as follows:

[0079] Given a two-dimensional tensor , where is the stripe width (the stripe widths adopted in the present invention are 1, 3, and 5), and where is a two-dimensional convolution operation with a convolution kernel size of , a stride of .

[0080] The calculation in the horizontal direction is as follows:

[0081]

[0082] The calculation in the vertical direction is as follows:

[0083]

[0084] For MSSA, for multiple sets of stripe widths 1, 3, 5, the output of MSSA is as follows:

[0085] ;

[0086] 6. Feature fusion: Input the feature map of stage2 into the FAPAN module, and combine feature maps of different scales to further improve the detection accuracy, especially in capturing the features of fine stripe defects.

[0087] 7. Model training settings: Use the processed training dataset, set the input image size to 1024x1024, batch size to 8, and the upper limit of the total training epochs to 500 to ensure sufficient training iterations to improve the model's stability.

[0088] 8. Adaptive learning rate dynamic adjustment: Adopt an adaptive learning rate strategy to automatically adjust the learning rate according to the validation set accuracy, enabling the model to reach the optimal performance state more quickly.

[0089] 9. Early Stopping strategy: Monitor the validation set accuracy during training. When the validation set accuracy has not improved for 100 consecutive rounds, stop training to prevent overfitting.

[0090] 10. Training process monitoring and evaluation: Monitor the changes in training loss and validation set accuracy, and save the model with the best performance on the validation set for use in actual detection.

[0091] 11. Real-time detection system deployment: Deploy the trained MSSA-YOLOv8 model to the film detection system. The system divides the film images on the production line into blocks of 1024x1024 size and inputs them into the model block by block for defect detection.

[0092] 12. Defect identification and classification: The model performs defect detection on each block of image and outputs information such as defect location, type, and confidence.

[0093] 13. Visualization of detection results: Display the detection results on the monitoring interface in real time, marking the defect location, category, and confidence to facilitate the operators to monitor the film quality.

[0094] 14. Defect handling and alarm mechanism: The system can trigger an alarm for serious defects, remind the operators to intervene, and save the detection records for subsequent quality analysis and tracking.

[0095] As Figure 3 shown is the comparison chart of the number of parameters and mAP@0.5 index for different models; it can be seen from the figure that compared with YOLOv8n, the number of parameters of the MSSA-YOLOv8 proposed by the present invention is reduced by 87.7%, and under the condition of using mAP@0.5 as the accuracy evaluation index, the accuracy index is increased by 0.219. That is, the lightweight MSSA-YOLOv8 model proposed by the present invention is both fast and accurate in terms of speed and accuracy.

[0096] An embodiment of a thin film defect detection device based on the MSSA-YOLOv8 lightweight model of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the thin film defect detection device based on the MSSA-YOLOv8 lightweight model of the present invention is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, the device where the embodiment is located in any device with data processing capabilities usually also includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here. The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0097] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0098] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a thin film defect detection device based on the MSSA-YOLOv8 lightweight model in the above embodiment.

[0099] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0100] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, or improvement made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A thin film defect detection method based on the MSSA-YOLOv8 lightweight model, characterized in that, Specifically, it includes the following steps: S1. Obtain thin-film image data containing thin-film surface defect information, preprocess it to obtain an annotated dataset, and divide the annotated dataset into a training set and a validation set; S2. Construct an initial MSSA-YOLOv8 lightweight model through the YOLOv8 network architecture and a multi-scale stripe attention module; The operation of the multi-scale stripe attention module is as follows: For a two-dimensional tensor , convolution operations with stripe widths of 1, 3, and 5 are used, and the calculation in the horizontal direction is as follows: ; The calculation in the horizontal direction is as follows: ; The output of MSSA is as follows: ; wherein is a two-dimensional convolution operation with a convolution kernel size of and a stride of ; S3. Input the training set in step S1 into the initial MSSA-YOLOv8 lightweight model for model training, and monitor the accuracy through the validation set until the model converges, then stop training to obtain the final MSSA-YOLOv8 network model; S4. Input the thin-film image data to be detected into the final MSSA-YOLOv8 network model to complete the detection of thin-film defects.

2. The thin-film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 1, characterized in that: The preprocessing in step S1 includes: S11. Crop and resize the thin-film image data to obtain image patches of 1024*1024; the image patches contain thin-film surface defect information; S12. Mark the area where the thin-film surface defect information is located in the image patches in step S11, and assign class labels according to the type of thin-film surface defect information to obtain an annotated dataset; the types of thin-film surface defect information include dirt, creases, and bright spots.

3. A thin film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 1, characterized in that: The ratio of the training set to the validation set is 8:

2.

4. A thin film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 1, characterized in that: Step S2 specifically includes the following sub-steps: S21. Set the number of input channels of the first convolutional layer of the YOLOv8 network architecture to 1; the backbone network includes 5 stages: stage1, stage2, stage3, stage4, and stage5. The number of channels of stage1 is 8 and it consists of 1 CONV_BN_SILU module; the numbers of channels of stage2, stage3, stage4, and stage5 are 8, 16, 32, and 64 respectively, and they are composed of CONV_BN_SILU modules connected to C2F modules; S22. Insert a multi-scale stripe attention module after stage2, stage3, stage4, and stage5, S23. Input the output feature maps of stage2, stage3, stage4, and stage5 into the FAPAN module; complete the construction of the initial MSSA-YOLOv8 lightweight model.

5. A thin film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 1, characterized in that: The specific operations in step S3 include: S31. Set training parameters; S32. Input the training set into the initial MSSA-YOLOv8 lightweight model for model training; S33. During the model training process, monitor the accuracy of the model through the validation set to determine whether the stop training standard is reached; if so, enter step S34, otherwise, loop step S33; S34. Stop training to obtain the final MSSA-YOLOv8 network model.

6. The thin film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 5, characterized in that: The training parameters in step S31 include an input image size of 1024*1024, a batch size of 8, and a total training epoch upper limit of 500.

7. A film defect detection method based on the MSSA-YOLOv8 lightweight model according to claim 5, characterized in that: The stopping training criteria in step S33 are as follows: the accuracy of the model does not increase for 100 consecutive epochs or reaches the upper limit of the training epochs.

8. A thin film defect detection device based on the MSSA-YOLOv8 lightweight model, characterized in that: It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement a film defect detection method based on the MSSA-YOLOv8 lightweight model according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A program is stored thereon. When the program is executed by a processor, it implements a film defect detection method based on the MSSA-YOLOv8 lightweight model according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Lightweight detection method, system and equipment for casting surface defects

    CN118196059A

  • Key video data extraction method based on multi-dimensional semantic information

    WO2024109308A1