Screw and screw hole detection method based on improved YOLOv8-Seg

By improving the YOLOv8-Seg model, combining data enhancement and feature fusion technology, the problem of inaccurate positioning of the YOLO algorithm in screw and hole detection is solved, achieving higher detection accuracy and real-time performance.

CN120071081APending Publication Date: 2025-05-30JIANGSU UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510140046.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The YOLO algorithm has problems of inaccurate detection and inaccurate positioning of screws and holes. Especially in the detection of small screws and holes on the circuit board, it is difficult to accurately separate multiple similar targets.

Method used

Based on the improved YOLOv8-Seg model, by constructing an image data set of multiple circuit board screws and holes, using mosaic data augmentation technology and a unified input image size standard, the SDI module and DBB module are introduced, and the CIoU loss function is replaced as the DIoU loss function to improve detection accuracy and positioning accuracy.

Benefits of technology

It significantly improves the positioning and segmentation capabilities of screws and holes, improves detection accuracy and real-time performance, and ensures real-time inspection requirements during circuit board assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071081A_ABST
    Figure CN120071081A_ABST
Patent Text Reader

Abstract

The invention discloses a circuit board screw and hole instance segmentation method based on improved YOLOv8-Seg, and relates to screw and hole assembly, and the method comprises the following steps: firstly introducing a multi-level feature fusion module (SDI), fusing features from different levels, and enhancing semantic information and detail information of screws and holes in an image; secondly, adding a diversity branch module (DBB) to replace common convolution operation, and enhancing multi-scale feature information by adding a plurality of parallel branches and reconstructing parameters; finally, a DIOU loss function is used for replacing a CIoU loss function, the distance between a prediction frame and the center point of a real frame is reduced, the position and edge features of a target are better positioned, the accuracy of screw and hole segmentation positioning is improved, and experimental results show that compared with a traditional YOLOv8-Seg method, the YOLOv8-SBD-Seg algorithm has the advantages that the segmentation precision is improved by 4.59%, and the accuracy of screw and hole segmentation positioning is improved. The screw and hole positioning precision and the assembling efficiency are remarkably improved in the screw and hole assembling occasion, and wide application prospects are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of screw and hole assembly, and particularly to a method for detecting screws and screw holes based on improved YOLOv8-Seg. Background Art

[0002] In industrial production, the detection of screws and screw holes are both basic tasks in the field of object detection. In 3C assembly, a large number of small targets need to be accurately positioned and aligned, and tiny deviations may damage the equipment. Traditional machine vision relies on industrial cameras to collect images and combines traditional image processing and machine learning to obtain results. With the development of deep learning, enterprises apply it to assembly vision detection to improve production efficiency.

[0003] Deep learning vision detection realizes the automatic extraction and classification of image features through a deep neural network model, thereby realizing the detection and recognition of target objects. Compared with traditional machine vision detection, deep learning vision detection does not require manual feature design and can automatically learn and extract image features. Therefore, it has higher detection accuracy and stronger generalization ability. The detection methods based on deep learning mainly include YOLO, Mask R-CNN, and SSD. For example, SSD and YOLO series directly classify and predict targets at each position on the feature map, with faster detection speed and stronger practicability.

[0004] In 2023, YOLOv8 released by Ultralytics is more excellent than previous YOLOs in terms of speed and accuracy, attracting wide attention from researchers in the field of object detection. YOLOv8 provides a model training framework and can perform basic tasks such as object detection, instance segmentation, image classification, and pose estimation;

[0005] The patent application number 202411084821.8 is based on improved YOLOv8, adding a P2 feature layer to improve the detection accuracy of small targets, introducing a TA layer to improve the recognition accuracy of the model, and introducing a CSC module to reduce model parameter redundancy;

[0006] The patent application number 202411106535.7 is based on improved YOLOv8, adding an attention mechanism and a lightweight module to improve the detection ability of small targets in drone images;

[0007] The patent application number 202411579522.1 is based on YOLOv8 for lightweight improvement to improve the inference speed and embed it into the computing terminal;

[0008] The patent application number 202410301072.3 is based on YOLOv8 to improve the feature pyramid and more accurately obtain image pixel information. The above methods all have a certain improvement in the detection accuracy of small targets, but do not emphasize the real-time performance of the algorithm.

[0009] In summary, the YOLOv8 algorithm has made great progress in the fields of object detection and instance segmentation. However, when it comes to detecting and positioning small screws and holes on circuit boards, when multiple similar objects appear in the image, the YOLO algorithm may consider them as one object or fail to accurately separate them. Using YOLOv8 for screw and hole detection and positioning is relatively fast, but there are problems of inaccurate detection and positioning for small objects. And there are also certain requirements for the real-time performance of instance segmentation in the production process. Based on this, the present invention designs a method for detecting screws and screw holes based on improved YOLOv8-Seg. Summary of the Invention

[0010] The purpose of the present invention is to solve the problems in the prior art, and a method for detecting screws and screw holes based on improved YOLOv8-Seg is proposed.

[0011] A method for detecting screws and screw holes based on improved YOLOv8-Seg includes the following steps:

[0012] S1: Construct a dataset: In order to accurately detect small screws and holes on circuit boards, the invention proposes a method for instance segmentation of screws and holes based on improved YOLOv8-Seg. First, an image dataset of various circuit board screws and holes is constructed. In the process of constructing the dataset, the existing public dataset WEEE screw disassembly dataset is utilized, and for 3C devices such as PCB boards, laptops, and single-chip microcomputers, a self-made dataset of circuit board screw holes is made, including the fine collection of device screw and hole images, and the manual annotation work through the labelme tool. Finally, a dataset containing multiple pictures is obtained, which includes small screws and screw holes on various circuit boards. In order to ensure the scientificity and practicality of the dataset, we use the Python module for random division, and divide the dataset into a training set, a validation set, and a test set according to 8:1:1.

[0013] S2: Image preprocessing: When performing the instance segmentation task of circuit board screw and screw hole images, first load the relevant image data into the memory. Subsequently, the mosaic data augmentation technology is adopted to randomly crop, scale, and re-stitch four images to generate a new image containing various types of screw and hole detections. This process aims to enhance the network's robustness to different screw and hole detections. At the same time, a unified input image size standard is set, and the initial anchor box size parameters are automatically verified and adjusted to ensure that the best recall rate of the anchor box on the screw and hole dataset is not less than 0.98.

[0014] S3: Improved YOLOv8-Seg Model: The images of screws and screw holes are input into the backbone network for feature extraction, feature fusion is achieved through the neck network, and the SDI module proposed in U-NET V2 is introduced to replace the original Concat operation. The DBB module is introduced into the head. Through the multi-branch structure and dilated convolution, the fine textures of screws and their holes and complex background details are extracted, and the key features are highlighted through weighted fusion. The original CIoU loss function is replaced with DIoU.

[0015] S4: Model Training and Testing: The improved YOLOv8-Seg model is used to train and evaluate the screw and screw hole dataset. Appropriate hyperparameters are selected during the training process, the trained model data is analyzed and judged, and then the model is adjusted and optimized. The weights of the optimal model are retained, and experimental tests are performed on the trained model.

[0016] Further, in step S3, the YOLO network uses the intersection over union (IOU) as the similarity metric for rectangular boxes. Its calculation method is to divide the intersection of the predicted box and the ground truth box by the union of the two, that is:

[0017]

[0018] where A is the area covered by the predicted box and B is the area covered by the ground truth box. IOU has the advantages of non-negativity and symmetry. However, when there is no intersection between the two rectangular boxes, the similarity cannot be calculated. In addition, IOU cannot give the details of the intersection between rectangular boxes and cannot measure the similarity of the center point distance, length, and width between rectangular boxes, resulting in different sensitivities of the IOU metric to targets of different scales and low detection performance for small-size targets. The present invention replaces the DIoU loss function, which can effectively solve such problems. Applying this loss function in the screw and hole segmentation task can greatly improve the detection accuracy and efficiency. The calculation formula of the DIoU loss function is:

[0019]

[0020] In the formula, b and b gt are the center points of the anchor box and the target box, and c is the diagonal distance of the smallest rectangle that can cover both the anchor and the target box. ρ is the Euclidean distance between these two center points, and its calculation formula is:

[0021]

[0022] Further, in step S3, the SDI module proposed in U-NET V2 is introduced into the neck of YOLOv8 to replace the original Concat operation, which specifically includes the following steps:

[0023] First, the hierarchical feature maps generated by the encoder are first applied with spatial and channel attention mechanisms to the features at each level i to enable the features to integrate local spatial information and global channel information:

[0024]

[0025] where is the feature map processed by the i-th layer, and are the i-th spatial and channel attention parameters. Using 1×1 convolution to reduce the channels to c, we get The processed feature map is sent to the decoder to as a reference, resize the feature map at each j-th layer to match the same size as :

[0026]

[0027] and a 3×3 convolution is used to smooth the feature map

[0028]

[0029] where θ ij is the parameter of the smoothing convolution, is the j-th smoothed feature map of the i-th layer. After adjusting all the feature maps to the same size resolution, the Hadamard product is applied to all the resized feature maps, thereby enhancing the features at the i-th level with more semantic information:

[0030]

[0031] Furthermore, in step S3, the DBB module is used to improve the original YOLOv8 detection head. During training, the DBB module extracts the fine textures of screws and holes and the details of complex backgrounds through a multi-branch structure and dilated convolutions, and highlights the key features through weighted fusion, enhancing the sensitivity of the model to boundaries and details. During the training process, it includes four branches: 1×1 convolution + BN, 1×1 convolution + K×K convolution + BN, 1×1 convolution + average pooling, K×K convolution + BN. Compared with a single K×K convolution, the initial multi-branch structure has stronger feature extraction and scale information fusion capabilities. During prediction, through reparameterization transformations such as BN fusion and branch summation, the four branches are merged into a single K×K convolution to ensure that the inference speed is not affected by the multi-branch structure.

[0032] Compared with the existing technology, the advantages of the present invention are as follows: an SDI module is introduced to replace the original Concat operation. The SDI module captures multi-scale features by using convolution kernels with different dilation rates, enabling the model to more comprehensively extract the features of targets with different sizes and shapes. The DBB module is introduced at the head, and the fine textures of screws and their holes and complex background details are extracted through a multi-branch structure and dilated convolution, and the key features are highlighted through weighted fusion. The original CIoU loss function is replaced with DIoU. Given that the position difference between the predicted box and the ground truth box of small targets such as screws and their holes is tiny, the consideration of the position relationship is increased to improve the positioning accuracy. Description of the Drawings

[0033] Figure 1 It is a structural diagram of a screw and screw hole detection algorithm based on YOLOv8-Seg.

[0034] Figure 2 It is a network model diagram of YOLOv8-Seg.

[0035] Figure 3 It is a network model diagram of YOLOv8-SBD-Seg.

[0036] Figure 4 It is a structural diagram of the SDI module.

[0037] Figure 5 It is a structural diagram of the DBB module.

[0038] Figure 6 It is a diagram of the dataset for collecting circuit board screws and holes.

[0039] Figure 7 It is a comparison diagram of the segmentation effects of screws and holes by different algorithms. Detailed Implementation Manner

[0040] Referring to Figures 1 to 7 , a screw and hole segmentation method based on improved YOLOv8-Seg includes the following steps:

[0041] S1: Construct a dataset: During the construction of the dataset, the existing public dataset WEEE screw disassembly dataset is utilized, and a self-made dataset of circuit board screw holes is made for 3C devices such as PCB boards, laptops, and single-chip microcomputers. The production process of this self-made dataset involves the fine collection of screw and hole images of these devices and the manual annotation work through the labelme tool. Finally, a dataset containing 1290 pictures is obtained, which includes small screws and screw holes on various circuit boards. To ensure the scientificity and practicability of the dataset, we use the Python module for random division, dividing the dataset into 1032 training sets, 129 validation sets, and 129 test sets.

[0042] S2: Image preprocessing: When performing the instance segmentation task of circuit board screws and screw holes, first load the relevant image data into memory. Subsequently, use the mosaic data augmentation technique to randomly crop, scale, and re - splice four images to generate a new image containing various screw and screw hole detection types. This process aims to enhance the network's robustness to different screw and screw hole detections. At the same time, set a unified input image size standard and automatically verify and adjust the initial anchor box size parameters to ensure that the best recall rate of the anchor box on the screw and screw hole dataset is not less than 0.98.

[0043] S3: Establish an improved YOLOv8 - Seg model: Refer to Figure 1 , the YOLOv8 - Seg model mainly consists of a backbone network, a neck network, and a head network, etc. Among them, the backbone network is responsible for feature extraction. The backbone network of YOLOv8 - Seg uses CSPDarknet53 as its basic architecture, and enhances the information flow between different stages of the network through Cross - Stage Partial (CSP) connections, while improving the gradient flow during training. Further upsampling operations to increase the spatial resolution of the feature map, and Concat operations to merge the features of different layers. The head network, based on the obtained feature maps of three sizes (large, medium, and small), adjusts the number of channels and predictions through a re - parameterized structure, and then through the processing of the CIOU (Complete intersection over union) loss function and Non - Maximum Suppression (NMS), obtains the final prediction result.

[0044] Refer to Figure 2 the YOLOv8 - Seg network structure diagram. When in use, the image to be detected first undergoes unified scaling and becomes 640×640×3 in size according to the set hyperparameters. The backbone network consists of Conv, C2f, and SPPF. The input image first passes through a Conv module, then four times through a Conv module and a C2f module, and finally through the SPPF module. Multiple C2f modules contain multiple residual connections, which can enrich the gradient flow and effectively extract features, and transfer the features of different levels to the Neck module. The Neck part consists of Concat, Upsample, C2f, and Conv. The features are fused through two processes: bottom - up and top - down. The bottom - up process uses the backbone network to generate a series of feature maps; the top - down process gradually upsamples the high - level feature maps. The head network consists of a re - parameterized convolution (Re - parameterized Convolution, RepConv) and a segmentation head Segment.

[0045] Replace the original CIoU loss function with DIoU. Given that the position difference between the predicted box and the ground truth box of small targets such as screws and their holes is tiny, the consideration of the position relationship is increased to improve the positioning accuracy. The YOLO network uses the intersection over union (IOU) as the similarity metric for rectangular boxes. Its calculation method is to divide the intersection of the predicted box and the ground truth box by the union of the two, that is:

[0046]

[0047] where A is the area covered by the predicted box and B is the area covered by the ground truth box. IOU has the advantages of non-negativity and symmetry. However, when there is no intersection between the two rectangular boxes, the similarity cannot be calculated. In addition, IOU cannot give the details of the intersection between rectangular boxes and cannot measure the similarity of the center point distance, length, and width between rectangular boxes, resulting in different sensitivities of the IOU metric to targets of different scales and relatively low detection performance for small-size targets. The present invention replaces the DIoU loss function, which can effectively solve such problems. Applying this loss function to the screw and hole segmentation task can greatly improve the detection accuracy and efficiency. The calculation formula of the DIoU loss function is:

[0048]

[0049] In the formula, b and b gt are the center points of the anchor box and the target box, and c is the diagonal distance of the smallest rectangle that can cover both the anchor and the target box. ρ is the Euclidean distance between these two center points, and its calculation formula is:

[0050]

[0051] Input the images of screws and screw holes into the backbone network for feature extraction, implement feature fusion through the neck network, and introduce the SDI module proposed in U-NET V2 to replace the original Concat operation, which specifically includes the following steps: First, apply the spatial and channel attention mechanisms to the hierarchical feature maps generated by the encoder for the features at each level i to enable the features to integrate local spatial information and global channel information:

[0052]

[0053] In the formula is the feature map processed at the i-th layer;

[0054] and are the i-th spatial and channel attention parameters;

[0055] Then use a 1×1 convolution to reduce the channels to c, obtaining Send the processed feature map to the decoder to using it as a reference, resize the feature map at each j-th layer to match the same size:

[0056]

[0057] and apply a 3×3 convolution to smooth the feature map

[0058]

[0059] where θ ij is the parameter of the smoothing convolution, is the j-th smoothed feature map of the i-th layer. After adjusting all the feature maps to the same size resolution, apply the Hadamard product to all the resized feature maps to enhance the features at the i-th level with more semantic information:

[0060]

[0061] The SDI module captures multi-scale features by using convolution kernels with different dilation rates, enabling the model to extract features of targets with different sizes and shapes more comprehensively.

[0062] Meanwhile, the DBB module is introduced in the head. Through a multi-branch structure and dilated convolution, it extracts the fine textures of the screws and their holes and complex background details, and highlights the key features through weighted fusion, enhancing the model's sensitivity to boundaries and details to capture a wider range of boundary and background features, thereby improving the localization and segmentation accuracy of the screws and their holes. It includes four branches: 1×1 convolution + BN, 1×1 convolution + K×K convolution + BN, 1×1 convolution + average pooling, K×K convolution + BN. Compared with a single K×K convolution, the initial multi-branch structure has stronger feature extraction and scale information fusion capabilities. During prediction, through reparameterization transformations such as BN fusion and branch summation, the 4 branches are merged into a K×K convolution to ensure that the inference speed is not affected by the multi-branch structure.

[0063] S4: Model Training and Testing: Use the improved YOLOv8-Seg model to train and evaluate the screw and screw hole datasets. All network models are based on the Pytorch 2.0.0 framework, use GPU acceleration on RTX 3070 devices, and adopt Python 3.8.19 as the programming language. Appropriate hyperparameters are selected during training. Analyze and judge the trained model data, and then adjust and optimize the model. Retain the weights of the optimal model and perform experimental tests on the trained model. In the YOLOv8-SBD-Seg method here, the detection scale and feature fusion operations are increased, enhancing the network's perception ability for small targets. Compared with YOLOv8-Seg, the number of model parameters has increased. Increasing the scale also leads to an increase in computational complexity, the neural network becomes deeper, and the time spent on convolution increases, resulting in a decrease in detection speed.

[0064] The software and hardware used in this experimental platform are shown in Table 1.

[0065] Table 1 Experimental Software and Hardware Configuration Table

[0066]

[0067] In this implementation, the average precision (AP), intersection over union (IoU), mAP 50 , mAP 50:95 , precision (P), recall (R), F1 value, number of parameters, computational complexity, and frames per second (FPS) are used. Among them, the average precision AP measures the detection accuracy of the model for a specific class, the intersection over union IoU evaluates the overlap degree between the predicted region and the real region, mAP50 represents the average precision when the IoU threshold is 0.50, and mAP50-95 is the average precision under different IoU thresholds. Precision P represents the proportion of correctly segmented positive samples by the model, recall R measures the proportion of positive samples detected and segmented by the model, and the F1 value combines precision and recall. The number of parameters reflects the complexity of the model, and the fewer the parameters, the simpler the model; the computational complexity measures the training complexity of the model, affecting the training time and resource consumption. FPS reflects the real-time performance of the model, and high FPS is crucial for real-time segmentation tasks. These metrics comprehensively evaluate the detection accuracy, segmentation precision, computational requirements, and real-time performance of the model.

[0068] Perform ablation experiments and comparisons between the method of the present invention and the original YOLOv8-Seg model on the circuit board screw and hole acquisition dataset to verify the gain of each improvement point for the network model. Five groups of ablation experiments were carried out, and the experimental results were statistically analyzed. The ablation experiment results are shown in Tables 2 and 3.

[0069] It can be obtained from Tables 2 and 3 that YOLOv8-Seg1 introduces the SDI module, which enhances the receptive field and preserves feature information, and the mAP 50:95 is significantly improved; YOLOv8-Seg2 introduces the DBB module, and through multi-scale feature fusion and re-parameterization, the mAP 50 is significantly improved; YOLOv8-Seg3 introduces the DIoU loss function to optimize the bounding box localization, and all indicators are generally improved. The improved model 4 is obtained by superimposing the three methods, and the P, R, and mAP values are 90.35%, 79.42%, and 83.00% respectively.

[0070] Through ablation experiments, by introducing the SDI, DBB modules and the DIoU loss function, the segmentation performance of YOLOv8-SBD-Seg on the screw dataset is significantly improved. Compared with YOLOv8-Seg, the P, R, 50 mAP 50:95 are improved by 4.59%, 2.47%, 1.23%, and 0.61% respectively. The improved model of the present invention effectively improves the localization and segmentation capabilities of screws and holes.

[0071] Table 2 Ablation Experiment Results 1

[0072]

[0073] Note: √ indicates the module added to the model.

[0074] Table 3 Ablation Experiment Results 2

[0075]

[0076] Note: √ indicates the module added to the model.

[0077] To verify the advantages of the improved YOLOv8-SBD-Seg network model in the instance segmentation of circuit board screws and holes, using the same dataset and parameter settings, the results of YOLO series instance segmentation methods proposed in recent years are compared as shown in Tables 4 and 5. Experiment 1 is the instance segmentation algorithm of YOLOv9 (Wang C Y, Yeh I H, Mark Liao H Y. Yolov9: Learning what you want to learn using programmable gradient information[C] / / European Conference on Computer Vision. Springer, Cham, 2025: 1-21.), and Experiment 2 is the instance segmentation algorithm of the original YOLOv8-Seg (Varghese R, Sambath M. YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness[C] / / 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS). IEEE, 2024: 1-6.)

[0078] Table 4 Comparison of segmentation performance of different algorithms and YOLOv8-SBD-seg

[0079]

[0080] It can be seen from Table 4 that the YOLOv8-SBD-Seg algorithm proposed in the present invention has significantly improved in terms of indicators such as P, R, mAP 50 and mAP 50:95 compared with Experiment 1 and Experiment 2. Specifically, YOLOv8-SBD-Seg is superior in the instance segmentation task of circuit board screws and holes, with a better segmentation effect than the algorithm in Experiment 1, and overcomes the problems of segmentation omission and error of the algorithm in Experiment 2, demonstrating stronger robustness. As Figure 7 shown, the comparison of the segmentation effects of different algorithms on screws and holes further proves this point.

[0081] Table 5 Comparison of model parameters of different algorithms and YOLOv8-SBD-seg

[0082]

[0083] As can be seen from Table 5: The YOLOv8-SBD-Seg method proposed in the present invention adds detection scales and feature fusion operations, enhancing the network's perception ability for small targets and increasing the number of model parameters compared to the algorithm in Experiment 2. However, at the same time, increasing the scales also leads to an increase in computational complexity, a deeper neural network, and a longer convolution time, resulting in a decrease in detection speed. Compared to the algorithm in Experiment 1, the number of model parameters is also increased, but the computational complexity is reduced and the sensitivity increases to a certain extent. In summary, compared to the algorithm in Experiment 2, the method of the present invention introduces the SDI module and the DBB module, and improves the DIoU loss function. Although there is indeed an increase in the number of parameters and computational complexity, more importantly, it brings a significant improvement in localization and segmentation accuracy, and the FPS data of YOLOv8-SBD-Seg can also ensure the real-time detection during the screw assembly process.

[0084] As is known by common technical knowledge, the present invention can be implemented through other embodiments that do not depart from its spiritual essence or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.

Claims

1. A screw and hole segmentation method based on improved YOLOv8-Seg, characterized by: The following steps are involved: S1: Construct a dataset, wherein the dataset is an image dataset containing a variety of screws and screw holes. The image dataset uses real circuit board screw and screw hole images and combines with public datasets to achieve image acquisition of small screws and screw holes on circuit boards in different environments. S2: Image preprocessing: the detection model loads the input screw and screw hole images, applies data augmentation technology, and generates new image data by rotating, flipping, cropping, adding noise, etc. to the original image in the image classification task; S3: Establish an improved YOLOv8-Seg model, which is mainly composed of a backbone network, a neck network and a head network. The backbone network is responsible for feature extraction and uses CSPDarknet53 as its basic architecture. The Cross-StagePartial connection is used to enhance the information flow between different stages of the network, while improving the gradient flow during training. Further upsampling operations are performed to increase the spatial resolution of the feature map, and Concat operations are performed to merge features of different layers. The head network adjusts the number of channels and predictions through a reparameterized structure based on the obtained feature maps of three sizes: large, medium and small. The final prediction result is obtained by replacing the DIoU loss function of CIoU and performing nonlinear maximum suppression. The images of the screws and screw holes are input into the backbone network for feature extraction. Feature fusion is achieved through the neck network, and the DBB module is introduced to capture a wider range of boundary and background features, thereby improving the positioning and segmentation accuracy of the screws and their holes. The DBB module is used to improve the original YOLOv8 detection head. S4: Model training and testing, use the improved YOLOv8-Seg model to train and evaluate the screw and screw hole dataset, select hyperparameters during the training process, and then analyze and judge the trained model data, then adjust and optimize the model, retain the weight of the optimal model, and perform experimental tests on the trained model.

2. The method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg according to claim 1, characterized in that: The backbone network consists of Conv, C2f, and SPPF. The input image first passes through a Conv module, then passes through a Conv module and a C2f module four times, and finally passes through the SPPF module. Multiple C2f modules contain multiple residual connections, which can enrich the gradient flow, effectively extract features, and transfer features at different levels into the Neck module. The Neck part consists of Concat, Upsample, C2f, and Conv. Features are fused through bottom-up and top-down processes. A series of feature maps are generated from the bottom up using the backbone network; high-level feature maps are gradually upsampled from the top to the bottom. The head network consists of re-parameterized convolution and segmentation head Segment.

3. The method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg according to claim 1, characterized in that: In step S1, the public dataset WEEE screw disassembly dataset and the self-made dataset of screw holes are used comprehensively. The self-made dataset mainly collects images of screws and holes of 3C devices such as PCB boards, laptops, and single-chip microcomputers. The collected images are manually labeled using the labelme tool to mark the location and category information of each screw and screw hole. The dataset is randomly divided into training set, test set, and validation set in a ratio of 8:1:1 through the Python module.

4. The method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg according to claim 1, characterized in that: In step S2, the input image size is set to a uniform size, and the initial anchor frame size parameters are automatically verified and adjusted to ensure that the best recall rate of the anchor frame on the screw and screw hole dataset is not less than 0.

98.

5. The method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg according to claim 1, characterized in that: In step S3, the CIoU loss function of the original YOLOv8 network is replaced by the DIoU loss function. The calculation formula of the DIoU loss function is: Where b and b gt It is the center point of the anchor box and the target box; ρ is the Euclidean distance between the two center points; c is the diagonal distance of the smallest rectangle that can cover both the anchor and the target box.

6. According to claim 1, the present invention relates to a method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg, which is characterized in that: in step S3, when the DBB module is used to improve the original YOLOv8 detection head, the multi-branch structure and dilated convolution are used to extract the fine texture and complex background details of the screws and their holes. In addition, through the weighted fusion technology, key features are highlighted, the model's sensitivity to boundaries and details is enhanced, and the loss of image information is reduced. For the original input image, a feature map X is obtained after feature extraction, wherein the small convolution kernel is responsible for focusing on subtle textures, the large convolution kernel is responsible for extracting the overall shape, and the dilated convolution is used to expand the receptive field, and finally an improved feature map X' is obtained.

7. As described in claim 1, the present invention relates to a method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg, which is characterized in that: in step S3, the SDI module proposed by U-NET V2 is introduced into the neck of YOLOv8, specifically comprising the following steps: First, the spatial and channel attention mechanisms are applied to the features f at each level i through the hierarchical feature maps generated by the encoder i 0 , so that the features can integrate local spatial information and global channel information: Where f i 1 It is the feature map processed by the i-th layer; and Attention parameters for the i-th space and channel; Then use 1×1 convolution to transform f i 1 The channel is reduced to c, and we get The processed feature map is sent to the decoder as f i 2 As a reference, adjust the feature map size at each jth layer to match f i 2 Same size: And use 3×3 convolution to smooth the feature map Where θ ij is the parameter of smoothing convolution, is the jth smoothed feature map of the i-th layer. After resizing all feature maps to the same resolution, the Hadamard product is applied to all resized feature maps to enhance the i-th level features with more semantic information:

8. As described in claim 1, the present invention relates to a method for instance segmentation of screws and screw holes based on improved YOLOv8-Seg, which is characterized in that: in the DBB module, during training, a multi-branch structure and dilated convolution are used to extract subtle textures of the target and complex background details, including four branches: 1×1 convolution + BN, 1×1 convolution + K×K convolution + BN, 1×1 convolution + average pooling, K×K convolution + BN; during prediction, the four branches are merged into one K×K convolution through heavy parameter transformations such as BN fusion and branch summation, ensuring that the reasoning speed is not affected by the multi-branch structure.

Citation Information

Patent Citations

  • Multi-object grabbing method based on improved YOLOv8 network instance segmentation

    CN118262108A

  • Lightweight small target detection method based on improved YOLOv8

    CN119091129A

  • PCB defect identification method based on YOLOv8 model

    CN119107532A

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv8

    CN119107568A