A Keycap Defect Detection Method Based on C-YOLOv3
By adopting a C-YOLOv3-based detection method in keycap defect detection, combining multi-scale attention mechanism and fast pyramid pooling residual blocks, the problems of low detection efficiency and poor accuracy in the prior art are solved, and high-precision and high-speed keycap defect detection are achieved.
Patent Information
- Application Number
- CN202311664562.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-12-06
AI Technical Summary
The prior art has low efficiency and poor accuracy in keycap defect detection, especially when detecting small defects, which is difficult to meet the needs of modern industrial production.
The keycap defect detection method based on C-YOLOv3 is adopted, and the C-YOLOv3 network model is constructed, combining the multi-scale attention mechanism residual block and the fast pyramid pooled residual block, the darknet-53 network is improved, and the model's feature extraction ability for small target defects is enhanced.
The detection accuracy and speed of small target defects on keycaps is significantly improved. Compared with traditional AOI optical detection, the detection accuracy is improved and the detection speed is greatly accelerated, which can more effectively detect tiny defects on keycaps.
Smart Images

Figure CN117808745B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of defect detection, and particularly relates to a keycap defect detection method based on C-YOLOv3, which is applicable to detecting defects on keycaps. Background Art
[0002] The keycap industry plays a crucial role in modern industrial production. There are numerous keycap manufacturers, but the overall automation level of the industry is relatively low. Especially in the defect detection process of finished keycaps, it mainly relies on manual observation currently, and the types of defects need to be manually recorded. This method is inefficient and has poor robustness. At the same time, it takes at least half a year and at most several years to train a skilled keycap detection worker, which greatly increases the production cost and is difficult to meet the requirements of modern industrial production. Therefore, for keycap defect detection, advanced technologies need to be introduced to improve the detection efficiency and accuracy.
[0003] In the field of keycap defect detection, the traditional AOI optical detection technology has high requirements for image quality and clarity. The accuracy of the detection result will decrease due to uneven image illumination. Secondly, AOI optical detection performs poorly in detecting tiny local defects and cannot accurately detect the defects on keycaps. In view of the characteristics of tiny, low-contrast, large-scale variation, and diverse types of defects on keycaps, traditional methods are difficult to achieve good detection results. Currently, the object detection algorithms in the field of artificial intelligence are developing rapidly. By improving the existing object detection algorithm framework for different detection targets, efficient and high-precision defect detection can be achieved.
[0004] Target detection technology is a technology that can identify and locate specific targets in an image. The mainstream target recognition algorithms are mainly divided into two categories according to the network architecture. One is the two-stage target recognition algorithm, such as Mask-RCNN, Faster-RCNN, etc. Since the two-stage network adopts the pre-selection box mechanism, the training result has high accuracy, but the disadvantage is that the detection speed is low. The other is the one-stage target recognition algorithm, such as SSD, YOLO, etc. Among them, the one-stage target detection algorithm of the YOLO series not only has good detection accuracy, but also has a fast detection speed. Since R. Joseph et al. proposed the YOLO (You Only Look Once) algorithm in 2015, target detection technology has developed unprecedentedly. The YOLO algorithm performs very well in position detection and object recognition. At present, the YOLOv3 algorithm has been widely used and has outstanding performance. In response to the similarity problem of some defects and dust in the keyboard production process, Xishi Huang et al. used a multi-layer deep neural network to first group the defects and then distinguish each defect. In order to solve the problem of small targets in remote sensing images being difficult to detect, W. Dong et al. introduced the SE structure into the backbone network, which greatly enhanced the network's ability to detect small targets in remote sensing images. In order to solve the common problems of multi-scale targets and occlusion in pedestrian detection, Y. Yu et al. introduced the CBAM attention mechanism to improve the YOLOv3 network structure and improve the detection accuracy. In order to solve the problems of pedestrian missed detection, false detection, and multi-posture changes of pedestrians, X. Gong et al. introduced the CSPNet structure and ECA attention mechanism into the original YOLOv3 algorithm to improve the accuracy of pedestrian detection. In order to improve the accuracy of foreign object detection on transmission lines in complex environments, P. Liu et al. proposed a new KGM-YOLO algorithm, which can achieve better detection accuracy by replacing the channel attention in the CBAM attention mechanism with the ECA module and adding the neck structure. Summary of the invention
[0005] The purpose of the present invention is to provide a keycap defect detection method based on C-YOLOv3 in view of the above-mentioned problems existing in the prior art.
[0006] The above-mentioned purpose of the present invention is achieved by the following technical means:
[0007] A keycap defect detection method based on C-YOLOv3, comprising the following steps:
[0008] Step 1, obtaining a keycap surface defect image, expanding the keycap surface defect image, and then annotating the defect features of the expanded keycap surface defect image, and finally dividing the annotated keycap surface defect image into a training set, a validation set, and a test set;
[0009] Step 2: Construct the C-YOLOv3 network model, which includes a backbone network and a detection network;
[0010] Step 3: Input the image of the surface defect of the keycap to be detected into the constructed C-YOLOv3 network model to obtain a predicted feature map, and then obtain the category of the defect in the image of the surface defect of the keycap to be detected and the four offset parameters of the corresponding prediction box through the predicted feature map;
[0011] Step 4: Use the training set and minimize the loss function Loss of the C-YOLOv3 network model to train the C-YOLOv3 network model. After the training is completed, save the parameters of the C-YOLOv3 network model to obtain the surface defect detection model of the keycap.
[0012] As described above, the backbone network includes a DBL structure and a residual network module. The residual network module includes 5 residual structures in the Darknet-53 network of the YOLOv3 network model and 3 residual structures embedded in the Darknet-53 network of the YOLOv3 network model;
[0013] The DBL structure includes a convolutional layer, a batch normalization layer, and a SiLU activation function layer arranged in sequence;
[0014] The 3 residual structures embedded in the Darknet-53 network of the YOLOv3 network model respectively include: a residual structure composed of multi-scale attention mechanism residual blocks connected behind the 3rd and 4th residual structures of the residual network module of the Darknet-53 network, and a residual structure composed of a fast pyramid pooling residual block connected behind the 5th residual structure of the residual network module of the Darknet-53 network;
[0015] The number of residual units in the residual blocks corresponding to the 8 residual structures of the residual network module of the backbone network is 1-2-8-1-8-1-4-1 in sequence. Each residual block includes a zero-padding structure, a DBL structure, and multiple residual units.
[0016] As described above, the detection network includes a pyramid feature network and an output layer. The output layer of the detection network includes three multi-channel convolutional modules and three 1×1 convolutional layers, and each multi-channel convolutional module is connected to a 1×1 convolutional layer.
[0017] As described above, the multi-scale attention mechanism residual block is based on the following formula:
[0018] Md = σ(f 1×1 (XAvg Pool(F)))+σ(f 1×1 (YAvg Pool(F)))+F
[0019] Mc(F) = σ(x c (σ(Avg Pool(Md)), f 3×3 (F)) + x c (σ(Avg Pool(f 3×3 (F))), Md)) + F
[0020] Wherein, Md is the attention feature map, F is the input feature map; avg Pool is the global average pooling encoding, XAvgPool(F) is the output encoding vector after the input feature map F undergoes the average pooling encoding in the X direction, and the X direction is the width direction of the feature map; YAvg Pool(F) is the output encoding vector after the input feature map F undergoes the average pooling encoding in the Y direction, and the Y direction is the height direction of the feature map; σ is the Sigmoid function, f 1×1 is a 1×1 convolution operation; Mc(F) is the output feature map, f 3×3 is a 3×3 convolution operation; x c is matrix multiplication.
[0021] As described above, the fast pyramid pooling residual block is to pass the output feature map of the 7th residual structure of the backbone network through a ConvBNMish module, then through 3 MaxPool layers with a size of 5×5, and then connect the outputs of the three MaxPool layers and pass them through a ConvBNMish module for output;
[0022] The ConvBNMish module includes Conv convolution, BN normalization, and Mish activation function.
[0023] As described above, the defect feature annotation is to label the defects in each defective keycap image, and the labels include the category of the defect and the location information of the defect. The categories of the defects are divided into paint spots, cotton fluffs, particles, and indentations.
[0024] As described above, the loss function Loss of the C-YOLOv3 network model includes three parts: classification loss, localization loss, and confidence loss. The loss function Loss of the C-YOLOv3 model is calculated based on the following formula:
[0025] Loss = Lbox + Lobj + Lcls
[0026] Lbox is the localization loss of the defect, Lobj is the confidence loss of the defect, and Lcls is the classification loss of the defect;
[0027] The localization loss Lbox of the defect adopts the CIoU loss function, and the CIoU loss function is based on the following formula:
[0028]
[0029] Where IoU is the intersection over union of the predicted bounding box and the ground truth bounding box, b is the center point of the predicted bounding box, and b gt is the center point of the ground truth bounding box, ρ is the distance between the center points of the predicted bounding box and the ground truth bounding box, c is the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box, α is a tuning factor, v is the aspect ratio penalty term, and α and v are respectively based on the following formulas:
[0030]
[0031]
[0032] Where w1 and h1 are the width and height of the predicted bounding box respectively, are the width and height of the ground truth bounding box respectively.
[0033] As described above, the optimizer of the C-YOLOv3 network model is the SGD model optimizer with momentum; during the training process of the C-YOLOv3 network model, transfer learning is adopted and the training process is divided into two stages. In the first stage, the parameters of the 8 residual structures in the C-YOLOv3 network model remain unchanged, the learning rate is set to 0.001, and the number of training epochs is set to 50 for training; in the second stage, the parameters of the 8 residual structures in the C-YOLOv3 network model are changed, the learning rate is 0.0001, the number of training epochs is set to 250, and the batch size is set to 8 for training.
[0034] The present invention has the following beneficial effects compared with the prior art:
[0035] 1. The C-YOLOv3 network model of the present invention improves the darknet-53 network by adding multi-scale attention mechanism residual blocks and fast pyramid pooling residual blocks, and can more effectively extract the features of small target defects on the keycaps.
[0036] 2. The C-YOLOv3 network model of the present invention improves the detection and localization ability of the model for small target defects by adding a multi-channel convolution module. Compared with traditional AOI optical detection, it greatly speeds up the detection speed while improving the detection accuracy.
[0037] 3. The present invention adopts the SiLU activation function to improve the activation function of the YOLOv3 network, and also adds the intersection over union, which alleviates the problem of gradient disappearance in deep neural networks, reduces the missed detection rate of defects, and improves the detection accuracy at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flowchart of the present invention;
[0039] Figure 2 It is a schematic structural diagram of the C-YOLOv3 network model of the present invention;
[0040] Figure 3 It is a schematic structural diagram of the multi-scale attention mechanism residual block of the present invention (H is the height direction of the feature map, W is the width direction of the feature map, and G is the number of channels);
[0041] Figure 4 It is a schematic structural diagram of the multi-channel convolution module of the present invention;
[0042] Figure 5 It is a schematic structural diagram of the fast pyramid pooling residual block of the present invention;
[0043] Figure 6 It is a diagram of the keycap image acquisition device in Embodiment 1 of the present invention;
[0044] Figure 7 It is an actual detection effect diagram of the keycap surface defect image of the C-YOLOv3 network model in Embodiment 1 of the present invention;
[0045] Figure 8 It is a schematic diagram of the evaluation index of the actual detection effect of the C-YOLOv3 network model in Embodiment 1 of the present invention; Detailed implementation manners
[0046] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with embodiments. The embodiments described herein are only used to illustrate and explain the present invention, and are not intended to limit the present invention.
[0047] Embodiment 1:
[0048] A keycap defect detection method based on C-YOLOv3 includes the following steps:
[0049] Step 1, obtain a keycap surface defect image, expand the keycap surface defect image, then perform defect feature annotation on the expanded keycap surface defect image, and finally use the labeled keycap surface defect image as a sample to divide the training set, validation set, and test set according to a ratio of 8:1:1;
[0050] Step 1.1, use a keycap image acquisition device to take pictures of defective keycaps to obtain keycap surface defect images. All keycap surface defect images form a keycap surface defect image set. The defects on the keycap surface include four categories: paint spots, cotton wool, particles, and indentations;
[0051] In this embodiment, a Hikvision industrial camera of model MV-CS200-10GC and a HIKROBOT lens of model MVL-KF1228M-12MP are used to take pictures of defective keycaps, and the obtained defective keycap surface image set contains 2046 defective keycap surface images;
[0052] Since the quantity distribution of various types of defects in the defective keycap surface image set is unbalanced (i.e., the quantity distribution of defective keycap surface images of paint spots, cotton wool, particles, and indentations is unbalanced), further processing is required.
[0053] Step 1.2: Adopt data augmentation technology to expand the data volume of defective keycap surface images, so that the quantity distribution of various types of defects is balanced (i.e., the quantity distribution of defective keycap surface images of paint spots, cotton wool, particles, and indentations is balanced);
[0054] The data augmentation technology is to use the image processing method of Alpha to process the defective keycap surface images to achieve the purpose of expanding the data volume;
[0055] Step 1.3: Then perform defect feature annotation on all the expanded defective keycap surface images;
[0056] Defect feature annotation is to use the open-source auxiliary image calibration software LabelImg to label the defects in each defective keycap image with corresponding labels. The labels include the category of the defect and the location information of the defect. The categories of the defects are divided into four categories: Grain (particles), Paint_spot (paint spots), Cotton (cotton wool), and Indentation (indentations), and the labels are generated in the form of txt files;
[0057] In this embodiment, the expanded defective keycap surface image set includes 7796 defective keycap surface images, and the quantity distribution of various types of defects is balanced.
[0058] Step 1.4: Make the labeled defective keycap surface images into a dataset according to the format of the VOC2007 dataset, that is, randomly divide the labeled defective keycap surface images into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0059] Step 2: Use the pytorch open-source deep learning framework to build a C-YOLOv3 network model. The C-YOLOv3 network model includes a backbone network and a detection network.
[0060] Step 2.1: Build a backbone network based on the Darknet-53 network of the YOLOv3 network model.
[0061] The backbone network includes a DBL structure (Darknetconv2D_BN_Leaky) and a residual network module. The residual network module includes 5 residual structures in the Darknet-53 network of the YOLOv3 network model and 3 residual structures embedded in the Darknet-53 network of the YOLOv3 network model.
[0062] The backbone network of the present invention embeds three residual structures on the basis of the Darknet-53 network of the YOLOv3 network model:
[0063] (1) A residual structure composed of a multi-scale attention mechanism residual block (EMA) is respectively connected behind the 3rd and 4th residual structures of the residual network module (ResNet) of the Darknet-53 network. The multi-scale attention mechanism residual block is as Figure 3 shown;
[0064] (2) A residual structure composed of a fast pyramid pooling residual block (SPPF) is connected behind the 5th residual structure of the residual network module of the Darknet-53 network. The fast pyramid pooling residual block is as Figure 5 shown.
[0065] The residual network module (ResNet) of the backbone network of the present invention includes a DBL structure (Darknetconv2D_BN_Leaky) and 8 residual structures;
[0066] The DBL structure of the backbone network includes a convolutional layer, a batch normalization layer (BN normalization layer), and a SiLU activation function layer arranged in sequence;
[0067] The number of residual units in the residual blocks corresponding to the 8 residual structures of the residual network module of the backbone network is 1-2-8-1-8-1-4-1 in sequence. Each residual block (Resblock_body) includes a zero-padding structure (zeropadding), a DBL structure, and multiple residual units (res_unit). Resn is the serial number of the residual block, and n takes 1-8, corresponding to 8 residual structures respectively.
[0068] Each residual unit processes the input through two DBL structures and then sums it with the input and outputs.
[0069] The residual structures composed of the multi-scale attention mechanism residual blocks (EMA) and the residual structures composed of the fast pyramid pooling residual blocks (SPPF) do not change the size of the input feature map. The other 5 residual structures will halve the size of the input feature map. For example, for an input feature map with a size of 416×416×3, after passing through 8 residual structures, the sizes of the output feature maps are 208×208×64, 104×104×128, 52×52×256, 52×52×256, 26×26×512, 26×26×512, 13×13×1024, and 13×13×1024 respectively. Among them, the output feature maps of the 4th residual structure, the 6th residual structure, and the 8th residual structure are used as the output of the backbone network and also as the input of the detection network.
[0070] The multi-scale attention mechanism residual block is based on the following formula:
[0071] Md = σ(f 1×1 (X Avg Pool(F))) + σ(f 1×1 (Y Avg Pool(F))) + F(1)
[0072] Mc(F) = σ(x c (σ(Avg Pool(Md)), f 3×3 (F))) + x c (σ(Avg Pool(f 3×3 (F))), Md)) + F(2)
[0073] In Equation (1), Md is the attention feature map and F is the input feature map. Avg Pool is the global pooling operation, that is, the average pooling encoding globally (including the X direction and the Y direction). Among them, the X direction is the width direction (W) of the feature map, and the Y direction is the height direction (H) of the feature map; X Avg Pool(F) represents the output encoding vector after the input feature map F undergoes average pooling encoding in the X direction; Y Avg Pool(F) represents the output encoding vector after the input feature map F undergoes average pooling encoding in the Y direction; σ represents the Sigmoid function, and f 1×1 represents the 1×1 convolution operation. The 1×1 convolution kernel enables the network to capture cross-channel features and represents the 1×1 branch.
[0074] In Equation (2), Mc(F) is the output feature map, and f 3×3 represents the 3×3 convolution operation, which is used to capture the multi-scale feature representation of the input feature map and represents the 3×3 branch; x c represents matrix multiplication.
[0075] The explanations of the formulas (i.e., Equation (1) and Equation (2)) of the multi-scale attention mechanism residual block are as follows:
[0076] The input feature map F is respectively encoded by average pooling in the X direction and the Y direction to obtain two output encoding vectors. Then, these two output encoding vectors are respectively mapped to attention scores in the range of 0 to 1 through the σ function (Sigmoid function). The input feature map F is adjusted by the attention scores (that is, adding the attention scores to the input feature map) to obtain the attention feature map Md. For the 3×3 branch, the vector f 3×3 (F) of the feature map output by the 3×3 convolution operation generates the attention weight score corresponding to the feature map through global pooling operation and then passes through the Sigmoid function. The attention weight score and the attention feature map Md are subjected to matrix multiplication x c operation to obtain the vector of the output feature map of the 3×3 branch. Similarly, for the 1×1 branch, the attention weight score corresponding to the feature map output after global pooling operation of the attention feature map Md and then passing through the Sigmoid function is subjected to matrix multiplication x c operation with the vector of the feature map output by the 3×3 convolution operation of the input feature map F to obtain the vector of the output feature map of the 1×1 branch. Finally, the vectors of the output feature maps of the two branches (3×3 branch and 1×1 branch) are added and then activated. The input feature map F is adjusted by the activated feature encoding vector to obtain the final output feature map Mc(F), and the output feature map Mc(F) is input to the next residual structure.
[0077] Fast Pyramid Pooling Residual Block: The output of the input feature map (i.e., the output feature map of the 7th residual structure of the backbone network) after passing through a ConvBNMish module passes through 3 MaxPool layers of size 5×5. Then, the outputs of the three MaxPool layers are connected and passed through a ConvBNMish module for output. Among them, the ConvBNMish module includes Conv convolution, BN normalization, and Mish activation function.
[0078] Step 2.2, construct a detection network, which includes a Pyramid Feature Network (FPN) and an output layer.
[0079] The input of the Pyramid Feature Network (FPN) is the input of the detection network, that is, the output feature maps of the 4th, 6th, and 8th residual structures in the backbone network. The Pyramid Feature Network (FPN) outputs three output feature maps.
[0080] The output layer of the detection network includes three multi-channel convolutional modules (C4) and three 1×1 convolutional layers (Conv). A 1×1 convolutional layer (Conv) is connected after each multi-channel convolutional module (C4). The three output feature maps output by the pyramid feature network (FPN) are respectively input into the three multi-channel convolutional modules (C4). The output layer of the detection network reduces the dimensions of the three output feature maps output by the pyramid feature network (FPN) through feature fusion operations to a fixed number of channels, obtaining three predicted feature maps with sizes of 52×52, 26×26, and 13×13 respectively.
[0081] The feature fusion operation is specifically as follows: As Figure 4 shown, the multi-channel convolutional module (C4) divides the output feature map of the pyramid feature network (FPN) into four groups according to the number of channels and inputs them into four channels respectively. Then, within each channel, residual convolutional operations are used to capture and process the features between different categories, obtaining a convolutional residual output. Finally, the convolutional residual outputs on the four channels are concatenated and passed through a 1×1 convolutional layer (Conv) to obtain the predicted feature map. Among them, the residual convolutional operation is to pass the feature map of each channel through two 1×1 convolutional layers (Conv), and then connect the input and output of the first 1×1 convolutional layer (Conv) with the output of the second 1×1 convolutional layer (Conv) to obtain the convolutional residual output.
[0082] The three output feature maps with different sizes output by the 4th residual structure, the 6th residual structure, and the 8th residual structure of the backbone network are used as the original feature maps and input into the detection network. The sizes of the original feature maps are 52×52, 26×26, and 13×13 respectively; after passing through the detection network, three predicted feature maps are obtained, and the sizes of the three predicted feature maps are 52×52, 26×26, and 13×13 respectively, which are the same as the sizes of the original feature maps.
[0083] Step 3: Input the image of the surface defect of the keycap to be detected into the constructed C-YOLOv3 network model to obtain the predicted feature map, and then obtain the category of the defect in the image of the surface defect of the keycap to be detected and the four offset parameters of the corresponding prediction box through the predicted feature map.
[0084] The specific method for obtaining the category of the defect in the image of the surface defect of the keycap to be detected and the four offset parameters of the corresponding prediction box through the predicted feature map is as follows:
[0085] After the detection network processes, the three predicted feature maps of the detection network are respectively segmented into N×N grids, where N is the size of the predicted feature map of the detection network. The K-means clustering algorithm is used for the data set to generate prior boxes. Each grid presets three prior boxes with different scales. Therefore, each predicted feature map needs to predict N×N×[S×(4 + 1 + C)] parameters, where S is the number of prior boxes predicted by each grid. In this embodiment, S is 3; for each prior box, 4 + 1 + C parameters need to be predicted. The 4 + 1 + C parameters include four offset parameters of each predicted box, a confidence score (the confidence is 0 when there is no defect in the grid, and when there is a defect in the grid, the confidence becomes the IoU value between the predicted box and the ground truth box), and C class scores, where C is the number of defect classes. In this embodiment, C = 4. The class with the highest score is the class of the finally predicted defect. Subsequently, by setting a confidence threshold, the predicted boxes with lower scores are filtered out. Finally, non-maximum suppression (NMS) is used to filter out redundant predicted boxes to obtain the final prediction result, that is, the four offset parameters of a predicted box and the class of a defect. The four offset parameters are the offsets between the four corners of the predicted box and the four corners of the ground truth box.
[0086] Step 4: Use the training set and train the C-YOLOv3 network model by minimizing the loss function Loss of the C-YOLOv3 network model. After the training is completed, save the C-YOLOv3 network model parameters to obtain the keycap surface defect detection model.
[0087] The loss function Loss of the C-YOLOv3 network model consists of three parts: classification loss, localization loss, and confidence loss. The loss function Loss of the C-YOLOv3 model is calculated based on the following formula:
[0088] Loss = Lbox + Lobj + Lcls (3)
[0089] Lbox is the localization loss of the four types of defects, Lobj is the confidence loss of the four types of defects, and Lcls is the classification loss of the four types of defects;
[0090] The localization loss Lbox of the four types of defects uses the CIoU loss function to optimize the bounding box regression. The CIoU loss function takes into account the complete intersection between bounding boxes, the aspect ratio of bounding boxes, and the center point distance between bounding boxes, and can more accurately measure the similarity between bounding boxes. Moreover, the CIoU loss function has the advantages of translational invariance, scale invariance, and greater sensitivity to overlapping target boxes, which can improve the accuracy and robustness of the object detection task.
[0091] The CIoU loss function is based on the following formula:
[0092]
[0093] Where IoU is the intersection over union of the predicted bounding box and the ground truth bounding box, b is the center point of the predicted bounding box, b gt is the center point of the ground truth bounding box, ρ is the distance between the center points of the predicted bounding box and the ground truth bounding box, and c is the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box. α is a tuning factor used to balance the relative weights of the center point distance and the aspect ratio penalty term. v is the aspect ratio penalty term. The formulas for α and v are shown as follows:
[0094]
[0095]
[0096] Where w1 and h1 represent the width and height of the predicted bounding box respectively, represent the width and height of the ground truth bounding box respectively.
[0097] The optimizer of the C-YOLOv3 network model is the Stochastic Gradient Descent with Momentum (Momentum SGD) optimizer, which helps to avoid getting stuck in local optima and find better global optima.
[0098] Transfer learning is adopted during the training process of the C-YOLOv3 network model, and the training process is divided into two stages. In the first stage, all layers in the C-YOLOv3 network model are frozen (i.e., the parameters of the 8 residual structures in the C-YOLOv3 network model remain unchanged), the learning rate is set to 0.001, and the number of training epochs is set to 50 for training; after completing the training task of the first stage, directly start the second stage of training. In the second stage, all layers in the C-YOLOv3 network model are unfrozen (i.e., the parameters of the 8 residual structures in the C-YOLOv3 network model are changed), the learning rate is 0.0001, the number of training epochs is set to 250, and the batch size is set to 8 for training. At the same time, the cosine learning rate decay strategy and the early stopping training strategy are also adopted during the training process of the C-YOLOv3 network model.
[0099] During the training process of the C-YOLOv3 network model, masico data augmentation and mixup data augmentation are also added, and mixup data augmentation is only performed on the images after mosaic data augmentation. The first 70% of the training epochs in each stage use masico data augmentation. One training epoch has 37 steps. By default, 50% of the images use mosaic data augmentation in each step, and there is a 50% probability of using mixup data augmentation after mosaic data augmentation. The total probability of mixup data augmentation is the product of the probabilities of mosaic data augmentation and mixup data augmentation. The images of keycap surface defects can be further expanded. The weight decay coefficient is set to 0.0005 to prevent overfitting.
[0100] The cos learning rate decay strategy is that after a certain number of training epochs of the C-YOLOv3 network model, when the loss function of the C-YOLOv3 network model no longer changes, the learning rate is reduced to obtain better training results.
[0101] The C-YOLOv3 network proposed in the present invention is compared with the implementation results of the original YOLOv3 network, YOLOv7, and Faster-rcnn. In this embodiment, recall, precision, average precision AP (Average Precision), and mean average precision mAP are used as evaluation indicators.
[0102] Recall represents the ratio of the images correctly detected in each category to all the images tested. Precision represents the average detection accuracy of the defects in each category. Average precision AP represents the prediction accuracy of each individual category. Mean average precision mAP represents the average value of the detection accuracies of all categories.
[0103] Precision is based on the following formula:
[0104]
[0105] In the formula, TP represents the number of samples predicted as positive samples by the keycap surface defect detection model and actually being positive samples, and FP represents the number of samples predicted as positive samples by the model but actually being negative samples;
[0106] Recall is based on the following formula:
[0107]
[0108] In the formula, FN represents the number of samples predicted as negative samples by the keycap surface defect detection model but actually being positive samples;
[0109] The average precision AP is based on the following formula:
[0110]
[0111] Where R is the proportion of defects correctly predicted by the keycap surface defect detection model among all true defects. P(R) is a curve with the recall as the abscissa and the precision as the ordinate;
[0112] The mean average precision mAP is based on the following formula:
[0113]
[0114] Where K is the total number of samples. In this embodiment, the samples are keycap surface defect images;
[0115] Table 1 is a result graph comparing the C-YOLOv3 network proposed by the present invention with the original YOLOv3 network, YOLOv7, and Faster-rcnn in terms of detection accuracy and detection speed
[0116]
[0117] In Table 1, P is the precision and fps is the frame rate, indicating the detection speed.
[0118] It should be noted that the embodiments described in the present invention are only illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A keycap defect detection method based on C-YOLOv3, characterized in that, It includes the following steps: Step 1: Obtain the surface defect image of the keycap, expand the surface defect image of the keycap, then perform defect feature annotation on the expanded surface defect image of the keycap, and finally divide the training set, validation set, and test set according to the labeled surface defect image of the keycap; Step 2: Construct a C-YOLOv3 network model, and the C-YOLOv3 network model includes a backbone network and a detection network; Step 3: Input the surface defect image of the keycap to be detected into the constructed C-YOLOv3 network model to obtain a predicted feature map, and then obtain the category of the defect in the surface defect image of the keycap to be detected and the four offset parameters of the corresponding prediction box through the predicted feature map; Step 4: Use the training set and minimize the loss function Loss of the C-YOLOv3 network model to train the C-YOLOv3 network model, and save the parameters of the C-YOLOv3 network model after training to obtain a surface defect detection model of the keycap; The backbone network includes a DBL structure and a residual network module, and the residual network module includes 5 residual structures in the Darknet-53 network of the YOLOv3 network model and 3 residual structures embedded in the Darknet-53 network of the YOLOv3 network model; The DBL structure includes a convolutional layer, a batch normalization layer, and a SiLU activation function layer arranged in sequence; The 3 residual structures embedded in the Darknet-53 network of the YOLOv3 network model respectively include: a residual structure composed of multi-scale attention mechanism residual blocks connected behind the 3rd and 4th residual structures of the residual network module in the Darknet-53 network, and a residual structure composed of a fast pyramid pooling residual block connected behind the 5th residual structure of the residual network module in the Darknet-53 network; The number of residual units in the residual blocks corresponding to the 8 residual structures of the residual network module of the backbone network is 1-2-8-1-8-1-4-1 in sequence, and each residual block includes a zero-padding structure, a DBL structure, and multiple residual units; The detection network includes a pyramid feature network and an output layer, and the output layer of the detection network includes three multi-channel convolutional modules and three 1×1 convolutional layers, and each multi-channel convolutional module is connected to a 1×1 convolutional layer; The multi-scale attention mechanism residual block is based on the following formula: Md = σ(f 1×1 (X Avg Pool(F))) + σ(f 1×1 (Y Avg Pool(F))) + f Mc(F) = σ(x c (σ(Avg Pool(Md)), f 3×3 (F)) + x c (σ(Avg Pool(f 3×3 (F))), Md)) + F Wherein, Md is the attention feature map, and F is the input feature map; Avg Poll is the global average pooling encoding, X AvgPool(F) is the output encoding vector after the input feature map F undergoes the average pooling encoding in the X direction, and the X direction is the width direction of the feature map; Y Avg Pool(F) is the output encoding vector after the input feature map F undergoes the average pooling encoding in the Y direction, and the Y direction is the height direction of the feature map; σ is the Sigmoid function, and f 1×1 is a 1×1 convolution operation; Mc(F) is the output feature map, and f 3×3 is a 3×3 convolution operation; x c is matrix multiplication; The fast pyramid pooling residual block is to pass the output feature map of the 7th residual structure of the backbone network through a ConvBNMish module, then through 3 MaxPool layers with a size of 5×5, and then connect the outputs of the three MaxPool layers and output through a ConvBNMish module; The ConvBNMish module includes Conv convolution, BN normalization, and Mish activation function.
2. The keycap defect detection method based on C-YOLOv3 according to claim 1, characterized in that, The defect feature annotation is to label the defects in each defective keycap image, and the label includes the category of the defect and the position information of the defect. The categories of the defects are divided into paint spots, cotton fluffs, particles, and indentations.
3. The keycap defect detection method based on C-YOLOv3 according to claim 2, characterized in that, The loss function Loss of the C-YOLOv3 network model consists of three parts: classification loss, localization loss, and confidence loss. The loss function Loss of the C-YOLOv3 model is calculated based on the following formula: Loss = Lbox + Lobj + Lcls Lbox is the localization loss of the defect, Lobj is the confidence loss of the defect, and Lcls is the classification loss of the defect; The localization loss Lbox of the defect adopts the CIoU loss function, and the CIoU loss function is based on the following formula: where IoU is the intersection over union of the predicted bounding box and the ground truth bounding box, b is the center point of the predicted bounding box, and b gt is the center point of the ground truth bounding box, ρ is the distance between the center points of the predicted bounding box and the ground truth bounding box, c is the diagonal distance of the smallest closed region that can contain both the predicted bounding box and the ground truth bounding box, α is a tuning factor, v is the aspect ratio penalty term, and α and v are based on the following formulas respectively: where w1 and h1 are the width and height of the predicted bounding box, respectively, which are the width and height of the ground truth bounding box, respectively.
4. The keycap defect detection method based on C-YOLOv3 according to claim 3, characterized in that, The optimizer of the C-YOLOv3 network model is the SGD model optimizer with momentum; during the training process of the C-YOLOv3 network model, transfer learning is adopted and the training process is divided into two stages. In the first stage, the parameters of the 8 residual structures in the C-YOLOv3 network model remain unchanged, the learning rate is set to 0.001, and the number of training epochs is set to 50 for training; In the second stage, the parameters of the 8 residual structures in the C-YOLOv3 network model are changed, the learning rate is 0.0001, the number of training epochs is set to 250, and the batch size is set to 8 for training.
Citation Information
Patent Citations
Metal surface defect detection method based on improved YOLOv3
CN113920400A