An improved clustering algorithm based on Shape-GIoU
By introducing Shape-GIoU improved clustering algorithm in the YOLOv4 algorithm and combining with the K-means++ method to cluster the anchor box, the problem that the existing IoU calculation method cannot effectively express the distance and aspect ratio of the center point of the box is improved, and the accuracy and recall rate of target detection are improved.
Patent Information
- Application Number
- CN202210007090.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-05
AI Technical Summary
The existing IoU parameters have defects in calculating the similarity between the anchor frame and the real frame, and cannot effectively express the distance and aspect ratio information of the center point of the two boxes, resulting in insufficient target detection accuracy and recall.
A clustering algorithm based on Shape-GIoU is proposed. By introducing the proportional coefficient λ/β, the calculation method of IoU is improved, and the impact of non-coverage area and aspect ratio is considered, the problem of GIoU degradation to IoU is solved, and it is combined with the K-means++ clustering method to cluster the initial Anchorbox.
The detection accuracy and recall rate of the YOLOv4 algorithm in the data set are improved. Through the improved Shape-GIoU calculation method, the similarity between the anchor box and the real box can be expressed more accurately, improving the performance of the model.
Smart Images

Figure CN114494756B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image recognition, and in particular relates to a clustering algorithm based on Shape-GIoU improvement. By improving the IoU calculation method Shape-GIoU, the K-means++ clustering method is combined with Shape-GIoU to cluster the initial Anchorbox, thereby improving the detection accuracy and recall rate of the YOLOv4 algorithm in the data set. Background Art
[0002] In recent years, deep learning models have gradually become the main algorithms in the field of target detection, among which convolutional neural networks have achieved remarkable results in the field of target detection. Target detection models based on deep learning can be divided into two types: one is a two-stage detection algorithm based on Region Proposal, such as R-CNN, Fast R-CNN, Faster R-CNN, etc. This algorithm needs to generate candidate regions first, and then classify and regress the target candidate regions. The other is a single-stage detection algorithm such as the YOLO series. Compared with the two-stage target detection algorithm, the YOLO series algorithm abandons the stage of generating candidate regions and uses anchor boxes to replace candidate regions for final regression. Before YOLOv4 was proposed, the parameter for evaluating the similarity between the anchor box and the real box was IoU, but the IoU parameter has many defects: first, when the two boxes do not intersect, IoU is always 0, and IoU cannot express the information of the distance between the center points of the two boxes; second, because IoU only contains the information of the overlapping area of the two boxes, when the IoU of the two boxes is equal, it is impossible to determine their specific shape information, such as aspect ratio.
[0003] Therefore, the present invention proposes a clustering algorithm based on Shape-GIoU improvement, which makes up for the problem that the original GIoU degenerates into IoU by embedding the improved IoU calculation method Shape-GIoU. By introducing a proportional coefficient λ / β, the influence of the aspect ratio is introduced into the formula, which can solve the problem that GIoU degenerates. The influence of the non-covered area and the aspect ratio is taken into account, and the detection accuracy and recall rate of the YOLOv4 network are improved. Summary of the invention
[0004] In view of the above problems, the purpose of the present invention is to provide an improved clustering algorithm for IoU calculation method.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] An improved clustering algorithm based on Shape-GIoU introduces a scaling factor λ / β, considers the influence of non-coverage area and aspect ratio, solves the problem of original GIoU degenerating into IoU, and improves the detection performance of YOLOv4 algorithm.
[0007] The improved clustering algorithm comprises the following steps:
[0008] Step 1: Download the public datasets AIZOO and RMFD for face mask detection. Since the quality of the dataset photos is also uneven, the present invention selects face photos with a resolution greater than 608*608 wearing masks or not wearing masks from the public AIZOO and RMFD face recognition datasets to construct the face mask dataset used by the present invention; create a corresponding xml file for each sample image in the dataset, and store the path of each sample image, the category information of the target of interest in the sample image, and the xml file. i The coordinate information of the upper left corner and lower right corner of the target, as well as the resolution of the sample image are stored in the corresponding xml file; the xml file data is preprocessed, first the coordinate values of the target of interest in the sample image are obtained, and then the coordinate data is converted into the width and height of the real frame. The calculation formula is: width = lower right corner horizontal coordinate - upper left corner horizontal coordinate, height = lower right corner vertical coordinate - upper left corner vertical coordinate, and finally the width and height values are normalized, and the normalized width and height values = width and height values / input image resolution; the download address of the AIZOO dataset is: https: / / github.com / AIZOOTech / FaceMaskDetection; the download address of the RMFD dataset is: https: / / github.com / X-zhangyang / Real-World-Masked-Face-Dataset;
[0009] Since there are few public data sets for face mask detection at present, and the quality of photos is also uneven, the present invention selects face photos with a resolution greater than 608*608 wearing masks or without masks from the public AIZOO and RMFD face recognition data sets, and constructs the face mask data set used by the present invention; the data set of the present invention is divided into two categories: face targets wearing masks and face targets without masks, including a total of 11,208 photos, and the data set contains 7,933 face targets wearing masks in different scenes and 13,651 face targets without masks, and the test set, validation set and training set are divided according to the ratio of 6:2:2;
[0010] Step 2: Select the normalized width and height data (w j ,h j ) as the initial cluster center y j(w j ,h j );
[0011] Step 3: For each sample x in the data set i (w i ,h i ), calculate its difference with the selected cluster center y j (w j ,h j ) distance; the distance used d ij The calculation method is:
[0012] d ij =1-Shape-GIoU ij
[0013] Shape-GIoU of the present invention ij The calculation method is as follows:
[0014]
[0015] x i ∩y j =min(w i ,w j )×min(h i ,h j )
[0016] x i ∪y j =w i ×h i +w j ×h j -min(w i ,w j )×min(h i ,h j )
[0017] A ij =max(w i ,w j )×max(h i ,h j )
[0018]
[0019]
[0020] where x i ∩y j represents the area of the union between the i-th ground-truth box and the j-th cluster center, x i ∪y jrepresents the area of the intersection between the i-th ground-truth box and the j-th cluster center, A ij represents the area of the minimum bounding rectangle between the i-th ground-truth box and the j-th cluster center, where i = 0, 1, ..., 21583, j = 1, 2, ..., 9;
[0021] It can be seen from the above formula that when A ij =(x i ∪y j ), the two boxes completely intersect. At this time, GIoU will degenerate into IoU, and a proportional coefficient is introduced. This coefficient will introduce the influence of the ratio of the aspect ratio of the two frames into the formula, which can solve the problem of GIoU degenerating into IoU;
[0022] Step 4: Calculate each sample x i (w i ,h i ) is selected as the next cluster center Randomly generate new cluster center y j+1 (w j+1 ,h j+1 );
[0023] Step 5: Repeat steps 3 and 4 until K cluster centers are selected, and use these K initial cluster centers to run the standard k-means algorithm to recalculate the cluster centers of each category;
[0024] Because the YOLOv4 algorithm has three detection scales and each scale has 3 anchors, we select K=9 to get 9 cluster centers. Multiply the values of the cluster centers by the resolution of the output image to get the 9 anchor values required by the YOLOv4 algorithm.
[0025] Step 6: Use K-means++ clustering algorithm to obtain anchor value, train through standard YOLOv4 network and save test result test1, use the method of the present invention to obtain new anchor value, train through standard YOLOv4 network and save test result test2; compare the test results of the two experiments, and the comparison indicators are mAP and Recall respectively; download the standard YOLOv4 network and compile it, the download address of YOLOv4 is: https: / / github.com / AlexeyAB / darknet, change the training set, validation set, and test set directories in the voc.data file in the data folder to download The address of the data set, and specify the number of categories and category names; modify the anchor, classes, and filters parameters in the YOLOv4.cfg file in the cfg folder, where the parameter anchor is the 9 anchor values obtained using the K-means++ clustering algorithm; classes is the number of categories of the data set constructed by the present invention, which are respectively face targets wearing masks and face targets not wearing masks, that is, classes = 2; filters = (number of categories + 5) * 3 = 21; start training the YOLOv4 network, and after the network training is completed, select the YOLOv4_best.weights file in the backup folder as the test weight Q 1 , perform the test and obtain the test result test1; use the method of the present invention to obtain a new anchor value, modify the anchor parameter in the YOLOv4.cfg file in the cfg folder, train the YOLOv4 network, and obtain the test weight Q 2 , perform the test and obtain the test result test2.
[0026] The method of the present invention proposes a clustering algorithm based on Shape-GIoU improvement. By improving the IoU calculation method Shape-GIoU, the K-means++ clustering method is combined with Shape-GIoU to cluster the initial Anchor box. Through experiments and comparison with the K-means++ clustering method, the method of the present invention improves the mAP value and Recall value of the YOLOv4 algorithm in the data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0028] Figure 1is a flow chart of the method of the present invention;
[0029] Figure 2 These are some sample images in the training set;
[0030] Figure 3 This is a comparison chart of various IoU values;
[0031] Figure 4 It is the Shape-GIoU value calculation graph;
[0032] Figure 5 It is a partial detection result diagram of the YOLOv4 model using the method of the present invention;
[0033] Figure 6 The overall performance of the YOLOv4 model using the K-means++ clustering algorithm and the YOLOv4 model using the method of the present invention on the validation data set; DETAILED DESCRIPTION
[0034] In order to make the above and other purposes, features and advantages of the present invention more obvious, the embodiments of the present invention are specifically cited below, and the accompanying drawings are used to provide a detailed description as follows:
[0035] Reference Figure 1 , the implementation steps of the present invention are as follows:
[0036] Step 1: Download the public datasets AIZOO and RMFD for face mask detection. Since the quality of the dataset photos is also uneven, the present invention selects face photos with a resolution greater than 608*608 wearing masks or not wearing masks from the public AIZOO and RMFD face recognition datasets to construct the face mask dataset used by the present invention; create a corresponding xml file for each sample image in the dataset, and store the path of each sample image, the category information of the target of interest in the sample image, and the xml file. i The coordinate information of the upper left corner and lower right corner of the target, as well as the resolution of the sample image are stored in the corresponding xml file; the xml file data is preprocessed, first the coordinate values of the target of interest in the sample image are obtained, and then the coordinate data is converted into the width and height of the real frame, the calculation formula is: width = lower right corner horizontal coordinate - upper left corner horizontal coordinate, height = lower right corner vertical coordinate - upper left corner vertical coordinate, finally the width and height values are normalized, the calculation formula is: normalized width and height value = width and height value / input image resolution; the download address of the AIZOO dataset is: https: / / github.com / AIZOOTech / FaceMaskDetection; the download address of the RMFD dataset is: https: / / github.com / X-zhangyang / Real-World-Masked-Face-Dataset;
[0037] The dataset of the present invention is divided into two categories: face targets wearing masks and face targets not wearing masks, including 11,208 photos in total. The dataset contains 7,933 face targets wearing masks in different scenes and 13,651 face targets not wearing masks. The test set, validation set and training set are divided in a ratio of 6:2:2.
[0038] Figure 2 It is a sample image of some training sets in the data set used in the present invention, which represents the universality of the target detection object, and trains different images in different scenes and angles;
[0039] Step 2: Select the normalized width and height data (w j ,h j ) as the initial cluster center y j (w j ,h j );
[0040] Step 3: For each sample x in the data set i (w i ,h i ), calculate its difference with the selected cluster center y j (w j ,h j ) distance; the distance used d ij The calculation method is:
[0041] d ij =1-Shape-GIoU ij
[0042] Reference Figure 4 , the Shape-GIoU calculation method of the present invention is shown in the following formula:
[0043]
[0044] x i ∩y j =min(w i ,w j )×min(h i ,h j )
[0045] x i ∪y j =w i ×h i +w j ×h j -min(w i ,w j )×min(hi ,h j )
[0046] A ij =max(w i ,w j )×max(h i ,h j )
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] where x i ∩y j represents the area of the union between the ground-truth box and the cluster center, x i ∪y j Represents the area of the intersection between the true box and the cluster center, A ij Represents the area of the minimum enclosing rectangle of the true box and the cluster center,
[0053] It can be seen from the above formula that when A ij =(x i ∪y j ), the two boxes completely intersect. At this time, GIoU will degenerate into IoU, and a proportional coefficient is introduced. This coefficient will introduce the influence of the ratio of the aspect ratio of the two frames into the formula, which can solve the problem of GIoU degenerating into IoU;
[0054] Step 4: Calculate each sample x i (w i ,h i ) is selected as the next cluster center Randomly generate new cluster center y j+1 (w j+1 ,h j+1 );
[0055] Step 5: Repeat steps 3 and 4 until K cluster centers are selected, and use these K initial cluster centers to run the standard k-means algorithm to recalculate the cluster centers of each category;
[0056] Because the YOLOv4 algorithm has three detection scales and each scale has 3 anchors, we select K=9 to get 9 cluster centers. Multiply the values of the cluster centers by the resolution of the output image to get the 9 anchor values required by the YOLOv4 algorithm.
[0057] Step 6: Use K-means++ clustering algorithm to obtain anchor value, train with standard YOLOv4 network and save test result test1, use the method of the present invention to obtain new anchor value, train with standard YOLOv4 network and save test result test2; compare the test results of the two experiments, and the comparison indicators are mAP and Recall respectively; download standard YOLOv4 network and compile it. The download address of standard YOLOv4 network is: https: / / github.com / AlexeyAB / darknet), change voc.data in the data folder The training set, validation set, and test set directories in the file are the addresses of the downloaded data sets, and the number of categories and category names are specified; modify the anchor, classes, and filters parameters in the YOLOv4.cfg file in the cfg folder, where the parameter anchor is the 9 anchor values obtained using the K-means++ clustering algorithm; classes is the number of categories of the data set constructed by the present invention, which are respectively face targets wearing masks and face targets not wearing masks, that is, classes = 2; filters = (number of categories + 5) * 3 = 21; start training the YOLOv4 network, and after the network training is completed, select the YOLOv4_best.weights file in the backup folder as the test weight Q 1 , perform the test and obtain the test result test1; use the method of the present invention to obtain a new anchor value, modify the anchor parameter in the YOLOv4.cfg file in the cfg folder, train the YOLOv4 network, and obtain the test weight Q 2 , perform the test and obtain the test result test2.
[0058] The invention is further described below in conjunction with a simulation example.
[0059] Simulation example:
[0060] The present invention uses a face mask dataset as a training set, a validation set, and a test set, and provides some detection effect diagrams using the method of the present invention.
[0061] Figure 2These are some sample images in the training set. Some test data in the face mask dataset are randomly selected as the result display. Images with different backgrounds, different types of masks, different target sizes, different angles, and different target densities are selected to demonstrate the universality of the test results.
[0062] Figure 3 The figure shows the calculation diagram of the method of the present invention based on Shape-GIoU, where the dotted box represents the predicted box and the solid box represents the real box. Figure 4 The parameter of (a) is w i =4,h i =8,w j =2,h j =4; the parameter of 4(b) is w i =4,h i =8, Figure 4 (a) Results using different calculation methods:
[0063] x i ∩y j =min(w i ,w j )×min(h i ,h j )=2×4=8
[0064] x i ∪y j =w i ×h i +w j ×h j -min(w i ,w j )×min(h i ,h j )
[0065] =4×8+2×4-2×4=32
[0066] A ij =max(w i ,w j )×max(h i ,h j )=4×8=32
[0067]
[0068]
[0069]
[0070]
[0071]
[0072] Figure 4 (b) Results using different calculation methods:
[0073]
[0074]
[0075]
[0076] It can be seen that when the predicted box is completely wrapped by the real box, for the case where the predicted box accounts for the same proportion of the real box but has a different aspect ratio, the Shape-GIoU proposed by the method of the present invention can distinguish them well, but the current calculation method can no longer distinguish them.
[0077] Figure 3 : This is a comparison diagram of the Shape-GIoU calculation method of the present invention and the current calculation method, where the dotted box represents the predicted box and the solid box represents the real box. It can be seen that when the predicted box is completely wrapped by the real box, the current calculation method can no longer distinguish it. For the case where the proportion of the predicted box to the real box is the same but the aspect ratio is different, the Shape-GIoU proposed by the method of the present invention can distinguish them well.
[0078] Figure 5 These are some of the detection results of the original YOLOv4 model. We selected detection images with different backgrounds, different types of masks, different target sizes, different angles, and different target densities to demonstrate the universality of the original detection model. It can be seen that the basic category detection effect of the objects in the image is good.
[0079] Figure 6 It is the overall performance of the YOLOv4 model using the K-means++ clustering algorithm and the YOLOv4 model using the method of the present invention on the test data set. It can be seen that the mAP and Recall of the method of the present invention on the test set are improved.
[0080] In summary, the simulation experiments show that the YOLOv4 model of the method of the present invention can distinguish the predicted box from the real box in the same proportion but with different aspect ratios when the predicted box is completely wrapped by the real box. The Shape-GIoU proposed by the present invention can improve the detection performance of the YOLOv4 algorithm model. The method of the present invention can also be used in combination with the classic algorithm model to improve the detection performance of the algorithm model.
Claims
1. A clustering method based on Shape-GIoU improvement, Features: Step 1: Download the public datasets AIZOO and RMFD for face mask detection, select face photos with or without masks with a resolution greater than 608*608, build the face mask dataset used, and preprocess the data in the dataset to obtain normalized width and height data; Step 2: Select the normalized width and height data (w j ,h j ) as the initial cluster center y j (w j ,h j ); Step 3: For each sample x in the data set i (w i ,h i ), calculate its difference with the selected cluster center y j (w j ,h j ) distance; the distance used d ij The calculation method is: d ij =1-Shape-GIoU ij Shape-GIoU ij The calculation method is as follows: x i ∩y j =min(w i ,w j )×min(h i ,h j ) x i ∪y j =w i ×h i +w j ×h j -min(w i ,w j )×min(h i ,h j ) A ij = max(w i , w j ) × max(h i , h j ) where x i ∩y j represents the area of the union between the ground-truth box and the cluster center, x i ∪y j Represents the area of the intersection between the true box and the cluster center, A ij Represents the area of the minimum enclosing rectangle of the true box and the cluster center, It can be seen from the above formula that when A ij =(x i ∪y j ), the two boxes completely intersect. At this time, GIoU will degenerate into IoU, and a proportional coefficient λ is introduced. ij / β ij , this coefficient will introduce the influence of the ratio of the aspect ratio of the two frames into the formula, which can solve the problem of GIoU degenerating into IoU; Step 4: Calculate each sample x i (w i ,h i ) is selected as the next cluster center Randomly generate new cluster centers y j+1 (w j+1 ,h j+1 ); Step 5: Repeat steps 3 and 4 until k cluster centers are selected, and use these k initial cluster centers to run the standard k-means algorithm to recalculate the cluster centers of each category; Step 6. Use K-means++ clustering algorithm to get anchor value, train with standard YOLOv4 network and save test result test1. Use improved clustering method based on Shape-GIoU to get new anchor value, train with standard YOLOv4 network and save test result test2. Compare the test results of the two experiments, and the comparison indicators are mAP and Recall respectively. In the above steps, i=0,1,...,21583 represents the label of the sample data, and j=1,2,...,9 represents the label of the cluster center.
2. According to claim 1, a clustering method based on Shape-GIoU improvement, It is characterized in that Step 1. Download the public datasets AIZOO and RMFD for face mask detection. Since the quality of the dataset photos is also uneven, select face photos with a resolution greater than 608*608 or without masks from the public AIZOO and RMFD face recognition datasets to build the face mask dataset used; create a corresponding xml file for each sample image in the dataset, and store the path of each sample image, the category information of the target of interest in the sample image, and the xml file. i The coordinate information of the upper left corner and lower right corner of the target, as well as the resolution of the sample image are stored in the corresponding XML file; Preprocess the xml file data. First, obtain the coordinate values of the target of interest in the sample image. Then convert the coordinate data into the width and height of the real frame. The calculation formula is: width = lower right corner horizontal coordinate - upper left corner horizontal coordinate, height = lower right corner vertical coordinate - upper left corner vertical coordinate. Finally, normalize the width and height values. The calculation formula is: normalized width and height values = width and height values / input image resolution. The download address of the AIZOO dataset is: https: / / github.com / AIZOOTech / FaceMaskDetection; the download address of the RMFD dataset is: https: / / github.com / X-zhangyang / Real-World-Masked-Face-Dataset. The dataset used is divided into two categories: face targets wearing masks and face targets not wearing masks, containing a total of 11,208 photos. The dataset contains 7,933 face targets wearing masks in different scenes and 13,651 face targets not wearing masks. The test set, validation set and training set are divided in a ratio of 6:2:
2.
3. According to claim 1, a clustering method based on Shape-GIoU improvement, It is characterized in that Step 5. Repeat steps 3 and 4 until K cluster centers are selected, and use these K initial cluster centers to run the standard k-means algorithm to recalculate the cluster centers of each category; because the YOLOv4 algorithm has three detection scales, each scale has 3 anchors, so we take K=9 to get 9 cluster centers, multiply the value of the cluster center by the resolution of the output image, and finally we get the 9 anchor values required by the YOLOv4 algorithm.
4. According to claim 1, a clustering method based on Shape-GIoU improvement, It is characterized in that Step 6. Use K-means++ clustering algorithm to get anchor value, train with standard YOLOv4 network and save test result test1, use improved clustering method based on Shape-GIoU to get new anchor value, train with standard YOLOv4 network and save test result test2; compare the test results of the two experiments, the comparison indicators are mAP and Recall, the standard YOLOv4 network download address is: https: / / github.com / AlexeyAB / darknet, download the standard YOLOv4 network and compile it, change the training set and validation set in the voc.data file in the data folder The proof set and test set directories are the addresses of the downloaded datasets, and the number of categories and category names are specified; modify the anchor, classes, and filters parameters in the YOLOv4.cfg file in the cfg folder, where the parameter anchor is the 9 anchor values obtained using the K-means++ clustering algorithm; classes is the number of categories for constructing the dataset, which are face targets with masks and face targets without masks, that is, classes = 2; filters = (number of categories + 5) * 3 = 21; start training the YOLOv4 network, and after the network training is completed, select the YOLOv4_best.weights file in the backup folder as the test weight Q 1 , test and get the test result test1; use the new method to get the new anchor value, modify the anchor parameter in the YOLOv4.cfg file in the cfg folder, train the YOLOv4 network, and get the test weight Q 2 , perform the test and obtain the test result test2.
Citation Information
Patent Citations
Railway overhead line system bird nest detection method
CN112949634A
Kiwi fruit leaf disease detection method based on improved YOLOv4-Tiny feature fusion
CN113379727A