A wheat ear detection method based on improved YOLOv5

By improving the YOLOv5 algorithm, adding four-scale feature detection, introducing CA attention mechanism and CIOU_Loss loss function, the problem of low recognition accuracy of small and medium-sized wheat ear detection is solved, and higher recognition accuracy and detection accuracy are achieved.

CN114973002BActive Publication Date: 2025-05-02ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210705045.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-05-02
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

The existing wheat ear detection methods have the problem of low accuracy and easy to be blocked when identifying small targets, making it difficult to achieve fast and accurate wheat ear detection.

Method used

By improving the YOLOv5 algorithm, four-scale feature detection is added, the CA attention mechanism and CIOU_Loss loss function are introduced, and the small object recognition accuracy and the accuracy of inspection box detection are improved.

Benefits of technology

It improves the recognition accuracy of small targets, improves the accuracy of inspection frame detection, and achieves better detection effect on wheat ears.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973002B_ABST
    Figure CN114973002B_ABST
Patent Text Reader

Abstract

The present invention relates to a wheat ear detection method based on improved YOLOv5, which is characterized by comprising: obtaining wheat ear images and annotating the wheat ear images to obtain a wheat ear image data set, and dividing the wheat ear image data set into a training set and a test set; constructing a YOLOv5 network model; improving the YOLOv5 network model to obtain an improved YOLOv5 network model; inputting the training set into the improved YOLOv5 network model to train the improved YOLOv5 network model; and evaluating and testing the improved YOLOv5 network model. The present invention uses four-scale feature detection and increases the shallow detection scale to improve the recognition accuracy of small targets; the present invention introduces a CA attention mechanism to improve the feature extraction capability of the algorithm; the present invention introduces CIOU_Loss as the Bounding Box Regression Loss of the algorithm loss function to improve the accuracy of the inspection frame detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of smart agriculture and information-based agriculture, and in particular to a wheat ear detection method based on improved YOLOv5. Background Art

[0002] Wheat is an important food crop with the largest trade volume in the world and one of the main food crops in my country. In order to ensure national food security, formulate reasonable food prices and macro-control policies, it is necessary to timely and accurately predict and estimate the expected yield of crops such as wheat. At present, wheat ear technology mostly uses manual survey methods, which has the disadvantages of being time-consuming, labor-intensive, costly, and having a small survey scope. How to accurately, efficiently, and non-destructively identify wheat ears is of great practical significance to wheat production and breeding.

[0003] In recent years, deep learning-based methods have been increasingly used in the fields of visual systems, speech detection, and document analysis. Compared with artificial feature extraction, deep learning technology can use multi-layer neural networks to process images, thereby obtaining local information and deep information of the image. In current deep learning, YOLOv5 is widely used because of its superior performance and the advantages of both real-time and accuracy. However, the original YOLOv5 still has problems such as being easily occluded and low accuracy in detecting small targets. Therefore, a fast target detection algorithm with high accuracy in recognizing occluded small targets is needed. Summary of the invention

[0004] The object of the present invention is to provide a wheat ear detection method based on improved YOLOv5, which can improve the recognition accuracy of small targets and enhance the accuracy of inspection frame detection.

[0005] To achieve the above object, the present invention adopts the following technical solution: a wheat ear detection method based on improved YOLOv5, the method comprising the following steps in order:

[0006] (1) obtaining wheat ear images and annotating the wheat ear images to obtain a wheat ear image dataset, and dividing the wheat ear image dataset into a training set and a test set;

[0007] (2) Build the YOLOv5 network model;

[0008] (3) Improve the YOLOv5 network model to obtain an improved YOLOv5 network model;

[0009] (4) Inputting the training set into the improved YOLOv5 network model to train the improved YOLOv5 network model;

[0010] (5) Evaluate and test the improved YOLOv5 network model.

[0011] The step (1) specifically comprises the following steps:

[0012] (1a) A DJI Phantom 4PRO drone equipped with a camera model FC6310 was used to collect 1500 wheat ear images with a resolution of 5472×3648 pixels;

[0013] (1b) Use the data annotation tool Labeling to annotate the wheat ear image, use a rectangular frame to mark the position of the wheat ear in the image, and annotate it in YOLO format to obtain a wheat ear image dataset;

[0014] (1c) The labeled wheat ear image dataset is divided into a training set and a test set in a ratio of 7:3.

[0015] In step (2), the YOLOv5 network model includes:

[0016] Input input end, used for Mosaic data enhancement and adaptive image scaling;

[0017] Backbone basic network, including CON structure, C3 structure and SPP structure;

[0018] Neck network adopts the structure of FPN+PAN, in which the FPN structure is used to transmit strong semantic feature information from top to bottom; the PAN structure adds an upward feature pyramid after the FPN structure to transmit strong positioning information from bottom to top;

[0019] The Prediction output layer includes a Bounding Box loss function and NMS non-maximum suppression, and the Bounding Box loss function adopts the GIOU_Loss loss function.

[0020] The step (3) specifically comprises the following steps:

[0021] (3a) Introduce the CA attention mechanism between the C3 structure and the SPP structure of the Backbone basic network of the YOLOv5 network model;

[0022] (3b) A 160×160 detection scale is added to the YOLOv5 network model, expanding the three-scale detection to a four-scale detection. After the 80×80 feature layer, convolutional layers and upsampling are added, and then the double upsampling feature layer is fused with the 160×160 feature layer to obtain the fourth 160×160 detection scale. The anchor frame setting is automatically generated using the K-Means algorithm provided by YOLOv5.

[0023] (3c) The Bounding Box loss function uses the CIOU_Loss loss function, and its formula is as follows:

[0024]

[0025] Among them, a is the weight coefficient, v represents the distance between the aspect ratio of the detection box and the real box, b and b gt They represent the center points of the prediction boxes of wheat ears and non-wheat ears, respectively. ρ represents the Euclidean distance. C represents the diagonal distance of the minimum bounding rectangle of the target. IoU represents the ratio of the intersection area of ​​two boxes to their union area. The expressions of a and v are:

[0026]

[0027]

[0028] In the formula, ω gt is the width of the real rectangular frame, h gt is the height of the real rectangular box, ω is the width of the detection rectangular box, and h is the height of the detection rectangular box.

[0029] The step (4) specifically comprises the following steps:

[0030] (4a) Setting training parameters, i.e., grid training parameters, including the number of iterations, batch size, learning rate, and momentum. The number of iterations is 200, the batch size is 8, the learning rate is 0.01, and the momentum is 0.937.

[0031] (4b) Set the YOLOv5 network model parameters, i.e., data enhancement parameters, including hsv_h, hsv_s, hsv_v, degrees, flipud, and fliplr, where hsv_h is 0.015, hsv_s is 0.7, hsv_v is 0.4, degrees is 1.0, flipud is 0.01, and fliplr is 0.5;

[0032] The improved YOLOv5 network model is trained using grid training parameters, data enhancement parameters and training sets.

[0033] The step (5) specifically refers to: using model evaluation indicators to evaluate the improved YOLOv5 network model, the model evaluation indicators include accuracy P, recall R and mean average precision mAP, and the formula is as follows:

[0034]

[0035]

[0036]

[0037]

[0038] Among them, AP is the average precision, TP is the number of wheat ears predicted correctly by the model, FP is the number of samples that identify non-wheat ears as wheat ears, FN is the number of samples that identify wheat ears as non-wheat ears, M is the number of categories, R is the recall rate, P(R) is the corresponding accuracy P under different recall rates R, AP i represents the average accuracy of the i-th iteration;

[0039] Finally, the test set is input into the trained improved YOLOv5 network model for testing.

[0040] In step (3a), the CA attention mechanism includes:

[0041] (3a1) Embed coordinate information: Use the pooling layer to encode each channel along the horizontal and vertical coordinates respectively, and obtain a pair of direction-aware feature maps;

[0042] (3a2) Generate coordinate information feature map: First, concatenate the extracted feature information, then use a 1×1 convolution transformation function to convert the information, and then obtain the intermediate feature map, which is decomposed into two separate tensors along the spatial dimension, and then used two convolution transformations to transform them into tensors with the same number of channels. Finally, expand the output results and use them as attention weight allocation values.

[0043] It can be seen from the above technical scheme that the beneficial effects of the present invention are: first, the present invention uses four-scale feature detection and increases the shallow detection scale to improve the recognition accuracy of small targets; second, the present invention introduces the CA attention mechanism to enhance the feature extraction ability of the algorithm; third, the present invention introduces CIOU_Loss as the Bounding Box Regression Loss of the algorithm loss function to improve the accuracy of inspection box detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flow chart of the method of the present invention;

[0045] Figure 2 This is the structural diagram of the YOLOv5 network;

[0046] Figure 3 This is the structural diagram of the improved YOLOv5 network;

[0047] Figure 4 This is the structural diagram of the CA attention mechanism;

[0048] Figure 5Schematic diagram of the changes of various parameters during training. DETAILED DESCRIPTION

[0049] like Figure 1 As shown, a wheat ear detection method based on improved YOLOv5 includes the following steps in order:

[0050] (1) obtaining wheat ear images and annotating the wheat ear images to obtain a wheat ear image dataset, and dividing the wheat ear image dataset into a training set and a test set;

[0051] (2) Build the YOLOv5 network model;

[0052] (3) Improve the YOLOv5 network model to obtain an improved YOLOv5 network model;

[0053] (4) Inputting the training set into the improved YOLOv5 network model to train the improved YOLOv5 network model;

[0054] (5) Evaluate and test the improved YOLOv5 network model.

[0055] The step (1) specifically comprises the following steps:

[0056] (1a) A DJI Phantom 4PRO drone equipped with a camera model FC6310 was used to collect 1500 wheat ear images with a resolution of 5472×3648 pixels;

[0057] (1b) Use the data annotation tool Labeling to annotate the wheat ear image, use a rectangular frame to mark the position of the wheat ear in the image, and annotate it in YOLO format to obtain a wheat ear image dataset;

[0058] (1c) The labeled wheat ear image dataset is divided into a training set and a test set in a ratio of 7:3.

[0059] like Figure 2 As shown, in step (2), the YOLOv5 network model includes:

[0060] Input input end, used for Mosaic data enhancement and adaptive image scaling;

[0061] Backbone basic network, including CON structure, C3 structure and SPP structure;

[0062] Neck network adopts the structure of FPN+PAN, in which the FPN structure is used to transmit strong semantic feature information from top to bottom; the PAN structure adds an upward feature pyramid after the FPN structure to transmit strong positioning information from bottom to top;

[0063] The Prediction output layer includes a Bounding Box loss function and NMS non-maximum suppression, and the Bounding Box loss function adopts the GIOU_Loss loss function.

[0064] Although the YOLOv5 network model has good performance in detection accuracy and speed, it still has big problems in detecting small targets such as occluders.

[0065] In view of the above problems, the present invention improves the YOLOv5 network model so that its accuracy, recall rate and average precision are improved, and wheat ears can be better detected.

[0066] The step (3) specifically comprises the following steps:

[0067] (3a) The CA attention mechanism is introduced between the C3 structure and the SPP structure of the Backbone basic network of the YOLOv5 network model, such as Figure 4 shown.

[0068] In step (3a), the CA attention mechanism includes:

[0069] (3a1) Embed coordinate information: Use the pooling layer to encode each channel along the horizontal and vertical coordinates respectively, and obtain a pair of direction-aware feature maps;

[0070] (3a2) Generate coordinate information feature map: First, concatenate the extracted feature information, then use a 1×1 convolution transformation function to convert the information, and then obtain the intermediate feature map, which is decomposed into two separate tensors along the spatial dimension, and then used two convolution transformations to transform them into tensors with the same number of channels. Finally, expand the output results and use them as attention weight allocation values.

[0071] (3b) Add a 160×160 detection scale to the YOLOv5 network model, expanding the three-scale detection to the four-scale detection, as shown in Figure 3 As shown in the figure, after the 80×80 feature layer, convolutional layers and upsampling are continued to be added, and then the double upsampling feature layer is fused with the 160×160 feature layer to obtain the fourth detection scale of 160×160. The anchor frame setting is automatically generated by the K-Means algorithm of YOLOv5.

[0072] (3c) The Bounding Box loss function uses the CIOU_Loss loss function, and its formula is as follows:

[0073]

[0074] Among them, a is the weight coefficient, v represents the distance between the aspect ratio of the detection box and the real box, b and b gt They represent the center points of the prediction boxes of wheat ears and non-wheat ears, respectively. ρ represents the Euclidean distance. C represents the diagonal distance of the minimum bounding rectangle of the target. IoU represents the ratio of the intersection area of ​​two boxes to their union area. The expressions of a and v are:

[0075]

[0076]

[0077] In the formula, ω gt is the width of the real rectangular frame, h gt is the height of the real rectangular box, ω is the width of the detection rectangular box, and h is the height of the detection rectangular box.

[0078] The step (4) specifically comprises the following steps:

[0079] (4a) Setting training parameters, i.e., grid training parameters, including the number of iterations, batch size, learning rate, and momentum. The number of iterations is 200, the batch size is 8, the learning rate is 0.01, and the momentum is 0.937.

[0080] (4b) Set the YOLOv5 network model parameters, i.e., data enhancement parameters, including hsv_h, hsv_s, hsv_v, degrees, flipud, and fliplr, where hsv_h is 0.015, hsv_s is 0.7, hsv_v is 0.4, degrees is 1.0, flipud is 0.01, and fliplr is 0.5;

[0081] The improved YOLOv5 network model is trained using grid training parameters, data enhancement parameters and training sets.

[0082] The step (5) specifically refers to: using model evaluation indicators to evaluate the improved YOLOv5 network model, the model evaluation indicators include accuracy P, recall R and mean average precision mAP, and the formula is as follows:

[0083]

[0084]

[0085]

[0086]

[0087] Among them, AP is the average precision, TP is the number of wheat ears predicted correctly by the model, FP is the number of samples that identify non-wheat ears as wheat ears, FN is the number of samples that identify wheat ears as non-wheat ears, M is the number of categories, R is the recall rate, P(R) is the corresponding accuracy P under different recall rates R, AP i represents the average accuracy of the i-th iteration;

[0088] Finally, the test set is input into the trained improved YOLOv5 network model for testing.

[0089] The changes of various evaluation indicators of the improved YOLOv5 network model during the training process are as follows: Figure 5 As shown in the figure, after the improved YOLOv5 network model is trained, it is tested with the test set, and the accuracy P reaches 91.1%, the recall R reaches 84.9%, the mAP(0.5) reaches 91.1%, and the mAP(0.5:0.9) reaches 48%. The accuracy P, the recall R, the mAP(0.5), and the mAP(0.5:0.9) are increased by 0.2%, 1.3%, 1.9%, and 1.8%, respectively, proving the feasibility of the present invention.

[0090] Table 2 Algorithm performance comparison

[0091] Precision means P Recall is R mAP(0.5) mAP(0.5:0.9) yolov5 0.909 0.836 0.892 0.462 Algorithm of the present invention 0.911 0.849 0.911 0.48

[0092] The results show that the algorithm of this invention has better effect than the original algorithm.

[0093] In summary, the present invention uses four-scale feature detection and increases the shallow detection scale to improve the recognition accuracy of small targets; the present invention introduces the CA attention mechanism to enhance the feature extraction capability of the algorithm; the present invention introduces CIOU_Loss as the Bounding Box Regression Loss of the algorithm loss function to improve the accuracy of inspection box detection.

Claims

1. A wheat ear detection method based on improved YOLOv5, characterized in that: The method comprises the following steps in order: (1) obtaining wheat ear images and annotating the wheat ear images to obtain a wheat ear image dataset, and dividing the wheat ear image dataset into a training set and a test set; (2) Build the YOLOv5 network model; (3) Improve the YOLOv5 network model to obtain an improved YOLOv5 network model; (4) Inputting the training set into the improved YOLOv5 network model to train the improved YOLOv5 network model; (5) Evaluate and test the improved YOLOv5 network model; In step (2), the YOLOv5 network model includes: Input input end, used for Mosaic data enhancement and adaptive image scaling; Backbone basic network, including CON structure, C3 structure and SPP structure; Neck network adopts the structure of FPN+PAN, in which the FPN structure is used to transmit strong semantic feature information from top to bottom; the PAN structure adds an upward feature pyramid after the FPN structure to transmit strong positioning information from bottom to top; Prediction output layer, including Bounding Box loss function and NMS non-maximum suppression, the Bounding Box loss function adopts GIOU_Loss loss function; The step (3) specifically comprises the following steps: (3a) Introduce the CA attention mechanism between the C3 structure and the SPP structure of the Backbone basic network of the YOLOv5 network model; (3b) A 160×160 detection scale is added to the YOLOv5 network model, expanding the three-scale detection to a four-scale detection. After the 80×80 feature layer, convolutional layers and upsampling are added, and then the double upsampling feature layer is fused with the 160×160 feature layer to obtain the fourth 160×160 detection scale. The anchor frame setting is automatically generated using the K-Means algorithm provided by YOLOv5. (3c) The Bounding Box loss function uses the CIOU_Loss loss function, and its formula is as follows: Among them, a is the weight coefficient, v represents the distance between the aspect ratio of the detection box and the real box, b and b gt They represent the center points of the prediction boxes of wheat ears and non-wheat ears, respectively. ρ represents the Euclidean distance. C represents the diagonal distance of the minimum bounding rectangle of the target. IoU represents the ratio of the intersection area of ​​two boxes to their union area. The expressions of a and v are: In the formula, ω gt is the width of the real rectangular frame, h gt is the height of the real rectangular box, ω is the width of the detection rectangular box, and h is the height of the detection rectangular box; The step (5) specifically refers to: using model evaluation indicators to evaluate the improved YOLOv5 network model, the model evaluation indicators include accuracy P, recall R and mean average precision mAP, and the formula is as follows: Among them, is the average precision, TP represents the number of wheat ears predicted correctly by the model, FP represents the number of samples that identify non-wheat ears as wheat ears, FN represents the number of samples that identify wheat ears as non-wheat ears, M represents the number of categories, R represents the recall rate P(R) represents the corresponding accuracy P under different recall rates R, AP i represents the average accuracy of the i-th iteration; Finally, the test set is input into the trained improved YOLOv5 network model for testing.

2. The wheat ear detection method based on improved YOLOv5 according to claim 1, characterized in that: The step (1) specifically comprises the following steps: (1a) A DJI Phantom 4PRO drone equipped with a camera model FC6310 was used to collect 1500 wheat ear images with a resolution of 5472×3648 pixels; (1b) Use the data annotation tool Labeling to annotate the wheat ear image, use a rectangular frame to mark the position of the wheat ear in the image, and annotate it in YOLO format to obtain a wheat ear image dataset; (1c) The labeled wheat ear image dataset is divided into a training set and a test set in a ratio of 7:

3.

3. The wheat ear detection method based on improved YOLOv5 according to claim 1, characterized in that: The step (4) specifically comprises the following steps: (4a) Setting training parameters, i.e., grid training parameters, including the number of iterations, batch size, learning rate, and momentum. The number of iterations is 200, the batch size is 8, the learning rate is 0.01, and the momentum is 0.

937. (4b) Set the YOLOv5 network model parameters, i.e., data enhancement parameters, including hsv_h, hsv_s, hsv_v, degrees, flipud, and fliplr, where hsv_h is 0.015, hsv_s is 0.7, hsv_v is 0.4, degrees is 1.0, flipud is 0.01, and fliplr is 0.5; The improved YOLOv5 network model is trained using grid training parameters, data enhancement parameters and training sets.

4. The wheat ear detection method based on improved YOLOv5 according to claim 1, characterized in that: In step (3a), the CA attention mechanism includes: (3a1) Embed coordinate information: Use the pooling layer to encode each channel along the horizontal and vertical coordinates respectively, and obtain a pair of direction-aware feature maps; (3a2) Generate coordinate information feature map: First, concatenate the extracted feature information, then use a 1×1 convolution transformation function to convert the information, and then obtain the intermediate feature map, which is decomposed into two separate tensors along the spatial dimension, and then used two convolution transformations to transform them into tensors with the same number of channels. Finally, expand the output results and use them as attention weight allocation values.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image wheat ear recognition method based on deep learning

    CN113435282A

  • YOLOv4 target detection algorithm for improving loss function

    CN114463718A