Banana target detection, maturity classification and counting method based on improved YOLOv5
By introducing the EMA attention mechanism, adding new feature layer and maturity classification tasks in the YOLOv5 model, and using the MPDIoU loss function, the problems of banana object detection, maturity classification and counting in complex environments are solved, and high-precision and efficient detection and classification capabilities are achieved.
Patent Information
- Application Number
- CN202510045118.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to efficiently and accurately realize banana target detection, maturity classification and counting in complex environments, especially in small target detection, dense target processing, occlusion and light changes.
The EMA attention mechanism is introduced to enhance feature extraction capabilities, a new feature layer that is adapted to small targets is added, a maturity classification task is added, and the maturity classification accuracy is optimized through cross-entropy loss. The MPDIoU loss function is used to combine object detection and maturity classification tasks to form a multi-task joint loss function.
It significantly improves the accuracy, maturity classification and counting capabilities of banana target detection, and is suitable for agricultural production and automation management systems in complex environments.
Smart Images

Figure CN119942078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and artificial intelligence, and in particular to a banana target detection, maturity classification and counting method based on improved YOLOv5. Background Art
[0002] Banana is one of the most widely grown and consumed fruits in the world, and its yield assessment and logistics management are of great significance in modern agriculture. However, due to the complex planting environment of bananas and the different sizes and shapes of fruits, it is difficult for existing technologies to efficiently and accurately detect and count banana targets.
[0003] Early target detection technologies relied on manually designed feature extraction methods, such as the Viola-Jones detector and HOG features. These traditional methods mainly complete target detection tasks by combining manually defined features with classifiers (such as support vector machines (SVMs)). Although these methods have certain interpretability, they have slow detection speed, low accuracy, and insufficient generalization ability, especially in complex backgrounds and diverse target scenes.
[0004] The development of deep learning has completely changed the landscape of object detection. In 2012, in the ImageNet Large Scale Visual Recognition Challenge (ILSVRC), AlexNet first introduced convolutional neural networks (CNNs) for image classification and achieved breakthrough results. This milestone research inspired the development of subsequent object detection methods, including the design of two-stage detectors and single-stage detectors. Object detection technology has gradually migrated from traditional methods to deep learning models and has been widely used in fields such as autonomous driving, medical diagnosis, and security monitoring.
[0005] Object detection is a natural extension of image classification, which aims to detect the categories and locations of all objects of interest in an image. Modern object detection technology regards the object detection problem as a supervised learning problem, and achieves efficient object recognition and localization by training on large-scale annotated images. Detectors usually handle multi-object detection tasks by predicting the bounding box and category label of the object, combined with non-maximum suppression.
[0006] Representative algorithms of two-stage detectors include Faster R-CNN and Mask R-CNN. These methods first generate a set of candidate regions, and then classify and regress these regions. Two-stage detectors have significant advantages in accuracy, but because they involve multiple steps of processing, they are slow and difficult to meet real-time requirements.
[0007] The YOLO family of algorithms represents the core progress of single-stage detectors. The YOLO model redefines the target detection task as a regression problem. Instead of relying on candidate region generation, it directly predicts the target category and bounding box parameters on the entire image. Taking YOLOv1 as an example, the algorithm divides the input image into S×S grids, and each grid predicts the bounding box, category, and confidence of the target. Subsequent versions such as YOLOv3 and YOLOv5 have further optimized the feature extraction network, Anchor mechanism, multi-scale detection, etc., achieving an excellent balance between speed and accuracy. However, early versions of YOLO performed poorly in small target detection and occluded scenes, and had limited processing capabilities for densely distributed targets.
[0008] Although object detection technology has made great progress in the past decade, it still faces many challenges in applications in complex environments (such as agricultural scenarios). For example: 1. The detection accuracy of small targets (such as a single banana) and densely distributed targets is not high; 2. Occlusion and illumination changes pose serious challenges to the robustness of detectors; 3. Existing models are difficult to balance detection speed and accuracy, especially when deployed on embedded devices.
[0009] Li Jiajun proposed an improved Faster R-CNN to improve the accuracy of strawberry fruit detection and counting. By optimizing the RPN structure and introducing the ResNet50 backbone network, the target detection performance in complex backgrounds was significantly improved. However, his method has shortcomings in dealing with banana detection tasks: first, Faster R-CNN is a two-stage detector with a slow detection speed, which is not suitable for scenarios with high real-time requirements; second, although it has optimized the recognition of small targets, it still has limitations in processing dense targets such as banana stacks.
[0010] With the rapid development of deep learning technology, target detection models represented by the YOLO series have been widely promoted in various application scenarios, among which YOLOv3 and YOLOv5 have performed particularly well in actual engineering applications. However, there are still many challenges in applying target detection models to banana maturity detection and counting tasks. In agricultural scenarios, banana targets usually have complex background environments, different lighting conditions, and occlusion problems, which put higher requirements on the detection accuracy of the model. In addition, banana maturity detection involves multi-dimensional feature analysis such as color and morphology, and existing target detection models are difficult to achieve high-precision comprehensive performance in these dimensions. At present, most banana counting and maturity judgment still rely on manual labor, which is not only a huge workload, but also prone to errors due to subjective judgment.
[0011] Therefore, technicians in this field are committed to developing a banana target detection, maturity classification and counting method based on improved YOLOv5. Summary of the invention
[0012] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to effectively improve the detection accuracy, maturity classification and counting accuracy of banana targets in a complex environment.
[0013] In view of the problems of low banana target detection accuracy, difficulty in maturity classification and inaccurate counting in the prior art, this application introduces the EMA (Efficient Multi-head Attention) attention mechanism to enhance the ability to extract banana target features; a feature layer adapted to small targets is added to the YOLOv5 structure, a maturity classification task is added to the detection head, and the maturity classification accuracy is optimized through cross entropy loss. The MPDIoU loss function is adopted, and the target detection and maturity classification tasks are combined to form a multi-task joint loss function, which improves the stability of bounding box regression and counts the number of banana targets of different maturity.
[0014] In the present invention, YOLOv5 refers to the YOLOv5 model, and improved YOLOv5 refers to the improved YOLOv5 model.
[0015] In one embodiment of the present invention, a banana target detection, maturity classification and counting method based on improved YOLOv5 is provided, comprising the following steps: S100, banana image acquisition, using an image sensor to collect banana images in different scenes, construct a banana target detection data set, and perform annotation; S200, image data processing and enhancement, performing image data processing and enhancement on the labeled banana object detection data set, and dividing the training set, validation set, and test set into specified proportions; S300, YOLOv5 improvement, add EMA module to the backbone network of YOLOv5 to enhance feature extraction capability, add detection head and feature layer to the feature fusion network of YOLOv5 to improve the detection accuracy of small targets and stacked targets, add classifier to the detection head of YOLOv5 to classify the maturity of bananas, replace the default loss function to improve the regression stability and positioning accuracy of the bounding box, combine the target detection and maturity classification tasks to synthesize the joint loss function, and get the improved YOLOv5; S400, improve YOLOv5 training, load training set, adjust training parameters, and train improved YOLOv5; S500, banana target detection, maturity classification and counting, uses the trained improved YOLOv5 to perform target detection, maturity classification and counting of bananas; S600, the results are output in real time, and the banana target frame and target quantity, maturity classification and counting results are output and displayed in real time.
[0016] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the image sensor uses a high-resolution industrial camera.
[0017] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, different scenes include different lighting conditions, occlusion situations and complex backgrounds.
[0018] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, different lighting conditions include strong light and shadow.
[0019] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the occlusion conditions include leaf occlusion and stacking.
[0020] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the complex background includes the ground, boxes and trellises.
[0021] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, labeling is performed using the LabelImg tool to generate a label file.
[0022] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the format of the label file is COCO format.
[0023] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the label file includes the coordinates, category, maturity category and location information of the target frame.
[0024] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, image data processing and enhancement include random rotation, mirroring, brightness adjustment, scaling and adding noise.
[0025] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the noise includes Gaussian noise.
[0026] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the angle range of random rotation is 0° to 180°.
[0027] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, the specified ratio is training set: validation set: test set = 7:1:2.
[0028] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, step S300 includes: S310, add EMA module, add EMA module to each C3 module in the YOLOv5 backbone network to enhance the global modeling ability and attention effect of the banana feature area; S320, added detection heads and feature layers, added detection heads and feature layers to the feature fusion network of YOLOv5 to additionally adapt to small targets, and improve the detection accuracy of stacked bananas and small targets; S330, adding a classifier, adding a classifier to the detection head of YOLOv5 to classify the ripeness of bananas; S340, optimize maturity classification, use cross-entropy loss function (Cross-Entropy Loss) to optimize maturity classification; S350, replace the loss function, replace the default CIOU loss function of YOLOv5 with the MPDIoU loss function to improve the regression stability and positioning accuracy of the bounding box; MPDIoU loss function The formula is as follows: in, i is a positive integer, indicating the number of the key point. n is a positive integer, indicating the total number of key points in the target box and the predicted box. IoU is the intersection-over-union ratio of the target box and the prediction box. The prediction box is the bounding box output by YOLOv5, which indicates the prediction result of YOLOv5 on the target position. The target box is the bounding box manually provided during data annotation, which indicates the actual position and range of the target in the image. The Euclidean distance between the center point of the prediction box and the target box: ( , )and( , ) are the center point coordinates of the prediction box and the target box respectively, is the diagonal length of the image in the banana object detection dataset after image data processing and enhancement; It is the sum of the Euclidean distances between the key points (including vertices and midpoints) of the prediction box and the target box; is an aspect ratio constraint used to optimize the shape of the box: , and , The width and height of the prediction box and target box respectively; α , β , γ are weight hyperparameters, which respectively control the influence of center point distance, key point distance and aspect ratio constraint. Preferably, their values are α =0.5, β =0.3, γ =0.2; S360, synthetic joint loss function, combining target detection and maturity classification tasks, and synthesizing the joint loss function of multi-task learning , the formula is as follows: in, is the target detection loss function, is the classification loss function for maturity classification, and It is a loss weight hyperparameter that controls the trade-off between detection and classification tasks. The target detection loss function uses the MPDIoU loss function: ; The classification loss function for maturity classification uses the cross entropy loss function: Where N is the number of samples. k is the number of maturity categories, k =3, is the true classification label of the target instance, indicating i objects (banana instances) and classification categories j , the value is 0 or 1, 0 means the target instance does not belong to the classification category, 1 means the target instance belongs to the classification category, is the predicted maturity probability.
[0029] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the size of the EMA module input feature map in step S310 is C×H×W , C Indicates the number of input feature map channels, H represents the input feature map height,W Indicates the input feature map width.
[0030] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, step S310 includes: S311, feature map grouping, the input feature map is grouped in the channel dimension, the number of channels in each group is C / N ,in C is the number of input feature map channels, N is the number of groups, the input feature map comes from the backbone network and the feature fusion network; S312, query vector generation, generate a query vector for each set of input feature maps through a shared linear projection matrix Q , key vector K , value vector V , the projection matrix is , the shared linear projection matrix refers to the one used in the EMA module to generate the query vector Q , key vector K , value vector V The projection matrix , the linear projection matrix The input feature maps of all groups are shared; S313, attention weight matrix calculation, use dot product to calculate the attention weight matrix of each group A , the formula is as follows: in is the dimension of the key vector, T represents transpose; S314, weighted feature map generation, attention weight matrix A Acting on a vector of values V , generate weighted feature map , the formula is as follows: ; S315: Weighted feature map concatenation: concatenate the weighted feature maps of all groups and restore them to the number of input feature map channels. C , and reconstruct the spatial dimension of the input feature map H×W ; S316: Output feature map generation: add the generated weighted feature map to the input feature map element by element to generate the output feature map ,in is a learnable parameter used to adjust the fusion strength of the attention feature. The size of the output feature map is the same as the input feature map. C×H×W .
[0031] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the feature map size of the newly added feature layer in step S320 is 80×80 pixels, corresponding to an input image size of 640×640.
[0032] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, a new feature layer is added to introduce the intermediate output P3 feature map of the YOLOv5 backbone network into the feature fusion stage to output a new feature map as the input of the new detection head.
[0033] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the intermediate output P3 feature map of the backbone network includes 128 output channels with a size of 80×80 pixels.
[0034] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the small target in step S320 refers to a banana target with a small area and a low pixel ratio relative to the entire image size, including bananas at a distance, partially blocked bananas, and closely distributed targets in stacked scenes.
[0035] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the newly added detection head in step S320 receives the feature map output by the newly added feature layer, the size of which is 80×80×256. The newly added detection head completes the banana maturity classification and bounding box regression prediction, and the bounding box regression prediction is used to determine the position, size and confidence of the target.
[0036] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the newly added detection head in step S320 performs convolution on the new feature map output by the newly added feature layer to further extract high-level features.
[0037] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the convolution layer parameters used in the convolution are set to: the convolution kernel size is 3×3, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 256.
[0038] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the size of the feature map after convolution is 80×80×256, which includes high-level features that have been further processed.
[0039] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, a new detection head is added to extract classification information from the convolutional feature map through two fully connected layers (defined as classification fully connected layers) to predict the banana maturity classification.
[0040] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the input size of the first classification fully connected layer is 256, and the output size is 128. The input size of the second classification fully connected layer is 128, and the output size is the number of banana maturity classifications. Finally, the classification probability distribution is generated by the Softmax function, and the output size is 80×80×the number of banana maturity classifications. The number of banana maturity classifications is 3, namely, unripe, ripe, and overripe.
[0041] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, a new detection head is added to extract regression information from the convolutional feature map through two fully connected layers (defined as regression fully connected layers) to predict the bounding box position and confidence of the banana.
[0042] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the input size of the first regression fully connected layer is 256, and the output size is 128. The input size of the second regression fully connected layer is 128, and the output size is 5×the number of Anchors. The output includes the center point coordinates (x, y), width and height (w, h), and confidence of the bounding box. The output size is 80×80×5×the number of Anchors, and the Anchor is the priori box.
[0043] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the number of output channels of the newly added detection head is = ( k + 5) x number of Anchors, where k is the number of banana maturity classifications, 5 represents the coordinates and confidence information of the bounding box, and the number of anchors refers to the number of anchors used for target prediction in each feature map cell, ranging from 1 to 6, and the default is 3.
[0044] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the number of Anchors is adjusted according to the accuracy requirement and the complexity of the environment. In the banana stacking scene, for high accuracy requirements, the number of Anchors is increased to 4-6 to adapt to targets of different sizes and aspect ratios, thereby improving the detection accuracy. In an environment where computing resources are limited, banana targets are relatively scattered and the targets are relatively large, the number of Anchors is reduced to 1-2 to improve the detection speed.
[0045] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, step S320 includes: S321, extract feature maps, add a new feature layer to extract P3 feature maps and P2 feature maps from the backbone network, perform a convolution on the P3 feature map, and then upsample it, and align it with the P2 feature map through the nearest neighbor interpolation method; S322, forming a new feature map, splicing the upsampled P3 feature map and the P2 feature map along the channel dimension to form a new feature map; S323, extracting features, performing feature extraction on the new feature map through two groups of convolutional layers in sequence; S324, extracting high-level features, using the new feature map after feature extraction as the input of the newly added detection head, further extracting high-level features, and performing banana maturity classification and regression prediction.
[0046] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the size of the P3 feature map in step S321 is 80×80×128, and the size of the P2 feature map is 160×160×64.
[0047] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the P3 feature map after one convolution in step S321 adjusts the number of channels to 256, and the convolution parameters are the convolution kernel size of 3×3, the stride of 1, and the padding of 1.
[0048] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the number of channels of the new feature map in step S322 is 320 (256 + 64), and the spatial resolution is 80×80.
[0049] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the parameters of the first group of convolutional layers in the two groups of convolutional layers in step S323 are convolution kernel size 1×1, stride 1, padding 0, reducing the number of channels from 320 to 128, reducing channel redundancy and improving computational efficiency; the parameters of the second group of convolutional layers are convolution kernel size 3×3, stride 1, padding 1, expanding the number of channels from 128 to 256, enhancing the characterization capability of features.
[0050] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the size of the feature map after extracting high-level features in step S324 is 80×80×256, which not only retains the spatial resolution required for small target and dense target detection, but also combines the semantic information of multi-scale features.
[0051] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the final output feature map size of the newly added feature layer = 80 x 80 x ( k + 5) x number of Anchors.
[0052] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the maturity classification of bananas includes unripe, ripe and overripe, that is, the number of maturity categories is 3.
[0053] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the classifier structure includes two fully connected layers, the output size is the number of maturity categories, and the classification probability distribution is generated by Softmax.
[0054] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, step S400 includes: S410, loading a training set, loading the training set into a training framework, and setting initial training parameters; S420, load pre-trained weights and use multi-task joint loss function , the initial weights λ1 and λ2 are set to 1:1 to balance the target detection loss and maturity classification loss; S430, improve YOLOv5 training, perform forward propagation and back propagation on the training set batch by batch, and calculate the target detection loss of the training set through forward propagation in each training round (epoch) and maturity classification loss , and combined with the multi-task joint loss function Perform back-propagation optimization; after each round of training, calculate and save the mean average precision (mAP), detection accuracy, and recall rate of the validation set; when the training end conditions are met, complete the training of the improved YOLOv5; S440, calculate the best training weight, compare the mean average precision (mAP) of the validation set, and the weight of the improved YOLOv5 when the mean average precision (mAP) is the largest is the best training weight.
[0055] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the training set includes banana images and their corresponding annotation files, and the training set ratio is set to 70%.
[0056] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the training parameters include input image size, batch size, learning rate, maximum number of training rounds (epochs), optimizer, weight decay parameter, and data enhancement strategy.
[0057] Further, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the input image size defaults to 640×640 pixels, which can be adjusted to 320×320-1280×1280 pixels according to the accuracy requirement; the batch size defaults to 16, which can be adjusted to 32 or 64 according to the GPU video memory; the initial value of the learning rate is set to 0.01, and the cosine annealing scheduling strategy is gradually reduced to improve stability; the maximum recommended range of training rounds is 100-500 rounds, which is adjusted according to the convergence of the model; the optimizer selects AdamW (an optimization algorithm with weight decay regularization added on the basis of the Adam optimizer); the weight decay parameter is set to 0.0001 to reduce overfitting; the data enhancement strategy includes random scaling, cropping, and flipping to improve the generalization ability of the model.
[0058] Preferably, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the maximum number of training rounds is set to 300 rounds.
[0059] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the pre-trained weights are downloaded from the official GitHub repository of YOLOv5 or other related resource websites.
[0060] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the training end condition includes reaching the maximum number of training rounds or satisfying the early stopping strategy or the loss function stably converges.
[0061] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the early stopping strategy is that the mean average precision (mAP) of the validation set or the detection accuracy has not improved in consecutive specified rounds, and the validation loss has no obvious decreasing trend.
[0062] Preferably, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the specified number of rounds is 10 rounds.
[0063] Furthermore, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the loss function converges stably when the training loss The decrease is below the preset threshold.
[0064] Preferably, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in the above embodiment, the preset threshold value = 0.01.
[0065] Optionally, in the banana target detection, maturity classification and counting method based on improved YOLOv5 in any of the above embodiments, step S500 includes: S510, optimal training weight loading, loading the optimal training weight, and using the improved YOLOv5 after training; S520, target detection and maturity classification, inputting the image or / and video stream of the banana to be detected into the trained improved YOLOv5 to perform target detection and maturity classification, and obtaining the bounding box of the banana to be detected, the detection confidence and the probability distribution of maturity classification; S530, bounding box optimization, uses the non-maximum suppression (NMS) algorithm to optimize the bounding boxes generated by the improved YOLOv5 prediction, removes redundant boxes and retains high-confidence bounding boxes; S540 , object counting, counting the optimized bounding boxes, counting the number of optimized bounding boxes according to maturity classification, and obtaining the count of each maturity classification.
[0066] This application introduces the EMA (Efficient Multi-head Attention) attention mechanism to enhance the ability to extract banana target features and improve target detection accuracy; a feature layer adapted to small targets is added to the YOLOv5 structure to improve the detection accuracy of stacked bananas and small targets; this application adds a maturity classification task to the detection head, divides banana targets into three categories: unripe, ripe, and overripe, and optimizes maturity classification accuracy through cross entropy loss; the MPDIoU loss function is used to combine target detection and maturity classification tasks to form a multi-task joint loss function, which improves the stability of bounding box regression. During the reasoning process, redundant frames are eliminated through non-maximum suppression NMS, the number of banana targets of different maturity is counted, and the detection results and counting information are displayed in real time. This application effectively improves the accuracy of banana target detection, maturity classification and counting capabilities, and is suitable for agricultural production and automated management systems in complex environments.
[0067] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is a flow chart illustrating a method for banana target detection, maturity classification, and counting based on improved YOLOv5 according to an exemplary embodiment; Figure 2 is a flow chart illustrating a YOLOv5 improvement according to an exemplary embodiment. DETAILED DESCRIPTION
[0069] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0070] In the drawings, components with the same structure are indicated by the same numerical reference numerals, and components with similar structures or functions are indicated by similar numerical reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of the components is schematically exaggerated in some places in the drawings.
[0071] The inventors provide a banana target detection, maturity classification and counting method based on improved YOLOv5, such as Figure 1 As shown, the following steps are included: S100, banana image acquisition, use image sensors to collect banana images in different scenes, build a banana target detection dataset, and annotate them. The image sensor uses a high-resolution industrial camera. Different scenes include different lighting conditions, occlusion conditions, and complex backgrounds. Different lighting conditions include strong light and shadows. Occlusion conditions include leaf occlusion and stacking. Complex backgrounds include the ground, boxes, and trellises. The LabelImg tool is used for annotation to generate a label file. The format of the label file is COCO format. The label file contains the coordinates, category, maturity category, and location information of the target box.
[0072] S200, image data processing and enhancement, image data processing and enhancement are performed on the labeled banana target detection data set, including random rotation, mirroring, brightness adjustment, scaling and noise addition, the noise includes Gaussian noise, the angle range of random rotation is 0° to 180°, and the training set, validation set and test set are divided into a specified ratio, training set: validation set: test set = 7:1:2.
[0073] S300, YOLOv5 improvement, add EMA module to the backbone network of YOLOv5 to enhance feature extraction capability, add detection head and feature layer to the feature fusion network of YOLOv5 to improve the detection accuracy of small targets and stacked targets, add classifier to the detection head of YOLOv5 to classify the maturity of bananas, replace the default loss function to improve the regression stability and positioning accuracy of the bounding box, combine the target detection and maturity classification tasks to synthesize the joint loss function, and obtain the improved YOLOv5; Figure 2 Specifically shown include: S310, add EMA module, add EMA module to each C3 module in the YOLOv5 backbone network to enhance the global modeling ability and attention effect of the banana feature area. The input feature map size of the EMA module is C×H×W , C Indicates the number of input feature map channels, H represents the input feature map height, W Indicates the width of the input feature map; specifically includes: S311, feature map grouping, the input feature map is grouped in the channel dimension, the number of channels in each group is C / N ,in C is the number of input feature map channels, N is the number of groups, the input feature map comes from the backbone network and the feature fusion network; S312, query vector generation, generate a query vector for each set of input feature maps through a shared linear projection matrix Q , key vector K , value vector V , the projection matrix is , the shared linear projection matrix refers to the one used in the EMA module to generate the query vector Q , key vector K , value vector V The projection matrix , the linear projection matrix The input feature maps of all groups are shared; S313, attention weight matrix calculation, use dot product to calculate the attention weight matrix of each group A , the formula is as follows: in is the dimension of the key vector, which should mean rotation; S314, weighted feature map generation, attention weight matrix A Acting on a vector of values V , generate weighted feature map , the formula is as follows: ; S315: Weighted feature map concatenation: concatenate the weighted feature maps of all groups and restore them to the number of input feature map channels. C , and reconstruct the spatial dimension of the input feature map H×W ; S316: Output feature map generation: add the generated weighted feature map to the input feature map element by element to generate the output feature map ,in is a learnable parameter used to adjust the fusion strength of the attention feature. The size of the output feature map is the same as the input feature map. C×H×W .
[0074] S320, added detection head and feature layer, added detection head and feature layer in the feature fusion network of YOLOv5 to additionally adapt to small targets, improve the detection accuracy of stacked bananas and small targets; the feature map size of the new feature layer is 80×80 pixels, corresponding to the input image size of 640×640, the new feature layer introduces the intermediate output P3 feature map of the YOLOv5 backbone network into the feature fusion stage to output a new feature map as the input of the new detection head, the intermediate output P3 feature map of the backbone network includes 128 output channels, the size is 80×80 pixels; small targets refer to banana targets with a small area and a low pixel ratio relative to the size of the entire image, including bananas at a distance, partially occluded bananas, and closely distributed targets in stacked scenes; the new detection head receives the feature map output by the new feature layer, the size of which is 80×80×256. The new detection head performs convolution on the new feature map output by the new feature layer. The convolution layer parameters used are set to: convolution kernel size is 3×3, stride is 1, padding is 1, input channel number is 256, output channel number is 256, and high-level features are further extracted. The size of the feature map after convolution is 80×80×256, which contains high-level features that have been further processed; the new detection head extracts classification information from the feature map after convolution through two layers of fully connected layers (defined as classification fully connected layers) to predict the banana maturity classification. The input size of the first classification fully connected layer is 256, and the output size is 128. The input size of the second classification fully connected layer is 128, and the output size is the number of banana maturity classifications. Finally, the classification probability distribution is generated through the Softmax function, and the output size is 80×80×number of banana maturity classifications. There are 3 banana maturity classifications, namely unripe, ripe, and overripe. The new detection head extracts regression information from the convolutional feature map through two fully connected layers (defined as regression fully connected layers) to predict the bounding box position and confidence of the banana. The input size of the first regression fully connected layer is 256 and the output size is 128. The input size of the second regression fully connected layer is 128 and the output size is 5×number of Anchors. The output includes the center point coordinates (x, y), width and height (w, h), and confidence of the bounding box. The output size is 80×80×5×number of Anchors. Anchor is the prior box. The number of output channels of the new detection head = ( k + 5) x number of Anchors, where kis the number of banana maturity classifications, 5 represents the coordinates and confidence information of the bounding box, and the number of anchors refers to the number of anchors used for target prediction in each feature map cell, ranging from 1 to 6, with a default of 3. The number of anchors is adjusted according to the accuracy requirements and environmental complexity. In the banana stacking scene, for high accuracy requirements, the number of anchors is increased to 4 to 6 to adapt to targets of different sizes and aspect ratios to improve detection accuracy. In an environment where computing resources are limited, banana targets are relatively scattered and the targets are large, the number of anchors is reduced to 1 to 2 to improve detection speed. Specifically, it includes: S321, extract feature maps, add a new feature layer to extract P3 feature maps and P2 feature maps from the backbone network, the size of P3 feature map is 80×80×128, the size of P2 feature map is 160×160×64, perform a convolution on the P3 feature map, and then upsample it, align it with the P2 feature map through the nearest neighbor interpolation method, adjust the number of channels of the P3 feature map after one convolution to 256, and the convolution parameters are the convolution kernel size of 3×3, the stride of 1, and the padding of 1; S322, forming a new feature map, concatenating the upsampled P3 feature map and the P2 feature map along the channel dimension to form a new feature map, the number of channels of the new feature map is 320 (256 + 64), and the spatial resolution is 80×80; S323, extract features, extract features from the new feature map through two groups of convolutional layers in sequence, the parameters of the first group of convolutional layers in the two groups of convolutional layers are convolution kernel size 1×1, stride 1, padding 0, reducing the number of channels from 320 to 128, reducing channel redundancy and improving computational efficiency; the parameters of the second group of convolutional layers are convolution kernel size 3×3, stride 1, padding 1, expanding the number of channels from 128 to 256, enhancing the representation ability of features; S324, extract high-level features, use the new feature map after feature extraction as the input of the newly added detection head, further extract high-level features, perform banana maturity classification and regression prediction, the feature map size after extracting high-level features is 80×80×256, which not only retains the spatial resolution required for small target and dense target detection, but also combines the semantic information of multi-scale features. The final output feature map size of the newly added feature layer = 80 x 80 x ( k + 5) x Anchor number. The maturity classification of bananas includes unripe, ripe and overripe, that is, the number of maturity categories is 3.
[0075] S330, adding a classifier. A classifier is added to the detection head of YOLOv5. The classifier structure includes two fully connected layers. The output size is the number of maturity categories. The classification probability distribution is generated through Softmax to classify the maturity of bananas. S340, optimize maturity classification, use cross-entropy loss function (Cross-Entropy Loss) to optimize maturity classification; S350, replace the loss function, replace the default CIOU loss function of YOLOv5 with the MPDIoU loss function to improve the regression stability and positioning accuracy of the bounding box; MPDIoU loss function The formula is as follows: in, i is a positive integer, indicating the number of the key point. n is a positive integer, indicating the total number of key points in the target box and the predicted box. IoU is the intersection-over-union ratio of the target box and the prediction box. The prediction box is the bounding box output by YOLOv5, which indicates the prediction result of YOLOv5 on the target position. The target box is the bounding box manually provided during data annotation, which indicates the actual position and range of the target in the image. The Euclidean distance between the center point of the prediction box and the target box: ( , )and( , ) are the center point coordinates of the prediction box and the target box respectively, is the diagonal length of the image in the banana object detection dataset after image data processing and enhancement; It is the sum of the Euclidean distances between the key points (including vertices and midpoints) of the prediction box and the target box; is an aspect ratio constraint used to optimize the shape of the box: , and , The width and height of the prediction box and target box respectively; α , β , γ are weight hyperparameters, which respectively control the influence of center point distance, key point distance and aspect ratio constraint. Preferably, their values are α =0.5, β =0.3, γ =0.2; S360, synthetic joint loss function, combining target detection and maturity classification tasks, and synthesizing the joint loss function of multi-task learning , the formula is as follows: in, is the target detection loss function, is the classification loss function for maturity classification, and It is a loss weight hyperparameter that controls the trade-off between detection and classification tasks. The target detection loss function uses the MPDIoU loss function: ; The classification loss function for maturity classification uses the cross entropy loss function: Where N is the number of samples. k is the number of maturity categories, k =3, is the true classification label of the target instance, indicating i objects (banana instances) and classification categories j , the value is 0 or 1, 0 means the target instance does not belong to the classification category, 1 means the target instance belongs to the classification category, is the predicted maturity probability.
[0076] S400, improve YOLOv5 training, load training set, adjust training parameters, and train the improved YOLOv5, specifically including: S410, load the training set, load the training set into the training framework, set the initial training parameters, including input image size, batch size, learning rate, maximum number of training rounds (epochs), optimizer, weight decay parameter, data enhancement strategy, the default input image size is 640×640 pixels, which can be adjusted to 320×320~1280×1280 pixels according to the accuracy requirements; the default batch size is 16, which can be adjusted to 32 or 64 according to the GPU memory; the initial value of the learning rate is set to 0.01, and the cosine annealing scheduling strategy is used to gradually reduce it to improve stability; the maximum number of training rounds is set to 300 rounds; the optimizer selects AdamW (an optimization algorithm based on the Adam optimizer with weight decay regularization added); the weight decay parameter is set to 0.0001 to reduce overfitting; the data enhancement strategy includes random scaling, cropping, and flipping to improve the generalization ability of the model, and the training set ratio is set to 70%; S420, load pre-trained weights, which are downloaded from the official GitHub repository of YOLOv5 or other related resource websites, using a multi-task joint loss function , initial weight and Set to 1:1 to balance the target detection loss and maturity classification loss; S430, improve YOLOv5 training, perform forward propagation and back propagation on the training set batch by batch, and calculate the target detection loss of the training set through forward propagation in each training round (epoch) and maturity classification loss , and combined with the multi-task joint loss function Perform back propagation optimization; after each round of training, calculate and save the key indicators of mean average precision (mAP), detection accuracy, and recall rate of the validation set; when the training end conditions are met, complete the training of the improved YOLOv5; the training end conditions include reaching the maximum number of training rounds or satisfying the early stopping strategy or the loss function is stably converged. The early stopping strategy is that the mean average precision (mAP) of the validation set or the detection accuracy has not improved for 10 consecutive rounds, and there is no obvious downward trend in the validation loss. The loss function is stably converged when the training loss The decrease is lower than the preset threshold, the preset threshold = 0.01; S440, calculate the best training weight, compare the mean average precision (mAP) of the validation set, and the weight of the improved YOLOv5 when the mean average precision (mAP) is the largest is the best training weight.
[0077] S500, banana target detection, maturity classification and counting, uses the trained improved YOLOv5 to perform target detection, maturity classification and counting of bananas; specifically includes: S510, optimal training weight loading, loading the optimal training weight, and using the improved YOLOv5 after training; S520, target detection and maturity classification, inputting the image or / and video stream of the banana to be detected into the trained improved YOLOv5 to perform target detection and maturity classification, and obtaining the bounding box of the banana to be detected, the detection confidence and the probability distribution of maturity classification; S530, bounding box optimization, uses the non-maximum suppression (NMS) algorithm to optimize the bounding boxes generated by the improved YOLOv5 prediction, removes redundant boxes and retains high-confidence bounding boxes; S540 , object counting, counting the optimized bounding boxes, counting the number of optimized bounding boxes according to maturity classification, and obtaining the count of each maturity classification.
[0078] S600, the results are output in real time, and the banana target frame and target quantity, maturity classification and counting results are output and displayed in real time.
[0079] In order to verify the technical effect of this patent, the applicant conducted an experiment to collect banana pictures through a high-definition industrial camera. The collection scenes include different lighting conditions (including strong light, shadow), occlusion (including leaf occlusion, stacking) and complex backgrounds (including ground, boxes, trellises). Collect banana image samples with different maturity, divided into three categories: unripe (green), ripe (ripe) and overripe (overripe). The image resolution is 1920×1080. A total of 3,000 sample images are collected, with about 1,000 samples of each category. The training set, validation set and test set are divided, and the training set: validation set: test set = 7:1:2. Use the training set to train the improved YOLOv5 model in the GPU environment (NVIDIA RTX 4060), determine the improved YOLOv5 after training through the validation set, and use the test set for banana target detection, maturity classification and counting. Compared with the prior art, this application effectively improves the accuracy of banana target detection, maturity classification and counting capabilities.
[0080] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A banana target detection, maturity classification and counting method based on improved YOLOv5, characterized in that: The following steps are involved: S100, banana image acquisition, using an image sensor to collect banana images in different scenes, constructing a banana target detection data set, and annotating them; S200, image data processing and enhancement, performing image data processing and enhancement on the labeled banana object detection data set, and dividing the training set, the validation set, and the test set into a specified proportion; S300, YOLOv5 improvement, adding EMA module to the backbone network of YOLOv5 to enhance feature extraction capability, adding detection head and feature layer to the feature fusion network of YOLOv5 to improve the detection accuracy of small targets and stacked targets, adding classifier to the detection head of YOLOv5 to classify the maturity of bananas, replacing the default loss function to improve the regression stability and positioning accuracy of the bounding box, combining the target detection and maturity classification tasks to synthesize the joint loss function, and obtaining the improved YOLOv5; S400, improving YOLOv5 training, loading the training set, adjusting training parameters, and training the improved YOLOv5; S500, banana target detection, maturity classification and counting, using the trained improved YOLOv5 to perform target detection, maturity classification and counting on bananas; S600, real-time output of results: real-time output and display of banana target frame and target quantity, maturity classification and counting results.
2. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 1, characterized in that, The different scenes include different lighting conditions, occlusion situations and complex backgrounds.
3. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 1, characterized in that, The image data processing and enhancement include random rotation, mirroring, brightness adjustment, scaling and adding noise.
4. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 1, characterized in that, The step S300 includes: S310, adding an EMA module, adding the EMA module to each C3 module in the YOLOv5 backbone network, to enhance the global modeling capability and attention effect of the banana feature area; S320, adding a detection head and a feature layer, adding a detection head and a feature layer to the feature fusion network of YOLOv5 to additionally adapt to small targets, thereby improving the detection accuracy of stacked bananas and small targets; S330, adding a classifier, adding a classifier to the detection head of the YOLOv5 to classify the maturity of the bananas; S340, optimize maturity classification, use the cross-entropy loss function Cross-Entropy Loss to optimize maturity classification; S350, replace the loss function, replace the default CIOU loss function of YOLOv5 with the MPDIoU loss function, and improve the regression stability and positioning accuracy of the bounding box; the MPDIoU loss function The formula is as follows: ; S360, synthetic joint loss function, combining target detection and maturity classification tasks, and synthesizing the joint loss function of multi-task learning , the formula is as follows: 。 5. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 4, characterized in that, The step S310 includes: S311, feature map grouping, grouping the input feature map in the channel dimension; the number of channels in each group is C / N, where C is the number of input feature map channels, N is the number of groups, and the input feature map comes from the backbone network and the feature fusion network; S312, query vector generation, generating a query vector for each group of input feature maps through a shared linear projection matrix Q , key vector K , value vector V , the projection matrix is ; S313, attention weight matrix calculation, use dot product to calculate the attention weight matrix of each group A , the formula is as follows: ; S314: generating a weighted feature map, converting the attention weight matrix A Acting on a vector of values V , generate weighted feature map , the formula is as follows: ; S315: weighted feature map splicing: splicing the weighted feature maps of all groups to restore the number of channels of the input feature map C , and reconstruct the spatial dimension of the input feature map H×W ; S316: generating an output feature map by adding the generated weighted feature map to the input feature map element by element to generate an output feature map .
6. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 4, characterized in that, The newly added feature layer introduces the intermediate output P3 feature map of the YOLOv5 backbone network into the feature fusion stage to output a new feature map as the input of the newly added detection head.
7. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 6, characterized in that, The newly added detection head performs convolution on the new feature map output by the newly added feature layer to further extract high-level features.
8. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 4, characterized in that, The step S320 includes: S321, extracting feature maps, adding a new feature layer to extract P3 feature maps and P2 feature maps from the backbone network, performing a convolution on the P3 feature map, and then upsampling it, and aligning it with the P2 feature map through the nearest neighbor interpolation method; S322, forming a new feature map, splicing the upsampled P3 feature map and the P2 feature map along the channel dimension to form a new feature map; S323, extracting features, performing feature extraction on the new feature map through two groups of convolutional layers in sequence; S324, extracting high-level features, using the new feature map after feature extraction as the input of the newly added detection head, further extracting high-level features, and performing banana maturity classification and regression prediction.
9. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in any one of claim 1, characterized in that, The step S400 includes: S410, loading a training set, loading the training set into a training framework, and setting initial training parameters; S420, load pre-trained weights and use multi-task joint loss function , initial weight and Set to 1:1 to balance the target detection loss and maturity classification loss; S430, improving YOLOv5 training, performing forward propagation and back propagation on the training set batch by batch, and in each training round, calculating the target detection loss of the training set through the forward propagation and maturity classification loss , and combined with the multi-task joint loss function Perform the back propagation optimization; after each round of training, calculate and save the average precision mean, detection accuracy, and recall rate key indicators of the verification set; when the training end condition is reached, complete the training of the improved YOLOv5; S440, calculating the best training weight, comparing the average precision means of the validation set, and the weight of the improved YOLOv5 when the average precision means is the largest is the best training weight.
10. The banana target detection, maturity classification and counting method based on improved YOLOv5 as claimed in claim 9, characterized in that, The step S500 includes: S510, optimal training weight loading, loading the optimal training weight, and using the improved YOLOv5 after training; S520, target detection and maturity classification, inputting the image or / and video stream of the banana to be detected into the trained improved YOLOv5 to perform target detection and maturity classification, and obtaining the bounding box of the banana to be detected, the detection confidence and the probability distribution of the maturity classification; S530, bounding box optimization, using a non-maximum suppression algorithm to optimize the bounding box generated by the improved YOLOv5 prediction, remove redundant boxes and retain high-confidence bounding boxes; S540: Target counting: counting the optimized bounding boxes, and counting the number of the optimized bounding boxes according to the maturity classification to obtain the count of each maturity classification.