Pineapple maturity target detection method based on YOLOv10 model
By improving the YOLOv10 model, ECA attention, deformable convolution and multi-scale residual dense blocks are introduced, which solves the accuracy and speed limitations of pineapple maturity detection in the prior art, as well as the accuracy of leaf occlusion and fruit eye position and size prediction, and achieves more efficient and accurate pineapple maturity detection.
Patent Information
- Application Number
- CN202510550730.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing YOLO-based pineapple maturity recognition method has limitations in detection accuracy and speed when facing the complexity and subtle characteristics of pineapple fruit eye maturity, and cannot effectively solve the accuracy of leaf occlusion and fruit eye position and size prediction.
Using the improved YOLOv10 model, a pineapple maturity detection model was constructed by introducing ECA attention mechanism, deformable convolution DCN and multi-scale residual dense blocks. This model improves the feature extraction network structure, improves the detection ability of pineapple eyes of different sizes, and enhances the accuracy of predicting the position and size of pineapple fruit eyes.
It significantly improves the accuracy and efficiency of pineapple maturity detection, enhances the robustness of the model, can more accurately detect the maturity of pineapple fruit eyes, and effectively solves the problem of leaf occlusion, meeting the demand for high-precision detection in agricultural production.
Smart Images

Figure CN120088774A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fruit and vegetable maturity detection, and particularly to a method for detecting the maturity of pineapples based on the YOLOv10 model. Background Art
[0002] Currently, the picking method of pineapples is mainly manual picking at one time according to the maturity degree. The detection of pineapple maturity mainly relies on manual experience, and depends on the skin color, appearance and fruit size of pineapples to judge, with low accuracy, often resulting in problems such as low picking efficiency and poor storage tolerance. Rapid maturity recognition of pineapples through machine vision is an important support for mechanized picking.
[0003] In recent years, with the rapid development of artificial intelligence technology, object detection algorithms based on deep learning have been widely used in the recognition of the maturity of fruits and vegetables. The fruit maturity detection algorithms based on deep learning are mainly divided into two categories. One category is the algorithms that first generate candidate regions and then classify and perform bounding box regression on them, mainly including R-CNN, Faster R-CNN, Mask-CNN, etc.; the other category is the algorithms that directly predict the category and bounding box of the object in a single neural network, mainly including the YOLO object detection algorithm, the SSD (Single Shot MultiBox Detector) object detection algorithm, etc. Comparing the two, algorithms such as R-CNN have higher accuracy, while algorithms such as YOLO have faster processing speed.
[0004] Previous YOLO-based fruit and vegetable maturity recognition mainly classified and recognized fruits based on the surface color of the fruit. However, when faced with fruits like pineapples where the skin color changes irregularly and is affected by multiple factors, there are situations where the skin color of the same pineapple shows multiple maturity levels at different angles. This situation leads to misdetection problems in the machine vision-based maturity recognition method of pineapples.
[0005] As an advanced single-stage object detection algorithm, YOLOv10 can effectively identify objects through its unique network architecture and optimization technology. Although YOLOv10 performs well in detection accuracy and speed, there are still some limitations in recognizing the complexity and subtle features of pineapple fruit eyes' maturity. At the same time, when faced with situations such as irregular arrangement of pineapple fruit eyes, large differences in size, different growth angles of pineapple fruit eyes, different maturity levels of fruit eyes on the same pineapple, and fruit eyes being blocked by leaves, there are problems of missed detection and misdetection, which cannot meet the needs of agricultural production. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for detecting the maturity of pineapples based on the YOLOv10 model, which can detect pineapple fruit eyes that cannot be detected by existing detection models and can effectively solve the problem of fruit eyes of pineapples being blocked by leaves.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A method for detecting the ripeness of pineapples based on the YOLOv10 model, comprising:
[0009] Obtaining a pineapple image to be detected;
[0010] Inputting the pineapple image to be detected into a pineapple ripeness detection model, outputting the ripeness score of the fruit eyes in the pineapple image to be detected, and obtaining the ripeness level of the pineapple in the pineapple image to be detected, wherein the pineapple ripeness detection model is constructed based on an improved YOLOv10 model and obtained through training with a training set, the improved YOLOv10 model is obtained by improving the feature extraction network structure of the traditional YOLOv10 model, and the training set is a pineapple image with annotations of the ripeness of pineapple fruit eyes.
[0011] Optionally, obtaining the training set includes: collecting original pineapple images during the growth process, eliminating the original pineapple images that do not conform to the preset rules, and performing ripeness annotation on the pineapple fruit eyes in the remaining original pineapple images to obtain the training set.
[0012] Optionally, improving the feature extraction network structure of the traditional YOLOv10 model includes: replacing the convolution of the neck network in the traditional YOLOv10 model with deformable convolution, embedding the ECA attention mechanism between the C2f layer and the Detect layer of the traditional YOLOv10 model, and embedding multi-scale residual dense blocks in the C2F-CIB module of the traditional YOLOv10 model.
[0013] Optionally, embedding the multi-scale residual dense blocks in the C2F-CIB module of the traditional YOLOv10 model includes:
[0014] Replacing the convolution in the C2F-CIB module with serial multi-scale convolution, and performing residual dense connection on the serial multi-scale convolution, wherein the multi-scale residual dense blocks perform multi-scale convolution operations on the preliminary features obtained by the backbone network and the neck network of the input pineapple ripeness detection model, extract the feature information of the pineapple fruit eyes, input the feature information of the pineapple fruit eyes and the initial features into a convolution layer for dimensionality reduction, and add the features after dimensionality reduction to the preliminary features to obtain the final output feature map.
[0015] Optionally, the multi-scale residual dense blocks performing multi-scale convolution operations on the preliminary features obtained by the backbone network and the neck network of the input pineapple ripeness detection model includes:
[0016] Input the preliminary features into the first-scale convolutional layer for convolutional operation to obtain the first feature map;
[0017] Input the first feature map and the preliminary features into the second-scale convolutional layer for concatenation and convolutional operation to obtain the second feature map;
[0018] Input the first feature map, the second feature map and the preliminary features into the third-scale convolutional layer for concatenation and convolutional operation to obtain the third feature map;
[0019] Input the first feature map, the second feature map, the third feature map and the preliminary features into the fourth-scale convolutional layer for concatenation and convolutional operation to obtain the fourth feature map.
[0020] Optionally, obtaining the maturity level of the pineapple in the to-be-detected pineapple image includes: assigning different weights to the fruit eyes in the mature stage, the mid-mature stage, the immature stage, and the development stage respectively, performing weighted averaging according to the weights and the fruit eye maturity scores in the to-be-detected pineapple image, obtaining the pineapple maturity score in the to-be-detected pineapple image, and obtaining the maturity level of the pineapple in the to-be-detected pineapple image according to the pineapple maturity score.
[0021] Optionally, the expression of the weighted average is:
[0022] ;
[0023] ;
[0024] where D is the pineapple maturity score, X is the total number of labeled fruit eyes, x 1 , x 2 , x 3 , x 4 are the numbers of fruit eyes in the mature stage, the mid-mature stage, the immature stage, and the development stage respectively, D1 k1 , D2 k2 , D3 k3 , D4 k4 are the fruit eye scores of the k 1 th fruit eye in the mature stage, the fruit eye scores of the k 2 th fruit eye in the mid-mature stage, the fruit eye scores of the k 3 th fruit eye in the immature stage, and the fruit eye scores of the k 4 th fruit eye in the development stage respectively.
[0025] The beneficial effects of the present invention are as follows: By introducing the ECA attention mechanism into the YOLOv10 model, the present invention improves the model's detection ability for pineapple eyes of different sizes and enhances the prediction accuracy of the position and size of pineapple fruit eyes. Introducing the deformable convolution network (DCN) to replace the ordinary convolution enables the model to better adapt to the scale and shape changes of pineapple fruit eyes, improving the robustness of the model. Introducing the multi-scale convolutional residual dense block in C2f-CIB to form C2f-RDB improves the reuse rate of features, strengthens the extraction of feature information of fruit eyes of different sizes, obtains more fruit eye feature representations, and significantly improves the detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 It is an example diagram of the calyx apex angle and fruit eye diameter of an embodiment of the present invention;
[0028] Figure 2 It is the fruit eye feature map of pineapples with different maturities in an embodiment of the present invention. Among them, (a) is the pineapple fruit in the development period, (b) is the pineapple fruit in the immature period, (c) is the pineapple fruit in the mid-ripe period, (d) is the pineapple fruit in the mature period, (e) is the side view of the fruit eye in the development period, (f) is the side view of the fruit eye in the immature period, (g) is the side view of the fruit eye in the mid-ripe period, (h) is the side view of the fruit eye in the mature period, (i) is the top view of the fruit eye in the development period, (j) is the top view of the fruit eye in the immature period, (k) is the top view of the fruit eye in the mid-ripe period, and (l) is the top view of the fruit eye in the mature period;
[0029] Figure 3 It is the structure diagram of C2f-RDB and C2f-CIB in an embodiment of the present invention. Among them, (a) is the structure diagram of C2f-RDB, and (b) is the structure diagram of C2f-CIB;
[0030] Figure 4 It is a schematic diagram of the test results of the pineapple maturity detection method based on the improved YOLOv10 in an embodiment of the present invention;
[0031] Figure 5 It is the structure diagram of the improved YOLOv10 model in an embodiment of the present invention;
[0032] Figure 6 It is the structure diagram of the ECA attention mechanism in an embodiment of the present invention;
[0033] Figure 7 It is the structure diagram of the DCN convolution in an embodiment of the present invention;
[0034] Figure 8 It is the structure diagram of the Residual Dense Block (RDB) of the embodiment of the present invention;
[0035] Figure 9 It is the Loss curve diagram of different models of the embodiment of the present invention;
[0036] Figure 10 It is the flowchart of a pineapple ripeness target detection method based on the YOLOv10 model of the embodiment of the present invention. Detailed implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0039] As Figure 10 shown, this embodiment provides a pineapple ripeness target detection method based on the YOLOv10 model, including:
[0040] Obtain the pineapple image to be detected;
[0041] Input the pineapple image to be detected into the pineapple ripeness detection model, output the fruit eye ripeness score in the pineapple image to be detected, and obtain the ripeness level of the pineapple in the pineapple image to be detected. Among them, the pineapple ripeness detection model is constructed based on the improved YOLOv10 model and obtained through training with a training set. The improved YOLOv10 model is obtained by improving the feature extraction network structure of the traditional YOLOv10 model, and the training set is a pineapple image with annotations of pineapple fruit eye ripeness.
[0042] This embodiment provides a pineapple ripeness recognition method based on the improved YOLOv10. By introducing the Efficient Channel Attention (ECA) attention mechanism, replacing ordinary convolutions with Deformable Convolutional Networks (DCN), and introducing multi-scale convolutions and residual dense blocks in C2f-CIB to obtain C2f-RDB, the detection efficiency and accuracy are significantly improved.
[0043] Further, obtaining the training set includes: collecting the original pineapple images during the growth process, removing the original pineapple images that do not meet the preset rules, labeling the maturity of the pineapple eyes in the remaining original pineapple images, and obtaining the training set.
[0044] Specifically, collect the pineapple images during the growth process and conduct screening, remove the original state images that do not meet the preset rules, obtain the target images as the data set, divide the maturity of the pineapple fruits into four levels according to the morphological patterns of the pineapple eyes at 14 days, 42 days, 70 days, and 91 days after fruit setting, label them, obtain the data set of the maturity images of the pineapple eyes, and divide the data set into a training set, a validation set, and a test set.
[0045] Further, improving the feature extraction network structure of the traditional YOLOv10 model includes: replacing the convolution of the neck network in the traditional YOLOv10 model with deformable convolution, embedding the ECA attention mechanism between the C2f layer and the Detect layer of the traditional YOLOv10 model, and embedding multi-scale residual dense blocks in the C2F-CIB module of the traditional YOLOv10 model.
[0046] Furthermore, embedding multi-scale residual dense blocks in the C2F-CIB module of the traditional YOLOv10 model includes:
[0047] Replacing the convolution in the C2F-CIB module with serial multi-scale convolution, and performing residual dense connection on the serial multi-scale convolution. Among them, the multi-scale residual dense block performs multi-scale convolution operations on the preliminary features obtained by the input pineapple maturity detection model through the backbone network and the neck network, extracts the feature information of the pineapple eyes, inputs the feature information of the pineapple eyes and the initial features into the convolutional layer for dimensionality reduction, and adds the features after dimensionality reduction to the preliminary features to obtain the final output feature map.
[0048] Specifically, the working steps of the C2f-RDB module include: First, replace the ordinary convolution in the C2f-CIB with multi-scale convolution of 1×1, 3×3, 5×5, and 7×7, perform serial operations on the multi-scale convolution of 1×1, 3×3, 5×5, and 7×7, and extract the feature information of the pineapple eyes at four maturity levels. Perform residual dense connection on the four serial multi-scale convolutions. The input end of each convolutional layer is the input combination of all previous layers. Concatenate all intermediate feature maps along the channel dimension, perform dimensionality reduction through a 1x1 convolutional layer, fuse the local features to obtain the final feature map, and add the result of the local feature fusion to the input feature map to obtain the final output feature map.
[0049] The working steps of DCN are as follows: First, use traditional convolution operations to generate feature maps. Second, while generating the feature maps, predict the offset of each convolution position through another convolutional layer to obtain a two-dimensional vector representing the offset of the convolution kernel center relative to the standard position. Use the predicted offset to adjust the position of the standard convolution kernel, and then apply the convolution operation to adaptively adjust the local set structure of the input feature map.
[0050] The ECA attention mechanism is embedded between the C2f layer and the Detect layer of the YOLOv10 model. By suppressing the irrelevant information output by the C2f layer, it enhances the final target localization and classification capabilities of the Detect layer, thereby improving the model's detection ability for pineapple eyes of different sizes and enhancing the prediction accuracy of the position and size of pineapple fruit eyes.
[0051] Furthermore, the multi-scale residual dense blocks perform multi-scale convolution operations on the preliminary features obtained by the input pineapple maturity detection model through the backbone network and the neck network, including:
[0052] Input the preliminary features into the first-scale convolutional layer for convolution operations to obtain the first feature map;
[0053] Input the first feature map and the preliminary features into the second-scale convolutional layer for splicing and convolution operations to obtain the second feature map;
[0054] Input the first feature map, the second feature map, and the preliminary features into the third-scale convolutional layer for splicing and convolution operations to obtain the third feature map;
[0055] Input the first feature map, the second feature map, the third feature map, and the preliminary features into the fourth-scale convolutional layer for splicing and convolution operations to obtain the fourth feature map, where the first feature map, the second feature map, the third feature map, and the fourth feature map are the extracted pineapple fruit eye feature information.
[0056] Further, obtaining the maturity level of the pineapple image to be detected includes: assigning different weights to the fruit eyes in the mature stage, mid-mature stage, immature stage, and development stage respectively, and performing weighted averaging on the pineapple fruits in the pineapple image to be detected according to the weights and the fruit eye maturity scores in the pineapple image to be detected to obtain the pineapple maturity score in the pineapple image to be detected, and obtaining the maturity level of the pineapple in the pineapple image to be detected according to the pineapple maturity score.
[0057] Furthermore, the expression of weighted averaging is:
[0058] ;
[0059] ;
[0060] Among them, D is the pineapple maturity score, X is the total number of marked fruit eyes, and x 1 , x 2 , x 3 , x 4 are the numbers of fruit eyes at the mature stage, mid-mature stage, immature stage, and development stage of the pineapple respectively. D1 k1 , D2 k2 , D3 k3 , D4 k4 are the fruit eye scores of the k 1 -th fruit eye at the mature stage, the fruit eye score of the k 2 -th fruit eye at the mid-mature stage, the fruit eye score of the k 3 -th fruit eye at the immature stage, and the fruit eye score of the k 4 -th fruit eye at the development stage respectively.
[0061] The following further explains the method of this embodiment with reference to the attached Figures 1-9 :
[0062] The method for identifying the maturity of pineapples in this embodiment is based on the number of days of fruit setting, the apical angle of the fruit eye calyx of the pineapple, and the fruit eye diameter; the apical angle of the calyx and the fruit eye diameter are as Figure 1 shown. E is the apical angle of the calyx, d is the width of the fruit eye. A method for detecting the maturity of pineapples based on the YOLOv10 model specifically includes the following steps:
[0063] Collect pineapple images during the growth process and screen them, eliminating the original state images that do not meet the preset rules to obtain the target images as the data set;
[0064] According to the fruit eye morphology of the pineapple, the fruit eyes of pineapples with different maturities are divided into four grades according to the fruit eye morphology at 14 days, 42 days, 70 days, and 91 days after fruit setting to obtain a data set of pineapple fruit eye maturity images. The characteristics of pineapple fruit eyes with different maturities are as Figure 2 (a)-(l) shown, and the maturity grade division is shown in Table 1.
[0065] Table 1
[0066]
[0067] Improve the traditional YOLOv10 model by replacing the ordinary convolution in the original model with deformable convolution, adding an ECA attention mechanism, and integrating the C2f-RDB module with multi-scale residual dense units to construct a pineapple maturity detection model. The pineapple maturity detection model is the improved YOLOv10 model as Figure 5 shown.
[0068] In this example, a new C2f-RDB module, namely a multi-scale residual dense block, is designed by combining the residual dense unit RDB (Residual Dense Block), multi-scale convolution, and the C2f module. The structural diagrams of the C2f-RDB module and the C2f-CIB module are shown in Figure 3 Figure (a) and Figure 3 Figure (b), and the residual dense unit RDB is shown in Figure 8 Figure
[0069] First, replace the ordinary convolution in the C2f-CIB module with multi-scale convolutions of sizes 1×1, 3×3, 5×5, and 7×7.
[0070] Perform convolution operations of 1×1, 3×3, 5×5, and 7×7 on the preliminary features obtained from the backbone network and the neck network for the input pineapple fruit eye dataset. The input end of each convolutional layer is the input combination of all previous layers. Add the output of the last layer of the dense convolution to the original input to obtain the final output.
[0071] The working steps of the C2f-RDB module include:
[0072] First, replace the ordinary convolution in the C2f-CIB with multi-scale convolutions of 1×1, 3×3, 5×5, and 7×7.
[0073] Perform serial operations on the multi-scale convolutions of 1×1, 3×3, 5×5, and 7×7 to extract the feature information of pineapple fruit eyes at four maturity levels;
[0074] Perform residual dense connections on the four serial multi-scale convolutions. The input end of each convolutional layer is the input combination of all previous layers;
[0075] Concatenate all intermediate feature maps along the channel dimension;
[0076] Reduce the dimension through a 1x1 convolutional layer to fuse the local features to obtain the final feature map;
[0077] Add the result of fusing the local features to the input feature map, i.e., the preliminary features, to obtain the final output feature map.
[0078] The splicing process of the multi-scale RDB module is as follows:
[0079] The input feature map is F 0 ;
[0080] The first layer: Input: F 0 , Convolution operation: F 1 = Conv 1×1 (F 0 );
[0081] Second layer: Input: concat(F 0 + F 1 ), Convolution operation: F 2 = Conv 3×3 (concat(F 0 + F 1 ));
[0082] Third layer: Input: concat(F 0 + F 1 + F 2 ), Convolution operation: F 3 = Conv 5×5 (concat(F 0 + F 1 + F 2 ));
[0083] Fourth layer: concat(F 0 + F 1 + F 2 + F 3 ), Convolution operation: F 4 = Conv 7×7 (concat(F 0 + F 1 + F 2 + F 3 ));
[0084] The local feature fusion process is: F local = Conv 1×1 (concat(F 0 + F 1 + F 2 + F 3 + F 4 ));
[0085] The residual connection output is: F out= F local + F 0 .
[0086] Among them, F 0 , F 1 , F 2 , F 3 , F 4 are the preliminary features, the first feature map, the second feature map, the third feature map, and the fourth feature map obtained through the backbone network and the neck network respectively. F local is the local feature fusion map after dimensionality reduction, and F out is the final output feature map. concat is the concatenation operation.
[0087] In this embodiment, the output of the last layer of the dense convolution part is added to the original input, which alleviates the problem of gradient disappearance, promotes feature reuse and gradient flow, and enhances the expressive power of the model.
[0088] Replace the ordinary convolution in the neck network of YOLOv10 with DCN. The DCN convolution structure diagram is as Figure 7 shown. The working steps of DCN are as follows:
[0089] First, use traditional convolution operations to generate feature maps.
[0090] Secondly, while generating the feature maps, predict the offset of each convolution position through another convolution layer. A two-dimensional vector is obtained, representing the offset of the convolution kernel center relative to the standard position.
[0091] Use the predicted offset to adjust the position of the standard convolution kernel, and then apply convolution operations to adaptively adjust the local set structure of the input feature map.
[0092] The deformable convolution calculation process is as follows:
[0093] ;
[0094] ;
[0095] In the formula, represents the feature at position in the output feature map . represents the pixel value of the input feature map at position . is the center position of the convolution kernel, is the enumeration value relative to the center position in the convolution kernel. represents the th weight of the convolution kernel. represents the added offset. defines the receptive field size and dilation. As shown in formula defines a convolution kernel with a size of 3×3 and a dilation rate of 1.
[0096] Embed the ECA attention mechanism between the C2f layer and the Detect layer of the YOLOv10 model. By suppressing the irrelevant information output by the C2f layer, enhance the final target localization and classification capabilities of the Detect layer, thereby improving the model's detection ability for pineapple eyes of different sizes and enhancing the prediction accuracy of the position and size of pineapple fruit eyes. The Detect layer is part of the Head layer and is responsible for specific detection tasks. In YOLOv10, the Detect layer adopts a dual-head design, namely One-to-many Head and One-to-one Head. These two heads participate in calculating the loss simultaneously during training, while only the One-to-one Head is used during inference and prediction. This design helps to solve the redundant prediction problem in post-processing and achieves end-to-end detection without Non-Maximum Suppression (NMS-fre).
[0097] As Figure 6 shown in the structure diagram of the ECA attention mechanism, the working steps of the ECA attention mechanism are as follows:
[0098] First, perform global average pooling (Global Average Pooling, GAP) on the input feature map to compress the features of each channel into a scalar value.
[0099] Next, use one-dimensional convolution (1D Convolution) to model the interdependence between channels: , where k is the convolution kernel and φ represents a series of operations in the ECA module.
[0100] The kernel size of the one-dimensional convolution is calculated through a fixed formula, which is based on the parity of the number of channels C.
[0101] The formula is as follows:
[0102] ;
[0103] where the input feature map has a shape of (B, C, H, W), where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively, γ is a scaling factor, b is a constant, and odd represents the odd number closest to k.
[0104] After the one-dimensional convolution, use an activation function Sigmoid to map the attention weights to the interval [0, 1].[[]] Figure 6 In is the Sigmoid activation function.
[0105] Finally, apply the calculated attention weights to the input feature map to adjust the importance of each channel. The shape of the output feature map output remains (B, C, H, W).
[0106] Calculate the comprehensive maturity score of the pineapple by using the above weighted average formula for pineapple fruit eyes prediction value.
[0107] Assign weights of 0.14, 0.42, 0.70, and 0.91 to the fruit eyes in the development stage (14 days after fruit setting), immature stage (42 days after fruit setting), mid-mature stage (70 days after fruit setting), and mature stage (91 days after fruit setting) respectively. The weighted average formula for pineapple fruit is as follows:
[0108] ;
[0109] ;
[0110] where D is the pineapple maturity score, X is the total number of labeled fruit eyes, x 1 , x 2 , x 3 , x 4 are the numbers of fruit eyes in the mature stage, mid-mature stage, immature stage, and development stage respectively, and D1 k1 , D2 k2 , D3 k3 , D4 k4 are the fruit eye scores of the k 1 th fruit eye in the mature stage, the k 2 th fruit eye in the mid-mature stage, the k 3 th fruit eye in the immature stage, and the k 4 th fruit eye in the development stage respectively. When the D value is in the ranges of 0.91 - 0.70, 0.70 - 0.42, 0.42 - 0.14, and 0.14 - 0, it corresponds to the mature, mid-mature, immature, and development stage pineapple grades respectively.
[0111] Feed the training set and validation set data into the model for training and validation. Perform validation using the validation set once for each round of training. After training a certain number of times, according to the results, if the weights obtained from training are used to test the test set and the effect does not meet the requirements, further adjust the training parameters and retrain until the results meet the usage standards. The obtained optimal detection model, as Figure 4 shown, is the pineapple maturity detection result using the test set for testing.
[0112] Input the pineapple image to be detected into the obtained optimal detection model, and the optimal detection model detects pineapples with different maturities from the pineapple image to be detected.
[0113] The specific implementation method is as follows:
[0114] In this embodiment, Precision (P), Recall (R), and mean Average Precision (mAP) are mainly used as the evaluation metrics of the model. The calculation of Precision P is shown in formula (1), and the calculation of Recall R is shown in formula (2). The formula for calculating the average precision is shown in (3), and the mean Average Precision mAP is the average precision X of all classes AP After summing and then taking the mean value.
[0115] (1);
[0116] (2);
[0117] (3);
[0118] Among them, TP (True Positive) is the true positive example, that is, the number of samples that the model correctly predicts as the positive class. FP (False Positive) is the false positive example, that is, the number of samples that the model wrongly predicts as the positive class. FN (False Negative) represents the number of unrecognized relevant information. To verify that each part of the improvement in this embodiment contributes to improving the model performance, the original YOLOv10 model is used as the baseline model, and ablation experiments are carried out and the model performance is evaluated using the evaluation metrics of mAP@0.5, mAP@0.5:0.95, P, and R. The results of the ablation experiments are shown in Table 2.
[0119] Table 2
[0120]
[0121] To verify the performance of the optimized YOLOv10 model in this embodiment, a comparative experiment is carried out on the optimized YOLOv10 model, the mainstream object detection models Faster-RCNN, SSD, RetinaNet, YOLO5, YOLOv7, YOLOv8, and the YOLOv10 baseline model on the same device and the same dataset. The comparison of the detection effects is shown in Table 3.
[0122] Table 3
[0123]
[0124] The experimental equipment configuration is as follows:
[0125] The operating system used in the experiment is Windows 10, the CPU model is Intel (R) Core (TM) i7-12700F CPU @ 2.10Ghz, and the GPU is NVIDIA GeForce RTX3080Ti (40G). It is trained using the PyTorch 1.10.0 deep learning framework, and the GPU acceleration library is CUDA 11.5. The SGD optimizer is used during learning, with an initial learning rate of 0.01, a weight decay of 0.0005, a learning momentum of 0.900, and a network input dimension of 640×640. The model is trained for 200 epochs with a batch size of 32. The loss value drops rapidly during the first 50 epochs of training and basically stabilizes after 100 epochs, as Figure 9 shown. After 200 epochs of training, the loss value converges to 0.027.
[0126] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A pineapple maturity target detection method based on the YOLOv10 model, characterized in that: include: Get the pineapple image to be detected; The pineapple image to be detected is input into a pineapple maturity detection model, and the fruit eye maturity score in the pineapple image to be detected is output to obtain the maturity level of the pineapple in the pineapple image to be detected, wherein the pineapple maturity detection model is constructed based on an improved YOLOv10 model and obtained through training with a training set, the improved YOLOv10 model is obtained by improving the feature extraction network structure of the traditional YOLOv10 model, and the training set is a pineapple image with pineapple eye maturity annotations.
2. The pineapple maturity target detection method based on the YOLOv10 model according to claim 1, characterized in that, Acquiring the training set includes: collecting original pineapple images during the growth process, removing the original pineapple images that do not meet preset rules, and marking the maturity of pineapple eyes in the remaining original pineapple images to obtain the training set.
3. The pineapple maturity target detection method based on the YOLOv10 model according to claim 1, characterized in that, The feature extraction network structure of the traditional YOLOv10 model is improved, including: replacing the convolution of the neck network in the traditional YOLOv10 model with a deformable convolution, embedding the ECA attention mechanism between the C2f layer and the Detect layer of the traditional YOLOv10 model, and embedding a multi-scale residual dense block in the C2F-CIB module in the traditional YOLOv10 model.
4. The pineapple maturity target detection method based on the YOLOv10 model according to claim 3, characterized in that, Embedding the multi-scale residual dense block in the C2F-CIB module in the traditional YOLOv10 model includes: The convolution in the C2F-CIB module is replaced with a serial multi-scale convolution, and the serial multi-scale convolution is subjected to residual dense connection, wherein the multi-scale residual dense block performs a multi-scale convolution operation on the preliminary features obtained by inputting the pineapple maturity detection model through the backbone network and the neck network, extracts the pineapple eye feature information, inputs the pineapple eye feature information and the initial features into the convolution layer for dimensionality reduction, adds the features after dimensionality reduction to the preliminary features, and obtains the final output feature map.
5. The pineapple maturity target detection method based on the YOLOv10 model according to claim 4, characterized in that, The multi-scale residual dense block performs a multi-scale convolution operation on the preliminary features input into the pineapple maturity detection model through the backbone network and the neck network to obtain the initial features, including: Inputting the preliminary features into a first-scale convolutional layer for convolution operation to obtain a first feature map; Inputting the first feature map and the preliminary feature into a second scale convolution layer for concatenation and convolution operations to obtain a second feature map; Inputting the first feature map, the second feature map and the preliminary feature into a third scale convolution layer for concatenation and convolution operations to obtain a third feature map; The first feature map, the second feature map, the third feature map and the preliminary feature are input into a fourth scale convolution layer for a concatenated convolution operation to obtain a fourth feature map.
6. The pineapple maturity target detection method based on the YOLOv10 model according to claim 1, characterized in that, Obtaining the maturity level of the pineapple in the pineapple image to be detected comprises: assigning different weights to ripe fruit eyes, mid-ripe fruit eyes, immature fruit eyes, and developmental fruit eyes, respectively, performing weighted averaging according to the weights and the maturity scores of the fruit eyes in the pineapple image to be detected, obtaining the maturity score of the pineapple in the pineapple image to be detected, and obtaining the maturity level of the pineapple in the pineapple image to be detected according to the pineapple maturity score.
7. The pineapple maturity target detection method based on the YOLOv10 model according to claim 6, characterized in that, The expression of the weighted average is: ; ; Among them, D is the pineapple maturity score, X is the total number of marked eyes, x1, x2, x3, x4 are the number of eyes in the mature stage, the middle mature stage, the immature stage, and the developing stage, respectively, D1 k1 、D2 k2 、D3 k3 、D4 k4 They are the fruit eye score of the k1th mature fruit eye, the fruit eye score of the k2th mid-ripe fruit eye, the fruit eye score of the k3th immature fruit eye, and the fruit eye score of the k4th developmental fruit eye.
Citation Information
Patent Citations
Lightweight YOLOv4-based pineapple maturity analysis method
CN116453111A
Fatigue driving identification method based on improved YOLOv5 neural network
CN118644841A
Infrared remote sensing image target detection method based on YOLOv8, electronic equipment and computer readable storage medium
CN118887511A
Citrus maturity detection method based on improved YOLOv8 and related device
CN118968502A
Electric transmission line nest detection method and device based on ER-YOLO
CN119359990A
Cited By
Pineapple accurate positioning and maturity discrimination method based on space and channel reconstruction
CN120853159A