Method for detecting maturity of tomatoes based on improved yo11

By improving the YOLOv11 model and introducing the wavelet transform convolution module, the problem of the existing tomato ripening algorithm being not robust in complex environments is solved, and higher detection accuracy and robustness are achieved.

CN120125908APending Publication Date: 2025-06-10ANHUI UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510282560.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing tomato ripening algorithm is not robust in complex environments, and the detection accuracy needs to be improved.

Method used

Using the improved YOLOv11 model, the C3k2_WTConv module is introduced, and the feature extraction capability is enhanced by using wavelet transform to improve the robustness of the model under multi-scale feature extraction and frequency domain noise interference.

Benefits of technology

It significantly improves the detection accuracy and robustness of tomato ripening, improves the mAP50 index, and enhances the model's adaptability and detection accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125908A_ABST
    Figure CN120125908A_ABST
Patent Text Reader

Abstract

The invention provides a tomato maturity detection method based on improved yo11, and relates to the technical field of tomato maturity detection.The tomato maturity detection method comprises the four steps of tomato image data set construction, maturity detection model construction, maturity detection model training and tomato maturity detection.A C3K2WTConv module is designed through an innovative wavelet transform convolution structure; therefore, a YOLOv11-WTConv algorithm is provided, the multi-scale feature extraction capability of a target is enhanced, the mAP50 index is improved, the detection accuracy is improved at the same time, the detection precision and robustness of tomato maturity grading are remarkably improved, and key technical support is provided for intelligent production management of facility agriculture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tomato ripeness detection, and in particular to a method for detecting tomato ripeness based on improved YOLO11. Background Art

[0002] For crops such as tomatoes with a short maturity period, harvesting has always relied on a large amount of manual labor, and currently there is a shortage of labor willing to engage in agricultural production. Therefore, it is particularly important to develop an automatic picking robot for cherry tomatoes. When the tomato picking robot is working, identifying the tomato ripeness becomes a major challenge. During the growth process of tomatoes, affected by the lighting conditions and the fruit development stage, the surface color and morphological characteristics of tomatoes show a dynamic evolution law in visual imaging.

[0003] Currently, the main research method for tomato ripeness detection is machine vision technology. This technology collects RGB images of tomato plants and, with the help of a deep neural network architecture such as a convolutional neural network, accurately identifies the mapping relationship between the morphology, color, and ripeness during the growth process of tomatoes, realizing intelligent grading determination of ripe and unripe tomatoes. Miao Ronghui et al. proposed a lightweight method for detecting tomato ripeness by improving YOLO v7. This method improves the feature expression ability of the network by introducing MobileNetV3 and a global attention mechanism module. The detection accuracy of this model is 0.1 percentage points higher than that of the original model. Zhang Yucheng et al. proposed an algorithm for detecting tomato ripeness by improving YOLOv4. While introducing the MobileNet network, a squeeze-and-excitation module (SE) and a convolutional block attention module (CBAM) are added, and the value of mAP is increased by 5.36%. Long Jiehua et al. proposed a method for segmenting tomatoes of different ripeness by improving MaskR-CNN. This method uses a cross-stage local network to fuse with the residual network in the Mask R-CNN network, and it is 2.29 percentage points higher than the method with ResNet50 as the backbone network.

[0004] However, although some progress has been made in tomato ripeness detection, there are still some deficiencies in the existing detection algorithms, such as weak robustness in complex environments and the need to improve detection accuracy. Therefore, the present invention proposes a method for detecting tomato ripeness based on improved YOLO11 to solve the problems existing in the prior art. Summary of the Invention

[0005] Aiming at the above problems, the purpose of the present invention is to propose a method for detecting tomato ripeness based on improved YOLO11. This method for detecting tomato ripeness based on improved YOLO11 has the advantage of enhancing the multi-scale feature extraction ability of the target and can solve the problems existing in the prior art.

[0006] To achieve the objectives of the present invention, the present invention is implemented through the following technical solutions: A method for detecting the maturity of tomatoes based on improved YOLO11, comprising the following steps:

[0007] Step 1: Construct a tomato image dataset

[0008] Based on the publicly available tomato dataset and the tomato image data obtained by photographing from a tomato agricultural base, perform enhancement processing and normalization processing on it, thereby constructing a tomato dataset, and then use the open-source software LabelImg for image annotation and category classification;

[0009] Step 2: Construct a maturity detection model

[0010] Use YOLOv11 as the basic model, replace the C3k2 module of YOLOv11 with the C3k2_WTConv module to obtain the YOLOv11-WTConv network model, where the C3k2_WTConv module applies wavelet transform on the basis of standard convolution;

[0011] Step 3: Training of the maturity detection model

[0012] Divide the tomato dataset constructed in Step 1 into a training set, a validation set, and a test set, use the training set to train the YOLOv11-WTConv network model constructed in Step 2, then use the validation set to evaluate the model performance, and finally use the test set to evaluate the generalization ability and actual performance of the model to complete the training;

[0013] Step 4: Detection of tomato maturity

[0014] Obtain the tomato image data to be detected for tomato maturity, perform preprocessing, and then input it into the YOLOv11-WTConv network model trained in Step 3, and the YOLOv11-WTConv network model outputs the category result.

[0015] Further improvement lies in: In the above Step 1, the enhancement processing includes random rotation and translation, random adjustment of HSV, and random horizontal flipping.

[0016] Further improvement lies in: In the above Step 1, the normalization processing is specifically: Adjust the resolution of all image data in the tomato dataset to 640x640, and the format is JPG.

[0017] Further improvement lies in: In the above Step 1, the classification standard of the category is: Determine by judging the color of the tomato. If the appearance is dark red, it is mature; if the color is light red or green, it is immature.

[0018] A further improvement lies in that: in the second step, the specific execution manner of the C3k2_WTConv module is as follows: perform wavelet transform on the input data to obtain four components, perform small kernel depth convolution transform on the four components to obtain multiple data, connect them along the specified dimension, then perform inverse wavelet transform on them, transmit the result back to the previous layer, add it to the data after convolution, and finally output the data.

[0019] A further improvement lies in that: in the third step, the training parameters of the network model are as follows: optimizer SGD, batch size 20, initial learning rate 0.01, learning rate factor 0.01, weight decay 0.0005, learning rate momentum 0.937, and number of training times 200.

[0020] A further improvement lies in that: in the fourth step, the preprocessing is to uniformly process the pixel values and input sizes of the input images.

[0021] A further improvement lies in that: in the fourth step, the output result is a visual detection result.

[0022] The beneficial effects of the present invention are as follows: The present invention designs the C3K2_WTConv module for the tomato maturity detection scenario in the agricultural automation scenario, and proposes the YOLOv11-WT Conv algorithm. Through the innovative wavelet transform convolution structure, the ability to extract multi-scale features of the target is enhanced, while improving the mAP50 index, the detection accuracy is also improved. It not only significantly improves the detection accuracy and robustness of tomato maturity grading, but also provides key technical support for the intelligent production management of facility agriculture.

[0023] This model integrates the ability of wavelet transform convolution to extract frequency domain information, combines the attention mechanism, and effectively overcomes the target overlap and occlusion interference in the complex field environment. In the application scenario of agricultural automation, it provides a reliable visual perception module for intelligent picking robots and precise water and fertilizer regulation systems, helps to build an "awareness - decision - execution" integrated digital agriculture system, and promotes the transformation and upgrading of traditional agriculture towards standardization, intensification, and sustainability. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is the schematic diagram of the step flow of the present invention.

[0025] Figure 2 is the schematic diagram of the detection effects in different scenarios of the present invention.

[0026] Figure 3 is the schematic diagram of the comparison of mAP50 changing with the number of iterations of the present invention.

[0027] Figure 4 is the schematic diagram of the comparison of Precision changing with the number of iterations of the present invention.

[0028] Figure 5 It is a schematic diagram of the YOLOv11 network structure.

[0029] Figure 6 It is a schematic diagram of the YOLOv11-WTConv network structure of the present invention.

[0030] Figure 7 It is a schematic diagram of the wavelet transform convolution process of the present invention. Detailed implementation manners

[0031] To deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation to the protection scope of the present invention.

[0032] Embodiment 1. According to Figure 1 As shown, this embodiment proposes a tomato maturity detection method based on improved yolo11, including the following steps:

[0033] Step 1: Construct a tomato image dataset

[0034] Based on the publicly available tomato dataset and the tomato image data obtained by shooting from a tomato agricultural base, and perform enhancement processing and normalization processing on it, so as to construct a tomato dataset. Then use the open-source software LabelImg for image annotation and category division. The enhancement processing includes random rotation and translation, random adjustment of HSV, and random horizontal flipping. The normalization processing is specifically as follows: the resolution of all image data in the tomato dataset is adjusted to 640x640, the format is JPG, and the category division standard is: judge by determining the color of the tomato. Those with a dark red appearance are mature, and those with a light red or green color are immature;

[0035] Step 2: Construct a maturity detection model

[0036] Use YOLOv11 as the basic model, and replace the C3k2 module of YOLOv11 with the C3k2_WTConv module to obtain the YOLOv11-WTConv network model, where the C3k2_WTConv module applies wavelet transform on the basis of the standard convolution;

[0037] The specific execution method of the C3k2_WTConv module is: perform wavelet transform on the input data to obtain four components, perform small kernel depth convolution transform on the four components to obtain multiple data, connect them along the specified dimension, then perform inverse wavelet transform on it, send the result back to the upper layer, add it to the data after convolution, and finally output the data;

[0038] Step 3: Training of the maturity detection model

[0039] The tomato dataset constructed in Step 1 is divided into a training set, a validation set, and a test set. The YOLOv11-WTConv network model constructed in Step 2 is trained using the training set. Then, the performance of the model is evaluated using the validation set. Finally, the generalization ability and actual performance of the model are evaluated using the test set to complete the training. The training parameters of the network model are as follows: optimizer SGD, batch size 20, initial learning rate 0.01, learning rate factor 0.01, weight decay 0.0005, learning rate momentum 0.937, and number of training epochs 200;

[0040] Step 4. Tomato maturity detection

[0041] Obtain tomato image data for detecting the maturity of tomatoes to be detected, and perform preprocessing. The preprocessing method is to unify the pixel values and input sizes of the input images, and then input them into the trained YOLOv11-WTConv network model in Step 3. The YOLOv11-WTConv network model outputs category results, and the output results are visualized detection results.

[0042] Example 2, according to Figures 1 - 7 As shown, this example proposes a tomato maturity detection method based on improved YOLO11, including the following steps:

[0043] Step 1. Construct a tomato image dataset

[0044] In this example, a free tomato dataset provided by PaddlePaddle and tomato images taken at a certain agricultural base in Lujiang County, Anhui Province are used, and they are enhanced to construct a tomato dataset. This dataset provides tomato images under different environments, lighting, and occlusion conditions, which helps train the model to handle diverse actual scenarios.

[0045] Specifically, the tomato image data was enhanced by random rotation and translation, random adjustment of HSV, and random horizontal flipping, resulting in a total of 6,111 tomato images covering different maturity states (such as mature and immature). Random rotation and translation include random rotation and random translation. Random rotation randomly rotates the image by an angle between -30° and 30° to simulate scenes with different shooting angles and improve the model's adaptability to different perspectives in the actual environment. Random translation translates the image randomly within the range of -50 to 50 pixels to simulate changes in the camera shooting position and enhance the model's robustness at different camera positions. Random adjustment of HSV (hue, saturation, value) changes the color characteristics of the image by randomly adjusting the hue (H), saturation (S), and value (V) of the image. This can simulate images under different lighting conditions and help the model identify maturity in different lighting environments. Random horizontal flipping randomly flips the image horizontally to enhance the model's ability to identify tomatoes in different directions, which helps improve the model's generalization ability, especially when encountering new environments or directions.

[0046] Next, the tomato images in the tomato dataset were normalized. That is, the resolution of all image data was 640x640, and the format was JPG. Uniform size facilitated model training. Then, the open-source software LabelImg was used for image annotation. The open-source software LabelImg allowed us to draw bounding boxes for the tomato regions in the image and assign labels to each bounding box to indicate the maturity of the tomatoes. Subsequently, the classification criteria were as follows: by determining the color of the tomatoes. Tomatoes with a dark red appearance were considered mature, while those with a light red or green color were considered immature.

[0047] Step 2: Build a maturity detection model

[0048] YOLOv11 was used as the base model. YOLOv11 is an excellent object detection model. It is based on the YOLO series architecture and uses a convolutional neural network (CNN) for feature extraction, featuring high efficiency and speed. The core structure of YOLOv11 includes the Backbone layer, the Neck layer, and the Head layer. As Figure 5 shown, the Backbone layer is the feature extraction part, usually using a lightweight network such as CSPDarknet. The Neck layer is used to fuse features of different scales, usually using the FPN (Feature Pyramid Networks) structure. The Head layer is the final detection part, responsible for generating the location, category, and confidence of the target. Thus, the C3k2_WTConv module was used to replace the C3k2 module of YOLOv11 to obtain the YOLOv11-WTConv network model, with the structure as Figure 6As shown, the C3k2 module in YOLOv11 consists of 3 convolutional layers and 2 convolutional kernels (C3 represents 3 convolutional layers, and k2 represents a convolutional kernel size of 2). It is mainly used to extract features in the network and enhance the network's expressive ability. In this embodiment, a wavelet transform convolution (WTConv) module is introduced. Utilizing the characteristics of wavelet transform, it can efficiently extract features from images in the frequency domain. By enhancing the convolutional layer, it can effectively capture low-frequency information and detailed features, and still operate stably in an environment with high noise. It is particularly suitable for tomato maturity detection in the farmland environment.

[0049] The specific execution method of the C3k2_WTConv module is as Figure 7 shown, which is:

[0050] Perform wavelet transform on the input data to obtain four components, then perform small kernel depth convolution transform on the four components to obtain multiple data, connect them along the specified dimension, then perform inverse wavelet transform on it, send the result back to the previous layer, add it to the data after convolution, and finally output the data.

[0051] Introducing WTConv in the convolutional neural network has two main advantages. First, for each additional layer of wavelet transform, the receptive field size of that layer will increase exponentially, but the number of trainable parameters will only increase slightly. The second advantage is that the WTConv layer can better capture low-frequency information compared to the standard convolutional layer. This is because repeated wavelet decomposition of the low-frequency part of the input will highlight these low-frequency parts and enhance the response of the corresponding layer. Through a series of calculations and operations, the image frequency domain information is introduced, and WTConv solves the problem of insufficient multi-scale feature extraction in the complex farmland environment, and at the same time improves the robustness of object detection under frequency domain noise interference.

[0052] Step 3: Training of the maturity detection model

[0053] Divide the tomato dataset constructed in Step 1 into a training set, a validation set, and a test set in the ratio of 8:1:1. Use the training set to train the YOLOv11-WTConv network model constructed in Step 2. Then use the validation set to evaluate the model performance and perform hyperparameter tuning. Finally, use the test set to evaluate the generalization ability and actual performance of the model. Specifically, match the image data in the training set with the label files one by one to ensure the data and labels match and avoid label errors. Then, in each training cycle, use the SGD optimizer to update the parameters, minimize the loss function, calculate the gradients of each batch of data, and update the weights to gradually improve the model's ability to recognize tomato maturity. Use the validation set to evaluate the model after each epoch, monitor metrics such as validation loss, accuracy, and recall to avoid overfitting. If the accuracy of the model on the training set gradually increases while the accuracy on the validation set stops improving or decreases, there may be an overfitting problem, and hyperparameters need to be adjusted, such as adjusting the learning rate to avoid unstable model training caused by too large a learning rate.

[0054] Furthermore, the parameter information of the training process is shown in Table 1 below:

[0055] Table 1

[0056]

[0057]

[0058] After training is completed, use the test set for the final model evaluation. The evaluation metrics include mean average precision (mAP), Precision (P), recall (R), and the number of parameters (Params). After training is completed, save the best model weights for subsequent application and deployment.

[0059] Step 4: Recognition of tomato categories

[0060] Obtain the tomato image data for detecting the maturity of the tomato to be detected, perform preprocessing, and then input it into the trained YOLOv11-WTConv network model in Step 3. The YOLOv11-WTConv network model outputs the category results, and the output results are visualized detection results, as shown in Figure 2 shown, Figure 2 shown. From top to bottom are the original image, YOLOv11, and YOLOv11-WTConv; from left to right are multi-classification, occlusion, and complex environment.

[0061] To accurately evaluate the performance of the model proposed in this invention, the research adopted a number of evaluation metrics, including mean average precision (mAP), Precision (P), Recall (R), and the number of parameters (Params). The specific explanations of the parameters are as follows:

[0062] (1) Mean Average Precision (mAP) is one of the commonly used metrics for evaluating object detection models. The mAP metric in this invention is mAP50. mAP50 is the average precision of the model when the Intersection over Union (IoU) threshold is 0.5. In object detection, the higher the mAP value, the better the detection performance of the model, as shown in the following formula:

[0063]

[0064] In the formula, AP represents the average precision of a single class, and N represents the number of classes of detection targets in the dataset.

[0065] (2) Precision (P) evaluates the probability that the classifier calculates the correct class. In object detection, the higher the P value, the better the classification and detection performance of the model, as shown in the following formula:

[0066]

[0067] In the formula, TP represents the true positive, that is, the prediction result matches the actual result, and FP represents the false positive, that is, the prediction result is true while the actual result is false.

[0068] (3) Recall (R) is the proportion of correctly predicted positive examples among all actual positive examples. The larger the R value, the fewer detection targets the model misses. As shown in the following formula:

[0069]

[0070] In the formula, FN is the false negative, that is, the prediction is negative while it is actually positive.

[0071] (4) The number of parameters (Params) represents the total number of parameters that need to be trained during model training.

[0072] To verify the superiority of this invention in the detection effect of tomato maturity, to ensure the rigor of the verification experiment, the dataset and training parameters are kept consistent, and the experimental effects of several detection networks, YOLOv5, YO-LOv8, YOLOv11, and the YOLOv11-WTConv model proposed in this invention, are compared. The comparison of experimental effects is shown in Table 2 below:

[0073] Table 2 Detection results of tomato maturity on the dataset for different models

[0074]

[0075] As can be seen from Table 2, in terms of mAP50, the YOLOv11-WTConv model significantly outperforms the original YOLOv11 model. The mAP50 of YOLOv11-WTConv reaches 97.8%, while the mAP50 of YOLOv11 is 93.6%. Compared with YOLOv11, it has increased by 4.2 percentage points. In terms of accuracy (P), YOLOv11-WTConv has increased by 5.5 percentage points compared with the original model, reaching 94.6%. In addition, the YOLOv11-WTConv model also shows good performance in recall rate (R). Although the number of parameters of YOLOv11-WTConv is slightly higher than that of the original model, considering the significant improvement in mAP50 and P, the increase in the number of parameters is also acceptable, especially for the tomato maturity detection task with extremely high requirements for detection accuracy.

[0076] The intuitive detection effects of the original model and the YOLOv11-WTConv model on maturity in multi-class, occluded, and complex environments are as Figure 2 shown. Due to the irregular growth of tomatoes, they are often occluded by other tomatoes and branches and leaves, which in turn interferes with the judgment of maturity. In the case of occlusion, the model of the present invention detects more targets than the original model. In complex environments, it can better identify fruit targets than the original model. It can be concluded that the model proposed by the present invention has better adaptability and robustness than the original model in complex environments.

[0077] To more intuitively evaluate the detection effect of the improved model, as Figure 3 and Figure 4 shown, the indicators of mAP50 and P of each model during the training process on the dataset are presented. Figure 3 For the data of mAP50 during the training process, it can be seen from the figure that the curve of the YOLOv11-WTConv model quickly converges to over 90% after 50 rounds of training and finally stabilizes after about 150 rounds of training, without obvious overfitting phenomenon, indicating that introducing the WTConv module will enhance multi-scale feature extraction and suppress noise interference, making the model more robust. Figure 4 For the data of P value during the training process, it can be seen from the figure that the YOLOv11-WTConv model quickly converges to over 85% after 50 rounds of training, and then the training steadily increases. Combining Figure 3 while the mAP50 value increases, the P value also increases simultaneously, and there is no problem of imbalance with high mAP and low P.

[0078] In summary, the YOLOv11-WTConv proposed by the present invention achieves significant improvements in both mAP50 and precision by introducing the wavelet transform convolution module, while maintaining a good convergence speed and stability.

[0079] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the framework and scope of application of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A tomato maturity detection method based on improved yolo11, characterized in that: The following steps are involved: Step 1: Build a tomato image dataset Based on the publicly available tomato dataset and tomato image data obtained from tomato agricultural bases, the tomato dataset was constructed by enhancing and normalizing the tomato image data. The open source software LabelImg was then used to annotate and classify the images. Step 2: Build a maturity detection model Using YOLOv11 as the base model, the C3k2 module of YOLOv11 is replaced by the C3k2_WTConv module to obtain the YOLOv11-WTConv network model, where the C3k2_WTConv module applies wavelet transform on the basis of standard convolution; Step 3: Training of maturity detection model Divide the tomato dataset constructed in step 1 into training set, validation set and test set. Use the training set to train the YOLOv11-WTConv network model constructed in step 2. Then use the validation set to evaluate the model performance. Finally, use the test set to evaluate the generalization ability and actual performance of the model to complete the training. Step 4: Tomato maturity test Obtain tomato image data of the tomato maturity to be detected, perform preprocessing, and then input it into the YOLOv11-WTConv network model trained in step 3, and the YOLOv11-WTConv network model outputs the category result.

2. A tomato maturity detection method based on improved yolo11 according to claim 1, characterized in that: In the step 1, the enhancement processing includes random rotation and translation, random adjustment of HSV and random horizontal flipping.

3. A tomato maturity detection method based on improved yolo11 according to claim 1, characterized in that: In the step 1, the normalization process is specifically as follows: the resolution of all image data in the tomato dataset is adjusted to 640x640, and the format is JPG.

4. A tomato maturity detection method based on improved yolo11 according to claim 1, characterized in that: In the step 1, the classification standard is: judging by the color of the tomato, dark red is ripe, and light red or green is unripe.

5. A tomato maturity detection method based on improved yolo11 according to claim 1, characterized in that: In the step 2, the specific execution method of the C3k2_WTConv module is: perform wavelet transform on the input data to obtain four components, perform small-kernel deep convolution transform on the four components to obtain multiple data, connect them along the specified dimension, perform inverse wavelet transform on them, pass the result back to the previous layer, add it to the convolved data, and finally output the data.

6. A tomato maturity detection method based on improved yolo11 according to claim 1, characterized in that: In step 3, the network model training parameters are: optimizer SGD, batch size 20, initial learning rate 0.01, learning rate factor 0.01, weight decay 0.0005, learning rate momentum 0.937, and training times 200.

7. The method for detecting tomato maturity based on improved YOLOL11 according to claim 1, characterized in that: In the step 4, the preprocessing specifically involves unifying the pixel values ​​and input size of the input image.

8. The method for detecting tomato maturity based on improved YOLOL11 according to claim 1, characterized in that: In the step 4, the output result is a visual detection result.

Citation Information

Cited By

  • Tomato image real-time detection method, tomato image real-time detection system, tomato picking method and tomato picking system

    CN120388367A

  • Real-time tomato image detection method and system, tomato picking method and system

    CN120388367B

  • Tomato fruit maturity detection method based on YOLOv11

    CN120472453A

  • Tomato fruit maturity detection method based on YOLOv11

    CN120472453B

  • Dual-module cooperative corn maturity detection method based on domain self-adaption

    CN121147912A