Improved YOLOV8-based strawberry picking maturity detection method and system

By improving the backbone network of the YOLOV8 model to EfficientViT, introducing SPPFCSPC and AlphaCIOU loss functions, combined with data enhancement technology, the problem of insufficient accuracy and robustness of strawberry ripening detection is solved, and efficient strawberry ripening detection is achieved in complex scenarios.

CN120260034APending Publication Date: 2025-07-04SOUTH CHINA AGRICULTURAL UNIVERSITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510338617.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing strawberry ripening detection methods have insufficient accuracy and robustness when dealing with strawberry ripening detection tasks with small size, multiple sizes and height variations, especially the traditional YOLOV8 model does not pay enough attention to long-distance dependence and different scale features in the image.

Method used

The improved YOLOV8 model is adopted, and the spatial pyramid pooling module SPPFCSPC is introduced by replacing the backbone network with the EfficientViT module, and the AlphaCIOU loss function is used, combined with data enhancement technology, to improve the detection capability of the model.

Benefits of technology

It realizes accurate identification of strawberry ripening in open-air farmland planting and greenhouses, reduces the false detection rate, improves detection accuracy and model robustness, and adapts to detection under different lighting, angles and background conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260034A_ABST
    Figure CN120260034A_ABST
Patent Text Reader

Abstract

The invention relates to an improved YOLOv8-based strawberry picking maturity detection method and system. The method comprises the following steps: acquiring an image data sample of strawberries in a greenhouse; performing maturity detection on the to-be-detected strawberry image based on an improved YOLOv8 model to obtain a strawberry maturity detection result; wherein the improved YOLOv8 model completes training through a training set and a verification set, and data of the training set and the verification set are strawberry image sets marked with maturity. According to the method, an Efficient ViT backbone network in an improved YOLOv8 model is used for respectively carrying out attention calculation in different groups, so that feature information of finer granularity of strawberries can be captured; meanwhile, a spatial pyramid pooling module in the model can extract and fuse features of different scales, and target detection and positioning are carried out by using the multi-scale features in the detection process. According to the method, the accuracy of identifying the strawberry picking maturity by the target detection model in an actual scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision, deep learning, and strawberry maturity detection, and particularly relates to a strawberry picking maturity detection method and system based on improved YOLOV8. Background Art

[0002] Currently, common strawberry maturity detection methods include those based on machine learning and those based on deep learning. Machine learning-based methods mainly rely on manually designed feature extraction algorithms and classifiers, but they have limited effectiveness for complex data, such as different maturity stages of strawberry images and varying lighting conditions, and the accuracy of machine learning models is relatively low. Deep learning-based methods widely use convolutional neural networks (CNNs), among which the YOLO series is a common object detection algorithm. However, there are still some challenges and deficiencies in dealing with strawberry maturity detection tasks with small sizes, multiple sizes, and highly variable heights.

[0003] Traditional YOLOV8 does not introduce a self-attention mechanism, and its attention to long-range dependencies in images and important features of images at different scales may not be sufficient. In strawberry maturity detection, specific maturity features are crucial for classification and localization, and special attention is needed to improve accuracy and robustness.

[0004] Prior Art 1 provides a method for quickly discriminating strawberry maturity based on machine learning: the strawberry picking area is divided into immature areas and high-maturity areas through a machine learning algorithm to reduce the processing volume; then local entropy analysis combined with region segmentation and clustering methods are used to evaluate the lighting uniformity, and the directionality of the moisture path is analyzed through the histogram of oriented gradients technology to evaluate the degree of sufficient light reception and the degree of moisture accumulation. Based on these two features, the predicted strawberry maturity is dynamically generated, and combined with regional correlation analysis, the picking order is automatically optimized to achieve the judgment of strawberry maturity.

[0005] Although Prior Art 1 can make discriminations from multiple dimensions such as light distribution and moisture accumulation, its adaptability to complex scenarios and the tracking of small and multiple targets are limited.

[0006] The prior art 2 provides a method for detecting the maturity of passion fruits based on YOLOV8: collecting images of passion fruits during the growth period, establishing an image dataset of passion fruits during the growth period; preprocessing the images of passion fruits during the growth period, calibrating the images of each stage of the passion fruit growth period and establishing a training set and a test set; constructing a deep learning network model for training, with the YOLOV8 framework; setting model training parameters, image processing parameters, classification recognition and detection parameters, etc.; using the adjusted object detection and recognition model for each stage of the passion fruit growth period to train the training set of images of each stage of the passion fruit growth period; using the test set to test the trained model, and reasonably outputting the model for the training results. The prior art 2 does not improve the algorithm but only applies the YOLOV8 algorithm, and the detection results are inaccurate. Summary of the Invention

[0007] To solve the above technical problems, the present invention proposes a method and system for detecting the picking maturity of strawberries based on improved YOLOV8, which can accurately identify and detect the maturity information of strawberries in images.

[0008] To achieve the above object, the present invention provides a method for detecting the picking maturity of strawberries based on improved YOLOV8, including:

[0009] Collecting images of strawberries to be detected;

[0010] Based on the improved YOLOV8 model, performing maturity recognition on the images of strawberries to be detected to obtain the strawberry maturity detection result; wherein, the improved YOLOV8 model is trained through a training set and a validation set, and the data of the training set and the validation set are strawberry image sets with marked maturity.

[0011] Optionally, before annotating the data of the training set, it further includes: collecting original strawberry image data, and preprocessing the original strawberry image data including:

[0012] Performing denoising processing on the original strawberry image, uniformly processing the denoised strawberry image to obtain the strawberry image. Optionally, the maturity annotation of the strawberry image set includes:

[0013] Mature, semi-mature and immature.

[0014] Optionally, the improved YOLOV8 model is:

[0015] Replacing the backbone network of YOLOV8 with the EfficientViT module, replacing the spatial pyramid pooling module SPPF with SPPFCSPC, connecting the spatial pyramid pooling module to the output end of the backbone network in the YOLOV8 model, and replacing the loss function CIOU with AlphaCIOU;

[0016] Based on the EfficentViT module, the spatial pyramid pooling module, and the AlphaCIOU loss function, an improved YOLOV8 model is obtained.

[0017] The EfficientViT module includes: a sandwich layout structure and a cascaded group attention sub-module;

[0018] The cascaded group attention sub-module is the cascaded group attention layer in the sandwich layout structure.

[0019] The sandwich layout structure includes:

[0020] Input the input features into N feed-forward network layers, then into a single cascaded group attention layer, and finally into N feed-forward network layers to obtain the output features;

[0021] The cascaded group attention sub-module includes:

[0022] Split the input features into j attention heads;

[0023] Calculate the attention through the output query vector, key vector, and value vector matrices;

[0024] Concatenate all the attention heads and multiply by the weight matrix to obtain the output features.

[0025] Optionally, the data augmentation process includes: performing operations such as scaling, cropping, rotating, translating, flipping, adjusting contrast, and brightness on the original strawberry image data.

[0026] Optionally, the spatial pyramid pooling layer includes:

[0027] Input the features of the backbone network EfficientViT module into the spatial pyramid pooling module, and perform pooling processing on the input features through the spatial pyramid pooling module to obtain scale features;

[0028] Combine the scale features to obtain multi-scale features.

[0029] Optionally, the loss function AlphaCIOU includes:

[0030] The weight calculation is based on the loss function AlphaCIOU of the model. During the training process, the model parameters are adjusted through the backpropagation algorithm to minimize the difference between the predicted value and the actual value.

[0031] The present invention also discloses a strawberry picking maturity detection system based on the improved YOLOV8, including:

[0032] An acquisition unit for collecting strawberry images to be detected;

[0033] The detection unit is used to identify the maturity of the strawberry image to be detected based on the improved YOLOV8 model, and obtain the strawberry maturity detection result; wherein, the improved YOLOV8 model is trained through a training set and a validation set, and the training set data and the validation set data are strawberry image sets with marked maturity.

[0034] Compared with the prior art, the present invention has the following advantages and technical effects:

[0035] (1) The present invention proposes a strawberry maturity detection method and application based on the improved YOLOV8, which can accurately identify and detect the maturity information of strawberries in the image, detect and locate the picking of strawberry fruits in open-field planting and greenhouse planting, and realize improving the accuracy of the target detection model in identifying the picking maturity of strawberries in the actual scenario, while reducing the false detection rate of the model.

[0036] (2) The EfficientViT module is used in the present invention. Through its hierarchical structure, EfficientViT can capture multi-scale features in the image. In the cascaded group attention module of EfficientViT, by calculating attention separately within different groups, the model can capture more fine-grained feature information of strawberries; through the cascaded method, the model can first capture the fine-grained dependence relationship of strawberries within a local range, and then summarize it within a more global range, which helps to better model long-range dependence. By introducing the EfficientViT module, the YOLOV8 model can more accurately identify the features related to strawberry maturity and improve the detection accuracy.

[0037] (3) The data augmentation module used in the present invention can increase the diversity of data by performing various random transformations and augmentations on the training data of three classifications. By randomly changing the size, rotation angle, contrast, etc. of the image during the training process, the model can have better robustness and transferability. The data augmentation module can help the YOLOV8 model better adapt to the maturity detection scenarios of strawberries under different lighting, angles, backgrounds and occlusion conditions, and improve the robustness and accuracy of the model.

[0038] (4) The spatial pyramid pooling module used in the present invention can extract and fuse features of different scales, and use multi-scale features for target detection and positioning during the detection process. By introducing the spatial pyramid pooling module, the YOLOV8 model can better process strawberry maturity targets of different sizes and proportions, and improve the accuracy and effect of target detection. Description of the Drawings

[0039] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0040] Figure 1 is a schematic flowchart of a strawberry picking maturity detection method based on improved YOLOV8 according to an embodiment of the present invention;

[0041] Figure 2 is a test schematic diagram of a strawberry picking maturity detection method based on improved YOLOV8 of the present invention;

[0042] Figure 3 is a schematic structural diagram of the EfficientViT module according to an embodiment of the present invention;

[0043] Figure 4 is a schematic structural diagram of the spatial pyramid pooling module SPPFCSPC according to an embodiment of the present invention;

[0044] Figure 5 is a schematic structural diagram of the improved YOLOV8 model according to an embodiment of the present invention. Detailed implementation manners

[0045] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will refer to the accompanying drawings and combine the embodiments to detail this application.

[0046] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0047] The present invention provides a strawberry picking maturity detection method based on improved YOLOV8, and the specific steps are as Figure 1 shown, including: S1. Collect image data samples of strawberries in the shed and make a data set: Collect strawberry images, after preprocessing the images, use the cosine similarity algorithm, with a similarity threshold of 0.95 (the maximum value is 1), to filter out similar images to eliminate data redundancy and enhance the generalization of the model. Specifically, use the open() function of the PIL library to read the pictures, and then use the cosine similarity algorithm where A and B respectively represent the n-dimensional vectors of image A and image B, calculate the similarity of two pictures, and evaluate the similarity of the images by comparing the cosine similarity of the images. Finally, remove the images with a similarity greater than the threshold of 0.95.

[0048] Use the labelimg annotation software to annotate the filtered images, obtaining a dataset containing three types of labels: mature strawberries, semi-mature strawberries, and immature strawberries.

[0049] S2. Divide the dataset: Divide the dataset into a training set, a validation set, and a test set in a ratio of 8:1:1. Before dividing the data, all data needs to be converted into the jpg format, and the data annotations need to be converted into the txt format.

[0050] S3. Improvements to the YOLOV8 model include: replacing the backbone network part with the EfficientViT module, adding an online data augmentation module, improving the spatial pyramid pooling module to SPPFCSPC, and improving the loss function to FoucsCIOU. The schematic diagram of its structure is as Figure 5 shown.

[0051] Replace the backbone network part with the EfficientViT module, as Figure 3 shown: The EfficientViT module consists of a memory-efficient sandwich layout structure and a cascaded group attention sub-module, which focus on improving the efficiency of the model in terms of memory and computation respectively. The EfficientViT module enhances the model's ability by simultaneously considering the relationship between memory-efficient and attention-head-efficient computations, applying an additional token interaction layer, using depthwise convolution DWConv, and introducing the inductive bias of local structural information. It can replace the backbone network of YOLOV8. The structure of the EfficientViT module is a sandwich layout, using fewer memory-constrained self-attention layers and more memory-efficient feed-forward network (FFN) layers for channel communication. Specifically, it applies a self-attention layer for spatial mixing, located between the FFN layers. The self-attention layer among them is the cascaded group attention sub-module, which splits different input feeds to each head and cascades the output features to the heads. This module not only reduces the computational redundancy in multi-head attention but also improves the model capacity by increasing the network depth.

[0052] Calculation formula for the sandwich layout structure:

[0053]

[0054] where X i is the complete input feature of the i-th block. Before and after a single self-attention layer, the block converts X i into X i+1 . This design reduces the memory and time consumption caused by the self-attention layer in the model and applies more FFN layers to effectively achieve communication between different feature channels.

[0055] Cascaded Group Attention Sub - module Calculation Formula:

[0056]

[0057] Among them, the j - th head calculates the self - attention on the input feature split X ij That is, X i =[X i1 , X i2 , …, X ih , where 1 ≤ j ≤ h. Here, h is the total number of heads. and are projection layers that map the input feature split to different sub - spaces. is a linear layer that projects the concatenated output features back to the same dimension as the input.

[0058] X i ′ j =X ij +X i(j-1) , 1 ≤ j ≤ h

[0059] Among them, X i ′ j is the sum of the j - th input split X ij and the output of the (j - 1) - th head . When calculating self - attention, it replaces X ij as the new input feature of the j - th head. In addition, a token interaction layer is applied after the Q projection, which enables self - attention to jointly capture local and global relationships, further enhancing the feature representation.

[0060] The EfficientViT module has the advantages of memory - efficient operation and cascaded group attention. This architecture includes three stages. Each stage stacks EfficientViT building blocks, as Figure 5 shown, enhancing the model's ability in multi - level visual representation learning. By replacing the backbone network of YOLOV8 with the EfficientViT module, the model's ability to detect strawberry maturity can be enhanced.

[0061] The design of the online data augmentation module includes: Online data augmentation refers to performing real - time data augmentation operations on each training sample during model training. Data augmentation can change the attributes of training samples such as size, translation, scaling, rotation, etc., thereby increasing the diversity of training samples. By increasing the number and diversity of training samples, the model can better learn the features and patterns of the data, improving the robustness and generalization ability of the model.

[0062] Data augmentation operations include:

[0063] Random translation: Translate the image randomly by a certain distance to simulate the maturity detection at different positions.

[0064] Random scaling: Randomly scale the size of the image to simulate the maturity detection at different distances.

[0065] Random flipping: Randomly flip the image horizontally or vertically to simulate the maturity detection of mirror images and asymmetric shapes.

[0066] Random rotation: Randomly rotate the image by a certain angle to simulate the maturity detection at different angles.

[0067] Random cropping: Randomly crop a part of the image to simulate the maturity detection in different regions.

[0068] Contrast and brightness adjustment: Randomly adjust the contrast and brightness of the image to simulate the maturity detection under different lighting conditions. During the training process, data augmentation is applied. In each training batch, an online data augmentation module is used to perform a series of random augmentation operations on the input image. This can use different augmented samples during each training, increase the diversity of the data, and improve the accuracy of the model for strawberry maturity detection under different lighting, angle, background, and occlusion conditions.

[0069] Add an improved Spatial Pyramid Pooling Fast with Context (SPPFCSPC) module after the EfficientViT module in the backbone network to extract multi-scale features.

[0070] As Figure 4 shown: SPPFCSPC decomposes the input feature map into grids of different sizes, and each grid corresponds to a set of different pooling ratios. Then, spatial pyramid pooling operations are performed on each pooling ratio to aggregate features with different receptive field sizes. In addition, SPPFCSPC also uses a context feature extraction mechanism to improve the detection accuracy by introducing global context information.

[0071] 1. Modify the network structure: Modify the network structure of YOLOv8. In the network structure of YOLOv8, find the end of the EfficientViT module in the backbone network, which is the part responsible for feature extraction. Generally, the extraction layer is located at an earlier stage of the entire network. According to the network structure of YOLOv8, the size of the features in the extraction layer is usually large and contains rich semantic information.

[0072] 2. Replace the original Spatial Pyramid Pooling module with SPPFCSPC: At the end of the EfficientViT module, replace the Spatial Pyramid Pooling module. This module can extract multi-scale features of the input features. Specifically, perform pyramid pooling on the input feature map and concatenate the feature maps of different scales together. In this way, the features of the target at different scales can be extracted. In addition, perform cross-stage partial connection on the pooled feature map, concatenate the feature map of the previous layer and the feature map of the next layer, halve the number of channels of the feature map, and improve the efficiency and accuracy of the network. Alleviate the problems of gradient disappearance and information bottleneck, and enhance the expression ability and generalization ability of the network.

[0073] S4. Train the improved YOLOv8 model and optimize it: Build the improved YOLOv8 network and use the training set to train the improved YOLOv8 model to obtain the strawberry maturity detection model; at the same time, use cross-validation to evaluate the model performance using the validation set after each training cycle, and adjust the model parameters according to the evaluation results to optimize the model performance.

[0074] During the training process, IOU_Loss is the ratio of the intersection between the predicted box and the ground truth box to the union of the predicted box and the ground truth box. There is a problem that non-intersecting boxes cannot be optimized. To solve this problem, GIOU_Loss is introduced, but it degrades to IOU when the boxes intersect and cannot distinguish the relative positions. Further improvement leads to DIOU_Loss, which considers the overlapping area and the distance between the center points, improving the convergence speed. CIOU_Loss adds the consideration of aspect ratio on the basis of DIOU. However, CIOU_Loss is not ideal when dealing with targets of different scales. AlphaCIOU_Loss designs a scale invariance on the basis of CIOU_Loss, which means it is more fair for bounding boxes of different sizes, is not affected by the target size, and significantly improves the mean average precision value. Therefore, this paper uses AlphaCIOU_Loss to calculate the model loss, and the specific calculation method is shown in the following formula:

[0075]

[0076] Among them, AlphaCIOU loss represents the Complete Intersection over Union Loss, CIOU represents the Complete Intersection over Union, ρ represents the Euclidean distance between two rectangular boxes, the role of c is to prevent the value of the loss function from being too large and improve the convergence speed, b represents the center point of the predicted box, b gt represents the center point of the ground truth box, and the role of α is to achieve different levels of bounding box regression accuracy, and v is used to measure the consistency of the aspect ratio, which is defined as follows:

[0077]

[0078] Among them, wgt represents the width of the ground truth box, h gt represents the height of the ground truth box, w represents the width of the predicted box, h represents the height of the predicted box, π is the pi, and β represents the weight coefficient, which are defined as follows:

[0079]

[0080] where IOU represents the intersection over union.

[0081] S5. Real-time detection of strawberry maturity: Deploy the model on the hardware, collect the real-time strawberry video stream, use the improved YOLOv8 algorithm to detect the strawberry maturity in real time and return the result to the front end. The detection result is as Figure 2 shown.

[0082] Example 2

[0083] The present invention also provides a strawberry maturity detection system based on the improved YOLOv8, including:

[0084] An acquisition unit, configured to collect the strawberry image to be detected;

[0085] A detection unit, configured to identify the maturity of the strawberry image to be detected based on the improved YOLOv8 model, and obtain the strawberry maturity detection result; wherein, the improved YOLOv8 model is trained through the training set and the validation set, and the data of the training set and the validation set are the strawberry image sets with labeled maturity.

[0086] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A strawberry picking maturity detection method based on improved YOLOv8, characterized in that, Including: Collecting strawberry images to be detected; Based on the improved YOLOv8 model, performing maturity recognition on the strawberry images to be detected to obtain strawberry maturity detection results; Among them, the improved YOLOv8 model is trained through a training set and a validation set, and the data of the training set and the validation set are strawberry image sets with labeled maturity.

2. The strawberry picking maturity detection method based on improved YOLOv8 according to claim 1, characterized in that Before labeling the training set data, it also includes: collecting original strawberry image data, and preprocessing the original strawberry image data, including: Performing denoising processing on the original strawberry images, and uniformly processing the denoised strawberry images to obtain strawberry image data.

3. A strawberry picking maturity detection method based on improved YOLOv8 according to claim 1, characterized in that The maturity labeling of the strawberry image set includes: Mature, semi-mature, and immature.

4. A strawberry picking maturity detection method based on improved YOLOv8 according to claim 1, characterized in that, The improved YOLOv8 model is: Replacing the backbone network of YOLOV8 with the EfficientViT module, replacing the Spatial Pyramid Pooling Module SPPF with SPPFCSPC, connecting the Spatial Pyramid Pooling Module to the output end of the backbone network EfficientViT in the YOLOv8 model, and replacing the loss function CIOU with AlphaCIOU; Based on the EfficientViT module, the Spatial Pyramid Pooling Module, and the AlphaCIOU loss function, an improved YOLOV8 model is obtained.

5. The strawberry picking maturity detection method based on improved YOLOv8 according to claim 4, wherein The EfficientViT module includes: a sandwich layout structure and a cascaded group attention sub-module; The cascaded group attention sub-module is the cascaded group attention layer in the sandwich layout structure.

6. The strawberry picking maturity detection method based on improved YOLOv8 according to claim 5, wherein, The sandwich layout structure includes: Inputting the input features into N feed-forward network layers, then into a single cascaded group attention layer, and finally into N feed-forward network layers to obtain output features; The cascaded group attention sub-module includes: Dividing the input features into j attention heads; Calculating attention through the output query vector, key vector, and value vector matrices; Concatenating all attention heads, multiplying by the weight matrix, and obtaining output features.

7. A strawberry picking maturity detection method based on improved YOLOv8 according to claim 2, characterized in that, The process of data augmentation includes: performing processing such as scaling, cropping, rotating, translating, flipping, adjusting contrast, and brightness on the original strawberry image data images.

8. A strawberry picking maturity detection method based on improved YOLOv8 according to claim 4, characterized in that The Spatial Pyramid Pooling layer includes: Inputting the features of the backbone network EfficientViT module into the Spatial Pyramid Pooling module, and performing pooling processing on the input features through the Spatial Pyramid Pooling module to obtain scale features; Combining the scale features to obtain multi-scale features.

9. A strawberry picking maturity detection method based on improved YOLOv8 according to claim 4, characterized in that, The loss function AlphaCIOU includes: The weight calculation is based on the loss function AlphaCIOU of the model. During the training process, the model parameters are adjusted through the backpropagation algorithm to minimize the difference between the predicted value and the actual value.

10. A strawberry picking maturity detection system based on improved YOLOv8, characterized in that, Including: An acquisition unit for collecting strawberry images to be detected; A detection unit for performing maturity recognition on the strawberry images to be detected based on the improved YOLOv8 model to obtain strawberry maturity detection results; among them, the improved YOLOv8 model is trained through a training set and a validation set, and the training set data and the validation set data are strawberry image sets with labeled maturity.

Citation Information

Cited By

  • Light-weight strawberry maturity detection method and device based on FDET-YOLOv11n

    CN121600311A