Strawberry maturity detection method based on improved YOLOv8n

By improving the YOLOv8n model and optimizing strawberry maturity detection using C2f-OREPA, EMA, and C2f-DCNv3 modules, the problems of long detection time and low accuracy in strawberry maturity detection are solved, achieving efficient and accurate strawberry maturity detection, which is suitable for low-computing-power devices.

CN121904580APending Publication Date: 2026-04-21CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
Filing Date
2025-12-25
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for detecting strawberry maturity rely on manual judgment or simple physical characteristics, which are time-consuming and prone to errors. Furthermore, existing target detection technologies have low accuracy and poor robustness in complex agricultural scenarios, making it difficult to meet the high-efficiency and precision requirements of modern agriculture.

Method used

An improved YOLOv8n model was adopted, which optimized the strawberry maturity detection model by replacing the C2f-OREPA module, introducing the EMA attention module and the C2f-DCNv3 module, enhancing multi-scale feature extraction and environmental adaptability, and reducing computational complexity and number of parameters.

Benefits of technology

It significantly improves the accuracy and robustness of strawberry maturity detection, adapts to the requirements of lightweight and real-time operation, is suitable for low-computing-power equipment, and enhances the efficiency of strawberry growth management and the performance of intelligent fruit thinning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904580A_ABST
    Figure CN121904580A_ABST
Patent Text Reader

Abstract

The invention discloses a strawberry maturity detection method based on improved YOLOv8n, and the method comprises the steps: photographing and marking strawberry images of different maturity in a real scene, and constructing a data set; a detection model is optimized on the basis of YOLOv8n; a C2f-OREPA module is used for replacing an original C2f module, so that the feature extraction efficiency and the reuse capability are improved; an EMA attention module is embedded in the backbone network, and extraction of important features is enhanced; and the C2f-DCNv3 module is adopted to replace the last two C2f modules, so that the adaptability to targets with different shapes and sizes is improved. According to the method, multi-scale strawberry targets can be efficiently detected in a complex agricultural scene, the problems of leaf shielding, fruit overlapping, different maturity of fruits in the same cluster and the like are solved, the detection precision and reliability are remarkably improved, meanwhile, the model complexity is simplified, and the deployment requirement of mobile terminal equipment is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision target detection technology, and relates to the automatic detection of fruit and vegetable maturity, and in particular to a strawberry maturity detection method based on an improved YOLOv8n. Background Technology

[0002] Traditional methods for detecting strawberry ripeness rely on manual judgment or mechanical detection based on simple physical characteristics such as color and size. These methods are time-consuming, prone to errors, and costly in terms of labor, failing to meet the demands of efficient and precise production in modern agriculture. Furthermore, existing target detection technologies perform poorly in complex agricultural scenarios, exhibiting low accuracy and poor robustness.

[0003] With the development of deep learning algorithms, traditional object detection algorithms are transforming towards higher efficiency, intelligence, and lightweight design. Deep learning-based object detection algorithms include two-stage detection and single-stage detection. Two-stage detection algorithms typically suffer from high computational complexity and a large number of parameters, requiring high-performance hardware, thus limiting their application in resource-constrained agricultural scenarios. In contrast, single-stage detection algorithms offer faster detection speeds, but their localization and recognition accuracy is relatively low, their ability to detect small targets is insufficient, and their real-time performance is not strong enough.

[0004] With the development of artificial intelligence and computer vision technologies, convolutional neural networks (CNNs) have been widely used in object detection, demonstrating good results in strawberry detection. CNNs can effectively extract features such as pose, color, and lighting of objects to adapt to the complex changes in the strawberry growing environment. However, existing research mainly focuses on detecting strawberries at maturity. Although the models used have high detection accuracy, they are usually accompanied by increased model complexity and computational resource consumption, which poses a challenge for practical applications, especially in agricultural scenarios with high requirements for lightweight design and real-time performance.

[0005] In real-world scenarios, strawberry fruit ripening is affected by natural conditions such as leaf shading, overlapping fruits, and uneven ripening within the same cluster. Advanced target detection technology is needed to address this critical issue and cope with complex environmental factors. Modern agriculture therefore urgently requires an intelligent and reliable technological approach that strikes a balance between lightweight design, precision, and efficiency to meet the needs of deploying low-computing-power equipment, while simultaneously improving the detection capabilities for multi-scale targets (especially small targets), thereby enabling automated monitoring of strawberry ripeness and precise harvesting.

[0006] To address this challenge, developing a method for detecting strawberry maturity is crucial. This method should significantly reduce the computational complexity and number of parameters of the model while maintaining detection accuracy, thereby enabling real-time detection of young strawberries. This is of great significance for improving the efficiency of strawberry growth management and the performance of intelligent fruit thinning systems. Summary of the Invention

[0007] To address the aforementioned problems, this invention provides a strawberry maturity detection method based on an improved YOLOv8n. This method aims to efficiently detect multi-scale strawberry targets in complex agricultural environments, particularly addressing practical issues such as leaf shading, fruit overlap, and maturity differences within the same cluster of fruits, significantly improving detection accuracy and reliability.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A method for detecting strawberry maturity based on an improved YOLOv8n includes the following steps:

[0010] Step 1: Capture images of strawberries at different ripeness levels in real-world scenes, and label strawberry ripeness category information and divide the dataset;

[0011] Step 2, construct a detection model, which is a YOLOv8n model with one or more of the following improvements:

[0012] Improvement 1: Replace the original C2f part of YOLOv8n with C2f-OREPA;

[0013] Improvement 2: Introducing the EMA attention module to enhance the robustness of multi-scale feature extraction;

[0014] Improvement 3: C2f-DCNv3 is introduced and added to the backbone network;

[0015] Step 3: Use the training set divided in Step 1 to train the detection model of Step 2, and input the collected strawberry images into the trained detection model to obtain images with strawberry maturity category information bounding boxes.

[0016] In one embodiment, step 1 includes: First, performing size unification and normalization processing on the acquired strawberry images. Then, adding bounding boxes to the strawberry fruits in each image. Next, based on variety and maturity, the acquired images are divided into six categories: bud stage, flowering stage, young fruit stage, fruit enlargement stage, color change stage, and ripening stage. Finally, the images are divided into training, validation, and test sets, and data augmentation operations are performed on the training set images.

[0017] The improved YOLOv8 model was trained and optimized using the strawberry flower and fruit dataset, and its robustness was enhanced through data augmentation.

[0018] The trained model is used to detect strawberry flowers and fruits in real time, and the detection results are output.

[0019] Preferably, the acquisition of the image data includes:

[0020] Select strawberry flowers and fruits, and use the device to take pictures at a distance of 45 degrees from the ground and 20 to 50 centimeters from the strawberry row;

[0021] Shooting should take into account sunny days, rainy days, and lighting conditions during the morning, noon, and evening.

[0022] Preferably, the step of annotating and processing the image data includes: resolving the image data to a resolution of 3024×4032, manually annotating the strawberry flowers and fruits using an annotation tool, and forming a strawberry fruit dataset containing different weather and lighting conditions.

[0023] Preferably, the step of improving the robustness of the model through data augmentation includes: performing data augmentation processing on the training dataset, wherein the data augmentation methods include rotation, brightness adjustment, contrast adjustment, and adding Gaussian noise to simulate different scene changes and enhance the robustness of the model.

[0024] Preferably, the C2f-OREPA module replaces the original C2f module in YOLOv8n. Its structure and execution flow include: for any given input feature map, the C2f-OREPA module adopts the dual-branch parallel feature extraction architecture of C2f, while replacing the 3×3 standard convolution in the bottleneck structure of the original C2f module with the OREPA online convolution reparameterization block, so as to achieve a balance between feature extraction capability in the training stage and lightweight efficiency in the inference stage.

[0025] The C2f-OREPA module employs a parallel structure of main and residual branches to optimize feature extraction and propagation efficiency. The main branch focuses on deep feature extraction, while the residual branch handles fast feature propagation and reuse. The OREPA core block is embedded in the main branch. This block first adjusts the channel dimensions of the input feature map using 1×1 convolutions to fit the input requirements of the OREPA block, and then performs online convolutional reparameterization.

[0026] During the execution of the C2f-OREPA block, non-linear layers such as ReLU and Batch Normalization (BN) in the prototype reparameterization structure are first removed, retaining only convolutional layers as basic feature extraction units. Simultaneously, a linear scaling layer is added at the end of the convolutional layer to dynamically adjust the weights of the convolutional output features, optimizing gradient flow stability during training. Furthermore, a normalization layer is added at the end of each branch to further improve the model's convergence. After linearization, the linear properties of convolution operations are utilized to compress and merge the multi-branch convolutional structures within the OREPA block, directly summing the weights of each branch's convolutional kernel and merging them into a single convolutional layer, thus transforming a complex training structure into a simplified convolutional layer.

[0027] Meanwhile, the residual branch adjusts the channel count of the input feature map using a 1×1 convolution without participating in complex feature extraction operations, thus ensuring rapid feature transfer. The feature map processed by the OREPA block in the main branch is concatenated with the feature map output by the residual branch along the channel dimension, and then a 1×1 convolution is used to achieve channel integration and dimension calibration of the concatenated features. The final output feature map maintains the same size as the input feature map, and achieves full reuse and enhancement of channel-dimensional features.

[0028] During the training phase, the C2f-OREPA module relies on the complex multi-branch structure of the OREPA block, which provides stronger feature representation capabilities and extraction accuracy. During the inference phase, the OREPA block performs an equivalent transformation from multiple branches to a single convolutional layer without adding extra inference computation overhead, thus balancing model training performance and inference efficiency and meeting the needs of modern deep learning network stacking deployment.

[0029] Preferably, the EMA module structure includes: for any given input feature map EMA divides X into G sub-features along the channel dimension for learning different semantics, where X = [X0, X...]. i ,...X G-1 ], Attention weight descriptors for grouped feature maps are extracted via three parallel paths. Two parallel paths are on a 1x1 branch, and the third path is on a 3x3 branch. In the 1x1 branch, two one-dimensional global average pooling operations are applied along two spatial directions, while in the 3x3 branch, only 3x3 KAN convolutional layers are stacked. On one hand, two encoded features are concatenated along the image height direction, sharing the same 1x1 KAN convolution. After decomposing the output of the 1x1 KAN convolutional layer into two vectors, two non-linear sigmoid functions are used to fit a two-dimensional binomial distribution after linear convolution. The two-channel intelligent attention maps within each group are aggregated together using simple multiplication, and then aggregated with the sub-feature maps. On the other hand, the 3x3 branch uses 3x3 KAN convolutional layers.

[0030] Two tensors are used, one containing the output of a 1x1 branch and the other the output of a 3x3 branch. Then, two-dimensional global average pooling is used to encode the global spatial information of the 1x1 branch output. Before the joint activation mechanism for channel features, the output of the smallest branch is directly transformed into its corresponding dimensional shape. The formula for two-dimensional global pooling is shown in (1).

[0031]

[0032] Where H represents the height of the feature map, W represents the width of the feature map, and (i,j) represents the pixel value of the feature map.

[0033] In the two-dimensional global output, a linear transformation is performed by fitting the natural nonlinear function of the two-dimensional Gaussian mapping using Softmax average pooling. The first spatial attention map is obtained by multiplying the output of the parallel processing above with a matrix dot product. Furthermore, two-dimensional global average pooling is also used to encode global spatial information in the 3x3 branch, while the 1x1 branch is directly converted to the corresponding dimensional shape before the joint activation mechanism of channel features. Based on this, a second spatial attention map that preserves the entire precise spatial location information was derived.

[0034] Finally, the output feature map within each group is computed as a set of two generated spatial attention weight values, which are then processed using the sigmoid function. This captures pixel-level pairwise relationships and highlights the global context across all pixels. The final output feature map size of the EMA is the same as the input feature map size, which is efficient for stacking in modern architectures.

[0035] Preferably, the C2f-DCNv3 module replaces the last two C2f modules in the backbone network. Its structure and execution flow include: for any given input feature map, the C2f-DCNv3 module retains the core of the dual-branch parallel architecture of C2f, replaces the 3×3 standard convolution in the bottleneck structure of the original C2f module with DCNv3 (deformable convolution v3), and enhances the flexibility and accuracy of deep feature extraction by dynamically adjusting the receptive field of the convolution kernel to adapt to the changes in the target spatial morphology.

[0036] This module consists of a main branch (deep feature extraction branch) and a residual branch (feature fast reuse branch). The two paths process the input features in parallel: the residual branch completes the channel dimension matching of the input features through 1×1 convolution, does not participate in complex feature transformation, and is only responsible for the fast transfer of features and the preservation of original information, ensuring the efficiency of feature reuse; the main branch sequentially goes through 1×1 convolution, DCNv3 module, and 1×1 convolution to build a complete process of "channel adjustment - deformation feature extraction - dimension calibration".

[0037] The 1×1 convolution in the main branch first adjusts the number of channels in the input feature map to match the input requirements of the DCNv3 module, while compressing redundant feature dimensions to improve computational efficiency. Then, the feature map is fed into the DCNv3 module, which learns the spatial deformation information of the target and dynamically generates the offset of the convolution kernel and attention weights. This allows the convolution kernel to adapt to the irregular shape, spatial position shift, and local scale changes of the target, breaking through the limitation of the fixed receptive field of standard convolution and accurately capturing key features such as the edges and textures of deformed targets. The feature map output by the DCNv3 module is then subjected to a second 1×1 convolution to complete the secondary calibration of the channel dimensions, integrate deep deformation features, and optimize feature representation.

[0038] The deep deformation feature map processed by the DCNv3 module in the main branch is spliced ​​and fused with the basic feature map output by the residual branch along the channel dimension to fully integrate the dual information of "deformation features + original features". Finally, the spliced ​​feature map is channel integrated and feature purified by 1×1 convolution to remove redundant information and output the final feature map.

[0039] The C2f-DCNv3 module combines the dynamic receptive field advantage of deformable convolution with the parallel feature reuse architecture of C2f, which significantly improves the model's feature extraction capabilities for irregular shapes, spatial offsets, and partially occluded targets without significantly increasing computational overhead. It is especially suitable for the high-precision feature representation requirements of complex targets in deep networks, providing more reliable deep feature support for subsequent detection tasks.

[0040] This invention provides a method for detecting strawberry maturity based on an improved YOLOv8n.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] This invention replaces the original YOLOv8 C2f module with the C2f-OREPA module, combining the advantages of C2f's branch-parallel feature reuse with OREPA's online convolution reparameterization technology. During the training phase, it retains the strong feature extraction capability of complex structures, and during the inference phase, it completes the compression and simplification of multi-branch to single convolutional layer. While significantly improving the accuracy of model feature representation, it greatly reduces the computation and storage overhead of the training process, balancing training efficiency and inference speed, effectively simplifying model complexity, and adapting to the deployment requirements of mobile devices and edge computing platforms.

[0043] This invention enhances the model's ability to capture multi-scale features of targets and fuse global-local contextual information by embedding an efficient multi-scale attention (EMA) module into the backbone network. At the same time, it optimizes the model's suppression effect on environmental noise and illumination changes in complex scenes, especially improving the feature extraction accuracy of small-sized targets and occluded targets, and solves the technical problem of insufficient accuracy of conventional models in detecting multi-scale targets in complex backgrounds.

[0044] This invention replaces the last two C2f modules in the backbone network with C2f-DCNv3 modules, and integrates deformable convolution (DCNv3) into the C2f parallel structure. By dynamically adjusting the receptive field of the convolution kernel to adapt to the irregular shape and spatial position changes of the target, it effectively enhances the model's feature capture ability for deformable and overlapping targets, further improves the localization accuracy and classification accuracy of target detection, and makes the model more adaptable to various targets in complex scenes.

[0045] The improved model of this invention relies on the lightweight and parametric design of C2f-OREPA, the multi-scale feature enhancement effect of the EMA module, and the deformation feature adaptation capability of C2f-DCNv3. While significantly reducing the number of parameters and computational load, it achieves a dual improvement in detection accuracy and inference speed, achieving the optimal balance between real-time performance and accuracy, and can fully meet the real-time and high-precision requirements of target detection tasks.

[0046] This invention, through the synergistic optimization of three improved modules, enables the model to have stronger environmental robustness and scene adaptability. It can effectively cope with the complex challenges in actual detection, such as drastic changes in lighting, target occlusion, shape deformation, and scale differences. It significantly improves the detection stability and reliability of the model in real scenes, and provides strong technical support for the engineering implementation and intelligent application in related fields.

[0047] The optimized structure and lightweight design employed in this invention effectively reduce computational resource consumption, making it particularly suitable for deployment on mobile devices. In the context of smart agriculture, intelligent devices can be used for real-time on-site detection to quickly assess strawberry maturity, thus assisting in agronomic management. Through multi-scale fusion, attention mechanisms, and optimized network design, the accuracy and robustness of strawberry maturity detection are significantly improved, maintaining high detection precision and providing more reliable support for agricultural production.

[0048] Verification has shown that, on the same real-world complex strawberry image dataset, compared with the baseline model, the detection model of this invention not only improves the performance of strawberry fruit detection but also significantly reduces the number of model parameters and computational load. It effectively solves the problems of low accuracy and high computational complexity in strawberry fruit recognition and detection in complex scenarios, providing a solution for application deployment on limited equipment in agricultural scenarios. Attached Figure Description

[0049] Figure 1 This is a flowchart of a strawberry maturity detection method based on an improved YOLOv8n according to the present invention.

[0050] Figure 2 Examples of the shooting environment and camera shooting location for the strawberry image dataset of the present invention are shown in the figures, where (a) strawberry greenhouse on a sunny day, (b) strawberry greenhouse on a cloudy or rainy day, and (c) mobile phone shooting location.

[0051] Figure 3 This is a schematic diagram of the detection model in this invention.

[0052] Figure 4 The diagram shows the C2f-OREPA module and its components of the present invention, wherein (a) C2f module, (b) C2f-OREPA module, (c) Conv module, and (d) OREPA-Bottleneck module.

[0053] Figure 5 This is a schematic diagram of the EMA module structure of the present invention.

[0054] Figure 6 This is a schematic diagram of the C2f-DCNv3 module structure of the present invention.

[0055] Figure 7 The results of this invention's detection model on strawberries at different stages of maturity are shown. Detailed Implementation

[0056] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples.

[0057] This invention provides a strawberry maturity detection method based on an improved YOLOv8n. This method acquires strawberry images at different maturity levels in real-world scenarios and designs an optimized detection model structure, enabling efficient detection of multi-scale strawberry targets in complex agricultural settings. It significantly improves detection accuracy and reliability, particularly addressing practical issues such as leaf occlusion, fruit overlap, and maturity differences within the same cluster of fruit. The method employs a lightweight single-stage target detection algorithm, optimizing model parameters and computational complexity for resource-constrained environments to meet the deployment requirements of low-computing-power hardware. Furthermore, by improving the backbone network, adding an adaptive attention mechanism, and optimizing the multi-scale feature fusion module, it effectively enhances the detection capability for small targets and robustness in complex scenarios. In addition, to improve the model's adaptability to different environmental conditions, we designed a training method based on multiple data augmentation strategies and multi-task learning, enabling the model to more accurately identify strawberry fruits at different maturity levels. This contributes to improving agricultural production efficiency, reducing labor costs, and providing reliable technical support for the automation and intelligentization of modern agriculture.

[0058] Specifically, this invention designs a strawberry fruit ripeness detection model based on an improved YOLOv8n network. First, a strawberry image dataset is manually collected; second, an improved network model is designed and trained on the image dataset; finally, the optimal model is used to validate the model on strawberry fruits in real-world scenarios. Figure 1 The technical roadmap of this invention is as follows:

[0059] Step 1: Obtain data images of strawberry fruits at different ripeness levels in real-world scenarios through methods such as manual photography.

[0060] In implementing this invention, the first step is to collect image datasets of strawberry fruits of different varieties and ripeness levels. These image datasets will be used to train and validate the improved YOLOv8n model. Since strawberry ripeness is affected by various factors, such as light intensity, shading, and overlap between fruits, image acquisition needs to be conducted in different scenarios to ensure the diversity and representativeness of the dataset. Image acquisition is carried out in a greenhouse environment to ensure consistency. Simultaneously, strawberry fruits at different ripeness levels may be affected by different light conditions within the greenhouse, which helps simulate light variation scenarios that may occur in real-world applications. Furthermore, strawberry fruits should be photographed against different backgrounds; for example, strawberry fruits may be shaded by other plants, strawberry leaves, or background objects during photography, to improve the robustness of the model. By simulating these complex scenarios, it can be ensured that the model can accurately identify shaded or overlapping strawberry fruits.

[0061] High-resolution smartphones are used to acquire high-quality image data to ensure that the images have sufficient clarity and detail, making it easier for the model to effectively learn the characteristics of strawberry fruits. Figure 2 The study showcases the shooting environments of strawberry fruits at different maturity levels, including strawberry greenhouses on sunny days, strawberry greenhouses on rainy days, and locations captured by mobile phones, to ensure the model can recognize strawberry fruits under different conditions. During the data collection process, shooting should be conducted at different times, such as morning, noon, and evening, to increase the diversity of the dataset.

[0062] Step 2: Preprocess all the images acquired in Step 1, and then use image annotation tools such as LabelImg to annotate the strawberry images with information on maturity categories. In this invention, the acquired images are divided into 6 categories according to variety and different maturity levels, and are divided into training set, validation set and test set according to proportion. Finally, data augmentation is performed on the training set images.

[0063] Before labeling the acquired images, preprocessing is necessary to meet the requirements of subsequent image annotation, data augmentation, and model training. First, all images need to be inspected, and blurry images are removed. Second, the dimensions of all images need to be standardized; in this embodiment, the image size is adjusted to 640×640 to ensure consistency and meet the model input requirements. This step not only reduces memory usage during training but also significantly improves training efficiency. Third, image normalization is performed to accelerate model training and prevent differences in data values ​​from affecting network convergence.

[0064] Supervised learning object detection algorithms in deep learning require a large number of labeled images. Therefore, labeling is a very important step in the process of building a dataset. By using the image labeling tool LabelImg, bounding boxes are added to the strawberry fruits in each image and the corresponding categories are labeled. According to the variety and different maturity, the collected images are divided into six categories: Bud, Flower, Young, Expansion, Color Change, and Ripening. These six categories represent the germination period, flowering period, young fruit period, swelling period, color change period, and ripening period, respectively.

[0065] After completing the annotation, to ensure a uniform distribution of data across all categories, we randomly sampled the strawberry dataset to guarantee the representativeness of the training, validation, and test sets and to avoid the impact of class imbalance on training and validation. Finally, we divided the dataset into training, validation, and test sets in an 8:1:1 ratio.

[0066] To improve the model's generalization ability and avoid overfitting, this invention performs data augmentation operations on the training set images. Data augmentation not only increases the sample size of the training set but also simulates various changes that may occur in the real environment, thereby improving the model's robustness. The data augmentation operations used in this invention include horizontal flipping, vertical flipping, brightness adjustment, contrast adjustment, and adding Gaussian noise. During the augmentation process, two of these techniques are randomly selected to transform the image. Furthermore, the improved YOLOv8n model retains the Mosaic data augmentation operations performed on the original model during the training phase. Therefore, the data augmentation operations used in this embodiment can fully meet the model's need to learn comprehensive and diverse features, enabling the model to better adapt to complex environmental changes in real-world scenarios.

[0067] Step 3: Construct the detection model.

[0068] The detection model of this invention is based on the improved YOLOv8n model, such as... Figure 3 As shown, it includes one or more of the following improvements to YOLOv8n:

[0069] Improvement 1: The original C2f structure of YOLOv8n is replaced by the C2f-OREPA structure, and the reparameterization mechanism is integrated to improve the feature expression capability while reducing the amount of computation, so as to achieve efficient feature extraction from shallow to medium layers.

[0070] Improvement 2: Introducing the EMA attention module to enhance the robustness of multi-scale feature extraction.

[0071] Improvement 3 replaces the C2f structure with C2f-DCNv3 in the backbone network and introduces deformable convolution DCNv3. By dynamically adjusting the receptive field of the convolution kernel, it enhances the feature capture capability of irregular targets.

[0072] The improvements mentioned above will be described one by one below.

[0073] For improvement 1, C2f-OREPA is based on online convolutional reparameterization technology. Its core objective is to reduce the training cost and structural complexity of deep learning models while improving their performance in computer vision tasks. This technology is widely applicable to scenarios such as image classification, object detection, and semantic segmentation. Its core principles revolve around three aspects: First, it constructs a two-stage process of "online optimization + training module compression" to achieve a balance between training efficiency and model accuracy; second, it introduces a linear scaling layer to optimize the training stability and feature learning capability of the online module; and third, through structural simplification, it compresses complex network blocks during training into a single convolutional layer, significantly reducing memory usage and computational overhead during the training phase, thereby improving training speed while ensuring performance.

[0074] OREPA's online reparameterization optimizes the training structure through a three-step progressive operation: First, it removes non-linear components, eliminating non-linear layers such as ReLU and Batch Normalization (BN) from the prototype block, retaining only convolutional layers as basic feature extraction units; second, it adds linear scaling layers, introducing a scaling mechanism at the output of the convolutional layers to dynamically adjust the feature weights, providing support for efficient weight updates; finally, it adds normalization layers, adding normalization modules at the end of branches to further improve the stability of model training, laying the foundation for subsequent structure compression and performance optimization.

[0075] OREPA's online reparameterization process proceeds systematically through two core stages: block linearization and block compression. The first stage, block linearization, linearizes the prototype reparameterization blocks, removing non-linear components and retaining only convolutional and batch normalization (BN) layers. Simultaneously, a linear scaling layer is embedded, transforming the block structure into a purely linear combination, ensuring the controllability and stability of the training process. The second stage, block compression, utilizes the linearity of convolutional operations to merge the linearized multi-branch structure into a single "OREPA Conv" convolutional layer. By directly summing the weights of the multi-branch convolutional kernels, the complex training structure is compressed into a simple convolutional layer, significantly reducing the computational and storage costs during training without sacrificing model performance.

[0076] The linear scaling layer is a key functional component in the OREPA framework. Its core value lies in optimizing the efficiency of weight updates and gradient stability during training. By dynamically scaling the output features of convolutional layers, this layer enhances the flexibility and adaptability of the model when learning features, better meeting the feature learning needs of different tasks. It also effectively maintains the stable flow of gradients during training, avoiding optimization problems such as gradient vanishing or exploding. Simultaneously, the linear scaling layer accelerates model convergence, reducing training resource consumption while ensuring effective model training. It is a core design element balancing training efficiency and performance.

[0077] The OREPA framework comprises four core components adapted to different scenarios, each meeting the feature extraction needs at different levels: First, a frequency prior filter, consisting of a 1×1 convolution, a 3×3 frequency prior filter, and a scaling layer, focuses on enhancing the ability to capture frequency domain information of features; second, a linear depthwise separable convolution, integrating a 3×3 depthwise separable convolution (DW Conv), a 1×1 pointwise convolution (PW Conv), and a scaling layer, achieving a balance between lightweight design and feature extraction efficiency; third, a reparameterized 1×1 convolution, consisting of two 1×1 convolutions with scaling layers, used for reparameterizing larger-scale convolutional structures, balancing structural complexity and training efficiency; and fourth, a linear deep stem cell, consisting of three 3×3 convolutional layers, suitable for initial feature extraction of input data, enhancing the model's early feature representation capabilities.

[0078] Module compression during training in the OREPA framework refers to a technique that transforms complex network blocks into single convolutional layers during the training phase through a series of operations. Specifically, this involves removing non-linear activation functions, merging multi-branch convolutional layers, and introducing scaling layers to adjust weight scales, thereby simplifying the training structure and significantly reducing memory usage and computational costs during training. From the perspective of structural adaptation between training and inference, OREPA is compatible with various complex structures during training, including standard convolutional layers and typical reparameterized blocks. After module compression, all training structures are uniformly simplified to a single convolutional layer during inference. This design ensures both feature learning capabilities during training and high speed and low resource consumption during inference. Furthermore, OREPA blocks can correspond to 3×3 convolutions during both training and inference, balancing structural consistency and engineering practicality.

[0079] Typically, complex multi-branch structures are used in network training. However, during the inference phase, the parameters of these complex structures are equivalently transformed into single-linear layer parameters, thereby improving network accuracy while maintaining inference speed. However, this also results in a large model size during training, impacting training speed. Therefore, this paper introduces an online convolutional reparameterization block to replace convolutions with a kernel size of 3, as shown in the following structure: Figure 4 As shown, employing an online convolution reparameterization strategy can effectively improve network performance while reducing the computational and storage overhead of reparameterizing the model during training, thereby increasing training speed.

[0080] The online reparameterization process mainly includes two steps: module linearization and compression. Module linearization involves three steps: First, the non-linear layers in the module are removed, i.e., the Batch Normalization (BN) layers in the module are reparameterized. Then, a linear scaling layer is added at the end to replace the original BN layer. Finally, to improve the stability of network training, a normalization layer is added at the end of each branch. After completing module linearization, the complex structure can be further simplified into convolutional layers through module compression, including simplified sequential and parallel structures.

[0081] The calculation process of convolution is shown in formula (2). The parallel structure is simplified through formulas (3) and (4).

[0082] Y = W * X (2)

[0083] Where: W represents the convolution kernel weight, X represents the input feature map, and Y represents the output feature map.

[0084] Y = WN(WN-1) * …(W2*(W1 * X)′)) (3)

[0085] Y = WN(WN-1) * …(W2 * W1)) * X = We * X (4)

[0086] When simplifying the parallel structure, due to the linear nature of convolution, the convolution kernels on multiple branches can be directly summed to obtain a unified convolution kernel, and then a single convolution operation can be performed. In this way, multiple branches can be merged into a single branch, and the calculation method is shown in formula (5).

[0087]

[0088] In formula (5): Wm is the weight of the m-th branch.

[0089] This paper combines the C2f module in YOLOv8 with the Online Convolution Reparameterization (OREPA) module to design the C2f-OREPA module, the structure of which is as follows: Figure 4 As shown in (b). The online convolution reparameterization module is used to replace the bottleneck structure of the 3×3 convolution in the C2f module, as follows: Figure 4 As shown in (d). Finally, this paper uses C2F-OREPA to replace the first two modules and all C2f modules in the YOLOv8n backbone to improve the model's accuracy without significantly increasing the additional inference time.

[0090] For Improvement 2, when detecting strawberries in complex environments, numerous non-target objects such as leaves, stems, and soil are encountered. These background elements affect the network model's ability to accurately identify ripe strawberries. Attention mechanisms are a commonly used neural network technique that allows the network to focus more on key information of the current task, thereby improving network efficiency and accuracy. Therefore, this paper proposes a method to introduce efficient multi-scale attention (EMA) into the backbone network. This attention mechanism, based on the concepts of feature grouping, parallel structure, and cross-space learning, can perform more refined pixel-level attention on high-level feature maps without reducing dimensionality, thus improving the network's detection performance. The overall architecture of the EMA attention mechanism is as follows: Figure 5 As shown.

[0091] Multi-scale channel attention (EMA) is a widely used technique in computer vision tasks, aiming to improve model performance through feature extraction and aggregation at different scales. This paper introduces an efficient EMA module that performs exceptionally well in image classification and object detection tasks.

[0092] The EMA module avoids more sequential processing and greater depth through parallel substructures. Its overall structure, as shown in the figure, mainly includes the following parts:

[0093] Feature grouping: For any given input feature map, EMA divides it into multiple sub-feature groups, each learning a different semantic meaning. This grouping method not only enhances feature learning of semantic regions but also compresses noise.

[0094] Parallel Subnetwork: EMA employs three parallel paths to extract attention weight descriptors for grouped feature maps. Two paths are 1x1 branches, and the third is a 3x3 branch. The 1x1 branch uses one-dimensional global average pooling to encode channel information in two spatial directions. The 3x3 branch captures multi-scale feature representations through 3x3 convolutions.

[0095] Cross-spatial learning: EMA enriches feature aggregation by providing cross-spatial information aggregation methods across different spatial dimensions. Specifically, the output of the 1x1 branch encodes global spatial information through two-dimensional global average pooling, while the output of the 3x3 branch is directly converted to the corresponding dimensional shape. Then, these outputs are aggregated through matrix dot product operations to generate the first spatial attention map.

[0096] The EMA attention mechanism divides the input feature map into G groups across channels, reducing computational overhead while preserving information from each channel. It employs a multi-branch structure, evenly distributing spatial semantic features within each feature group as two 1×1 convolutional branches and one 3×3 convolutional branch. The 1×1 branches use two global average pooling operations to encode features along two spatial dimensions, facilitating the aggregation of global information in the feature map. The 3×3 branches capture spatial feature information at multiple scales through convolutional kernels. Furthermore, cross-spatial learning establishes information dependencies between channels and spatial locations, achieving rich feature aggregation. Next, both 1×1 and 3×3 branch feature maps are introduced, and a 2D global average pooling method is used to encode global spatial information into the outputs of the two branches, as shown below. Figure 5 As shown. Finally, each set of output features is computed as a set of two spatial attention weights. Then, the sigmoid function is used to capture pixel-level pairwise relationships to highlight the global context of each pixel.

[0097] An EMA attention mechanism was added after each of the four C2f modules in the backbone network to enhance the network's feature representation of strawberry flower and fruit region information in the image, thereby further improving the overall detection accuracy of strawberry flowers and fruits.

[0098] For Improvement 3, traditional convolutional methods typically use fixed-size kernels (e.g., 1×1, 3×3, 5×5) to sample the input feature map, resulting in poor feature extraction performance for irregularly shaped targets. However, the research object in this paper is ground-grown strawberries, whose shape features exhibit some irregularity. Therefore, traditional convolutional methods struggle to achieve effective feature extraction. To better extract features from ground-grown strawberries, this paper introduces the state-of-the-art deformable convolution DCNv3. By adding an offset variable to each sampling position in the convolution kernel, it can more closely approximate the shape and size of the object. Furthermore, a modulation mechanism is utilized to enhance the network's ability to focus on the target region.

[0099] Based on DCNv2, DCNv3 was redesigned and adjusted, and the specific improvements include the following parts.

[0100] Shared Projection Weights. Similar to regular convolution, different sampling points in DCNv2 have independent projection weights, so the parameter size is linearly related to the total number of sampling points. To reduce parameter and memory complexity, borrowing the idea of ​​separable convolution, position-independent weights are used instead of group weights, and projection weights are shared among different sampling points, thus preserving all sampling position dependencies.

[0101] A multi-group mechanism was introduced. The multi-group design was first introduced in grouped convolutions and widely used in multi-head self-attention in Transformers. It can be combined with adaptive spatial aggregation to effectively improve feature diversity. Inspired by this, researchers divided the spatial aggregation process into several groups, each with an independent sampling offset. Thus, different groups within a single DCNv3 layer possess different spatial aggregation patterns, resulting in rich feature diversity.

[0102] Sampling point modulation scalar normalization. To alleviate the instability problem when the model capacity increases, the researchers set the normalization mode to Softmax normalization per sampling point. This not only makes the training process of large-scale models more stable, but also establishes the connection relationship of all sampling points.

[0103] Deformable Convolution DCNv3 offers several improvements over DCNv2. First, it implements shared convolutional weighting (wg), reducing computational complexity. Second, a multi-group mechanism is introduced to enhance the expressiveness of deformable convolutions. Finally, the modulation scalar is normalized, making the network training process more stable.

[0104] The Deformable Convolutional DCNv3 structure first divides the input feature map into G groups, performs convolution operations on each group of feature maps to obtain a set of prediction results for the convolution kernel offset and modulation factor, and finally calculates the output feature map.

[0105] To further improve the network's feature extraction capability for mature strawberry targets, deformable convolutional layers DCNv3 were added to the last two C2f modules of the YOLOv8n backbone network. Since 1×1 convolutions are mainly used to change the number of channels, this paper replaces the 3×3 convolutions in the C2f modules with deformable convolutional layers DCNv3, constructing a C2f-DCNv3 module, as shown below. Figure 6 As shown in the diagram, this module enables the network model to better adapt to the needs of strawberry target detection in complex natural environments, improving the model's detection performance in actual harvesting operations.

[0106] Step 4: Input the training set images obtained after data augmentation in Step 2 into the detection model constructed in Step 3 for iterative training, and finally output the model with the best detection effect.

[0107] In this embodiment, a preprocessed and augmented dataset was used to train the model. Before training begins, the training parameters of the improved model need to be configured, including the input image size (640×640), training epochs (300 epochs), optimizer (stochastic gradient descent, SGD), and batch size, among other initial training information. During training, the model performs forward propagation based on the input image and labels, calculating the predicted value for each target. Then, by comparing the difference between the loss function and the actual labels, the model performs backpropagation to update the weights. The optimizer updates the model's weights based on the gradients to minimize the loss function. Hyperparameters such as learning rate and momentum affect the convergence speed and stability of the optimization process. Therefore, a dynamic learning rate decay strategy is introduced to ensure rapid convergence in the early stages and fine-tuning in the later stages, thus achieving a balance between detection speed and accuracy. In addition, mixed-precision training can improve training speed and reduce memory usage while ensuring model accuracy. Simultaneously, the model employs an early stopping mechanism to prevent overfitting and improve generalization ability. At the end of each training cycle, the model is evaluated on the validation set, and metrics such as precision, recall, and mean average precision (mAP) are calculated. If the validation set performance is poor, optimization can be achieved by adjusting the learning rate, modifying the model structure, or increasing the dataset.

[0108] This example provides a wealth of information about the model. The first is the weights folder, containing two weight files: one is the final model weight file at the end of training; the other is the model weight file that performed best on the validation set. These two files store the model's weight parameters during training. The second is the args.yaml file, which mainly stores the parameters specified during training. The third is the confusion matrix, used to analyze and evaluate the model's performance. The fourth is the F1_Curve file, showing the changes in F1 scores under different classification thresholds. The fifth is the P_Curve file, showing the relationship between precision and confidence. Generally, as confidence increases, detection accuracy also increases. The sixth is the R_curve file, showing the relationship between recall and confidence. Generally, as confidence increases, recall decreases. The seventh is the PR_curve file, showing the relationship between model precision and recall under different classification thresholds. The closer the PR curve is to the upper right corner, the better the model performance and the more correctly it identifies positive samples. The eighth file is results.csv, which records the model's parameter information during training. The ninth file is results.csv, which displays graphs of training loss, validation loss, mAP50, mAP95, metrics / precision, and metrics / recall during training, as well as the detection results during training. This information provides a comprehensive understanding of the model's performance during training, allowing for further performance analysis and optimization to improve the strawberry ripeness detection model.

[0109] The evaluation metrics used in this embodiment are common in current object detection tasks, including precision (P), recall (R), mean average precision (mAP), and frames per second (FPS). Their calculation methods are as follows:

[0110]

[0111] In this context, TP (True Positive) represents the number of correctly detected positive samples, i.e., the number of regions containing the target that the model correctly identifies; FP (False Positive) represents the number of falsely detected negative samples, i.e., the number of regions without targets that the model incorrectly identifies as targets; and FN (False Negative) represents the number of missed positive samples, i.e., the number of regions containing targets that the model fails to detect correctly. Precision refers to the proportion of correctly predicted positive samples among the identified samples. Higher precision indicates better recognition performance. Recall refers to the proportion of correctly predicted positive samples among the identified positive samples. mAP50 refers to the average accuracy across all categories when the IoU threshold is 0.5. Furthermore, FPS represents the number of frames processed per second by the algorithm, i.e., the number of images detected per second. A higher value indicates faster detection speed. When evaluating the algorithm, the ultimate goal is to make it more lightweight while maintaining detection accuracy, taking into account all the above metrics.

[0112] Step 5: Input the strawberry image to be detected into the optimal detection model obtained in Step 4 to obtain an image with a bounding box labeled with strawberry maturity category information.

[0113] This embodiment inputs the image to be detected into the improved YOLOv8n model, performs forward inference to obtain the target detection results (including target category, confidence level, and other information), and intuitively displays this information in the image for easy inspection and analysis. Figure 7 This embodiment demonstrates the visualization results of the detection of different types of strawberries based on the improved YOLOv8n model. As can be seen from the figure, the model can accurately identify the ripeness of different strawberry varieties.

Claims

1. A method for detecting strawberry maturity based on an improved YOLOv8n, characterized in that, Includes the following steps: Step 1: Capture images of strawberries at different ripeness levels in real-world scenes, and label strawberry ripeness category information and divide the dataset; Step 2: Construct a detection model, which is a YOLOv8n model with one or more of the following improvements: Improvement 1: C2f-OREPA is adopted to replace C2f, which integrates the dual-branch architecture with OREPA. During training, features are extracted through multiple branches, and during push, they are compressed into a single convolution. Gradients are optimized, with no additional overhead and improved accuracy and speed. Improvement 2: The backbone is embedded with an EMA module, which, through feature grouping, combines cross-space learning and global pooling to enhance multi-scale capture and noise reduction, thereby improving the accuracy of small targets and occluded targets. Improvement 3 replaces the last two C2f functions with C2f-DCNv3, integrates the dual-branch architecture with DCNv3, dynamically adjusts the convolution kernel offset and weights, enhances deformation feature extraction, and improves the accuracy of occluded and overlapping targets. Step 3: Use the training set divided in Step 1 to train the detection model of Step 2, and input the collected strawberry images into the trained detection model to obtain images with strawberry maturity category information bounding boxes.

2. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that, In step 1, the image data is acquired in the following way: Strawberry flowers and fruits were selected, and images were taken at a distance of 20 to 50 centimeters from the strawberry rows with the camera at a 45-degree angle to the ground. The images covered different lighting conditions, including sunny days, rainy days, and different times of day, with a resolution of 3024×4032, forming a dataset of strawberry fruits under different weather and lighting conditions.

3. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that, Step 1 specifically includes: The acquired strawberry images were resized and normalized; bounding boxes were added to the strawberry fruits in each image, and the images were divided into 6 categories according to variety and maturity: bud stage, flowering stage, young fruit stage, swelling stage, color change stage, and ripening stage. The training set, validation set, and test set are divided according to a preset ratio; data augmentation operations are performed on the training set images, including rotation, brightness adjustment, contrast adjustment, and addition of Gaussian noise.

4. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that, In Improvement 1, the structure and execution flow of the C2f-OREPA module include: The C2f-OREPA module adopts the dual-branch parallel feature extraction architecture of C2f, and replaces the 3×3 standard convolution in the bottleneck structure of the original C2f module with the OREPA online convolution reparameterization block. The module includes a main branch and a residual branch. The main branch is responsible for deep feature extraction, and the residual branch completes the fast transfer and reuse of features. The main branch embeds the OREPA core block, which first adjusts the channel dimension of the input feature map through 1×1 convolution, and then sends it into the OREPA block to perform online convolution reparameterization processing. During the execution of the OREPA block, non-linear layers such as ReLU and Batch Normalization (BN) in the prototype reparameterization structure are first removed, leaving only the convolutional layers; a linear scaling layer is added at the end of the convolutional layer, and a normalization layer is added at the end of the branch; after the linearization process is completed, the linear properties of convolution are used to directly sum the weights of the multi-branch convolutional kernels and merge them into a single convolutional layer. The residual branch performs a 1×1 convolution on the input feature map to complete channel matching; the feature map processed by the OREPA block of the main branch is concatenated with the feature map output by the residual branch along the channel dimension, and then a 1×1 convolution is performed to complete channel integration and dimension calibration, outputting a feature map with the same size as the input.

5. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that, In Improvement 2, the structure and execution flow of the EMA module include: When the EMA module processes the input feature map, it first divides the feature map into G sub-features along the channel dimension to learn differential semantics; then it extracts attention weights through three parallel branches: the two 1×1 branches perform one-dimensional global average pooling along two spatial directions respectively, and the 3×3 branches stack 3×3 KAN convolutional layers. The two encoded features of the 1×1 branch are concatenated along the image height direction and share a 1×1 KAN convolution. The output is decomposed into two vectors and activated by Sigmoid. Then, the intra-channel attention map is aggregated by multiplication and fused with the sub-feature map. At the same time, the outputs of the 1×1 and 3×3 branches are subjected to two-dimensional global average pooling to encode global spatial information. After being activated by Softmax and fitted with a Gaussian mapping, two spatial attention maps are generated by matrix dot product. Finally, the two sets of attention weights are aggregated and activated by Sigmoid to output a feature map with the same size as the input, thus achieving multi-scale feature capture.

6. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that, In improvement 3, the structure and execution flow of the C2f-DCNv3 module include: The C2f-DCNv3 module retains the dual-branch parallel architecture of C2f, and replaces the 3×3 standard convolution in the bottleneck structure of the original C2f module with the DCNv3 deformable convolution. The module includes a main branch and a residual branch. The residual branch completes channel dimension matching through 1×1 convolution, which is responsible for fast feature transfer and preservation of original information. The main branch sequentially goes through 1×1 convolution, DCNv3 module, and 1×1 convolution; the 1×1 convolution adjusts the number of channels and compresses redundant dimensions; the DCNv3 module dynamically generates convolution kernel offset and attention weight by learning spatial deformation information to adapt to the irregular shape and spatial offset of the target. The deep deformation feature map processed by the main branch and the feature map output by the residual branch are concatenated along the channel dimension. Channel integration and feature purification are performed through 1×1 convolution to output the final feature map.

7. The strawberry maturity detection method based on the improved YOLOv8n according to claim 1, characterized in that: In step 3, the model training parameter configuration includes the following: the input image size is set to 640×640, the training epochs are set to 300, the optimizer uses the stochastic gradient descent (SGD) algorithm, and a dynamic learning rate decay strategy, mixed precision training, and early stopping mechanism are introduced. Model evaluation metrics include: precision, recall, average accuracy (mAP), and frames per second (FPS).