Lightweight target maturity detection method and system and application thereof

By improving the YOLOv8 model to SL-YOLO, combining spatial depth decomposition convolution and large separable kernel attention mechanism, and incorporating layer adaptive amplitude pruning, the problems of missed detection and high computational complexity in strawberry maturity detection are solved, achieving efficient and lightweight maturity detection, which is suitable for automated management of strawberries.

CN120976916APending Publication Date: 2025-11-18XIAMEN UNIV OF TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511052351.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing strawberry maturity detection technologies suffer from missed detections in complex natural environments. Furthermore, traditional algorithm models have a large number of parameters and high computational complexity, making them difficult to deploy on resource-constrained mobile or embedded devices, and thus failing to meet the needs of large-scale agricultural production.

Method used

The SL-YOLO model is adopted, and the YOLOv8 model is improved by using a spatial depth decomposition convolution module and a fast spatial pyramid pooling module with a large separable kernel attention mechanism. Combined with a layer adaptive amplitude pruning strategy, a lightweight SLP-YOLO model is constructed to achieve efficient maturity detection.

Benefits of technology

It achieves high-precision, lightweight detection of strawberry maturity in complex environments, reduces reliance on hardware resources, improves detection efficiency and robustness, and is highly adaptable, making it suitable for automated strawberry cultivation and harvesting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976916A_ABST
    Figure CN120976916A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight target maturity detection method and system and application of the lightweight target maturity detection method and system. The method comprises the steps that real-time monitoring data of a to-be-detected target is collected; based on a YOLOv8 network structure, improving a backbone network and a neck network of a YOLOv8 model by utilizing a space depth decomposition convolution module and combining a rapid space pyramid pooling module of a large separable kernel attention mechanism to form an SL-YOLO model, and constructing an SLP-YOLO model by utilizing a layer adaptive amplitude pruning strategy; the SLP-YOLO model is trained; and inputting real-time monitoring data of a to-be-detected target into the trained SLP-YOLO model to obtain a maturity detection result of the to-be-detected target. According to the method, through structure optimization and pruning strategies, the model complexity is remarkably reduced while the detection precision is ensured, and the method is suitable for actual application scenes with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, deep learning, and fruit maturity detection technology, specifically to a lightweight target maturity detection method, system, and its application. Background Technology

[0002] Accurate detection of strawberry ripeness is crucial for automated harvesting and refined planting management. Traditional manual detection methods rely on experience-based judgment, resulting in low efficiency and high costs, making them unsuitable for large-scale agricultural production. With the development of deep learning technology, detection methods based on YOLO series models are gradually being applied to strawberry ripeness detection, with YOLOv8 becoming a research hotspot due to its balance between accuracy and speed.

[0003] Existing technologies for strawberry maturity detection in complex natural environments still have many shortcomings. On the one hand, strawberry growing environments present challenges such as leaf shading, light variations, and complex backgrounds, leading to frequent missed detections. On the other hand, traditional algorithm models have large parameter sets and high computational complexity, making them difficult to deploy on resource-constrained mobile or embedded devices, thus limiting their widespread application in practical planting scenarios. For example, patent application CN117315473A discloses a strawberry maturity detection method based on an improved YOLOv8, which improves detection accuracy to some extent, but still has room for improvement in terms of model lightweighting and robustness in detecting small targets in complex environments. Therefore, how to achieve strawberry maturity detection that combines high accuracy, lightweight design, and strong adaptability has become a key problem that urgently needs to be solved in this field. Summary of the Invention

[0004] The purpose of this application is to provide a lightweight target maturity detection method, system, and application to solve the above-mentioned technical problems.

[0005] According to one aspect of this application, a lightweight method for detecting target maturity is proposed, the method comprising the following steps:

[0006] S1. Collect real-time monitoring data of the target to be detected;

[0007] S2. Based on the YOLOv8 network structure, the backbone and neck networks of the YOLOv8 model are improved by using a spatial depth decomposition convolution module and a fast spatial pyramid pooling module with a large separable kernel attention mechanism to form an SL-YOLO model. The SLP-YOLO model is then constructed using a layer adaptive amplitude pruning strategy.

[0008] S3. Train the SLP-YOLO model;

[0009] S4. Input the real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target.

[0010] In the above technical solution, by structurally improving and pruning YOLOv8, both model lightweighting and detection accuracy are taken into account, achieving efficient detection of target maturity, reducing dependence on hardware resources, and ensuring rapid processing and output of real-time monitoring data.

[0011] Furthermore, the backbone network of the SL-YOLO model includes one convolutional layer, several spatial depth decomposition convolutional modules, several feature fusion modules, and a fast spatial pyramid pooling module at the end of the backbone network that combines a large separable kernel attention mechanism.

[0012] In the above technical solution, a backbone network is used through multi-module collaborative design to improve feature extraction efficiency while preserving target detail features, providing high-quality feature input for subsequent maturity detection and enhancing the model's ability to perceive targets. Specifically, convolutional layers, as initial feature extraction units, perform preliminary feature mapping on the input real-time target monitoring data, reducing redundant noise in the original data and improving the quality of the starting point for feature extraction. Several spatial depth decomposition convolutional modules are cascaded, using a pooling-free downsampling method with interval partitioning and channel concatenation. This reduces computational complexity and parameter size while accurately preserving detailed features such as target edges and textures, avoiding feature loss due to pooling operations and enhancing the ability to capture subtle features of small targets. The feature fusion module can fuse features from different levels, combining shallow detail features with deep semantic features to compensate for the insufficient expressive power of single-level features, ensuring that the output features contain both target details and semantic features. The system provides detailed information and global semantic associations, offering comprehensive feature support for subsequent maturity assessment. Located at the end of the backbone network, the fast spatial pyramid pooling module, which combines a large separable kernel attention mechanism, further enhances the comprehensive capture capability of global and local features of the target through multi-scale feature fusion and enhanced direction awareness. The large separable kernel attention mechanism can focus on key areas of the target, while the fast spatial pyramid pooling can adapt to targets of different sizes. The combination of the two makes the backbone network output features more adaptable to changes in target scale and spatial position, laying a high-quality feature foundation for the neck network feature processing, and ultimately improving the accuracy and robustness of target maturity detection.

[0013] Furthermore, the neck network of the SL-YOLO model includes a feature fusion module, an upsampling layer, a spatial depth decomposition convolutional module, and a connection layer, while the detection head of the SL-YOLO model has several spatial depth decomposition convolutional modules. Through the application of spatial depth decomposition convolutional modules, the neck network and detection head enhance feature fusion and target localization accuracy while reducing computational load, thereby improving the response speed and accuracy of target maturity detection.

[0014] Furthermore, the spatial depth decomposition convolution module performs interval partitioning and channel concatenation on the input feature map to achieve pooling-free downsampling, preserving the target details in the real-time monitoring data of the target to be detected. Pooling-free downsampling avoids the detail loss problem caused by traditional pooling operations, and preserves key features such as the target's edges and textures through interval partitioning and channel concatenation, providing richer detailed evidence for maturity assessment.

[0015] Furthermore, methods for constructing fast spatial pyramid pooling modules incorporating large separable kernel attention mechanisms include:

[0016] Step a: Perform an initial convolution operation on the input feature map to generate an intermediate feature map, perform a max pooling operation on the intermediate feature map, and concatenate the intermediate feature map and the pooled feature map along the channel dimension to form a fused feature map.

[0017] Step b: Use depthwise separable convolutional layers to capture the spatial features of the fused feature map in the horizontal and vertical directions to form a direction-aware feature map; and use depthwise separable dilated convolutional kernels to convolve the direction-aware feature map to expand the feature perception range.

[0018] Step c: Adjust the number of channels in the convolutional layer and output the processed feature map.

[0019] In the above technical solution, by enhancing multi-scale feature fusion and direction perception capabilities, the model's ability to capture global and local features of the target is improved, while reducing computational complexity, thus providing more comprehensive feature support for target maturity detection.

[0020] Furthermore, when pruning the SL-YOLO model, the layer-adaptive amplitude pruning strategy adaptively allocates the pruning ratio based on the absolute value of the weights of the parameters in each layer of the SL-YOLO model. By adaptively allocating the pruning ratio, redundant parameters are removed to achieve model lightweighting while retaining the key feature extraction capability to the maximum extent, ensuring a balance between the detection accuracy and efficiency of the pruned model.

[0021] Furthermore, the layer adaptive magnitude pruning strategy also includes fine-tuning the pruned model. The layer adaptive magnitude pruning strategy calculates the pruning score according to the following formula and prunes the weight with the lowest score; when u < v, |W[u]| ≤ |W[v]| holds, and its calculation formula is as follows:

[0022]

[0023] In the formula, u and v represent weight indices; W[u] and W[v] represent the mapped weights corresponding to u and v respectively; the numerator (W[u]) represents the squared value of the current weight, reflecting its absolute importance; the denominator ∑ v≥u (W[v]) 2 represents the sum of the squares of all subsequent weights starting from the current weight, which is used to normalize the importance of the current weight.

[0024] In the above technical solution, precise pruning is achieved by quantifying the weight importance, and the fine-tuning step is combined to reduce the impact of pruning on the model performance. While significantly reducing the model parameter quantity, the detection accuracy is maintained, and the lightweight level and operation efficiency of the model are improved.

[0025] In the second aspect, the present application proposes an application of a lightweight target maturity detection method, which is applied to the cultivation and picking of strawberries based on the lightweight target maturity detection method described in the first aspect. Applying this method to the strawberry cultivation and picking scenario can achieve automatic and high-precision detection of strawberry maturity, provide technical support for precise cultivation management and picking timing judgment, reduce labor costs and improve production efficiency.

[0026] Furthermore, the training parameters in step S3 are set as follows: the input image size is 640 pixels × 640 pixels, the training batch is 100, the batch size is 16, the initial learning rate is 0.01, the initial momentum is 0.937, the confidence threshold and IoU threshold of the maturity detection result are set to 0.5 and 0.7 respectively; the pruning rate of the layer adaptive magnitude pruning strategy is 50%.

[0027] In the above technical solution, through the optimized training parameter and pruning rate settings, the convergence speed and detection accuracy of the model in the strawberry detection scenario are ensured, and at the same time, the model is lightweighted to meet the real-time requirements of the actual cultivation and picking scenarios.

[0028] In the third aspect, the present application proposes a lightweight target maturity detection system, which includes:

[0029] A data acquisition module, configured to collect real-time monitoring data of the target to be detected;

[0030] The model building module is configured based on the YOLOv8 network structure. It uses a spatial depth decomposition convolution module and a fast spatial pyramid pooling module combined with a large separable kernel attention mechanism to improve the backbone and neck network of the YOLOv8 model to form the SL-YOLO model. It also uses a layer adaptive amplitude pruning strategy to build the SLP-YOLO model.

[0031] The model training module is configured for training the SLP-YOLO model;

[0032] The maturity detection module is configured to input real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target to be detected.

[0033] The maturity detection module includes a visualization detection unit built using PySide6 to achieve real-time display of the target maturity level.

[0034] Compared with the prior art, the beneficial results of the present invention are as follows:

[0035] (1) To address the need for lightweight and efficient strawberry maturity detection in complex natural environments, this invention makes targeted improvements based on YOLOv8n: by replacing most of the original convolutions with spatial-to-depth convolutions, efficient pooling-free downsampling is achieved, significantly improving the feature loss problem that is easily caused by traditional convolution downsampling; by fusing large separable kernel attention with SPPF, the receptive field is effectively expanded to capture global information, while suppressing background noise interference, solving the problem of missed detection of target strawberry maturity in complex environments; by combining layer adaptive amplitude pruning strategy to perform channel pruning on the improved model, the model size, number of floating-point calculations and number of parameters are greatly reduced while maintaining detection accuracy, and finally a fast strawberry maturity detection model with low computational cost, low complexity and high accuracy is formed, meeting the lightweight requirements of automated detection.

[0036] (2) The experimental data of this application verify the performance advantages of the present invention: the improved model has an accuracy of 93.40%, an average accuracy of 86.00%, and a detection frame rate of 344.83. Compared with existing models such as Faster-RCNN, YOLOv5n, YOLOv9t, and YOLOv11n, it performs better in key indicators such as accuracy, detection frame rate, and floating-point computation. Compared with the basic model YOLOv8n, its accuracy, average accuracy, and detection frame rate are improved by 2.5, 1.0, and 24.1 percentage points, respectively, while the number of floating-point computations, the number of parameters, and the model memory usage are reduced by 27.16%, 61.87%, and 59.02%, respectively, which fully demonstrates the dual breakthrough in accuracy and efficiency.

[0037] (3) This application builds a visualization detection system integrating SLP-YOLO based on PySide6. Through multi-modal design, it achieves the same detection effect as the improved model, improving the practicality and convenience of the detection process. This model can accurately identify strawberries of different maturity levels in complex environments and has good robustness. It provides strong technical support for the intelligent cultivation and harvesting of strawberries in real-world scenarios and promotes the development of automated detection technology in the strawberry planting industry. Attached Figure Description

[0038] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and, together with the description, serve to explain the principles of the invention. Many anticipated advantages of the embodiments and other embodiments of the invention will be readily recognized as they become better understood through reference to the following detailed description. Elements in the drawings are not necessarily to scale. The same reference numerals refer to corresponding similar parts.

[0039] Figure 1 This is a flowchart of a lightweight target maturity detection method according to an embodiment of this application;

[0040] Figure 2 This is a schematic diagram of the network structure of the SLP-YOLO model according to an embodiment of this application;

[0041] Figure 3 This is a structural diagram of the Spatial Depth Decomposition Convolutional Module (SPDConv) according to an embodiment of this application;

[0042] Figure 4 This is a structural diagram of the SPPF-LSKA module according to an embodiment of this application;

[0043] Figure 5 This is a schematic diagram of the LAMP pruning strategy according to an embodiment of this application;

[0044] Figure 6 These are partial strawberry ripeness images according to embodiments of this application;

[0045] Figure 7 This is a schematic diagram of a portion of the data-enhanced image of S0C-V14 according to an embodiment of this application;

[0046] Figure 8 These are comparison images of the detection effects of four models according to embodiments of this application on S0D-v14;

[0047] Figure 9 This is a framework diagram of a lightweight target maturity detection system according to an embodiment of this application;

[0048] Figure 10This is an architectural design diagram of a strawberry maturity visualization detection system according to an embodiment of this application;

[0049] Figure 11 This is a visualization effect diagram of the strawberry maturity visualization detection system according to an embodiment of this application; Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0051] refer to Figure 1 , Figure 1 A flowchart of a lightweight target maturity detection method according to an embodiment of this application is shown. As shown in the figure, the method includes the following steps:

[0052] S101. Collect real-time monitoring data of the target to be detected.

[0053] In some specific embodiments, the real-time monitoring data of the target to be detected is local data, including images, videos, and real-time detection data from cameras.

[0054] S102. Based on the YOLOv8 network structure, the backbone and neck networks of the YOLOv8 model are improved by using a spatial depth decomposition convolution module and a fast spatial pyramid pooling module with a large separable kernel attention mechanism to form an SL-YOLO model. The SLP-YOLO model is then constructed using a layer adaptive amplitude pruning strategy.

[0055] In some specific embodiments, combined with Figure 2 , Figure 2A schematic diagram of the network structure of the SLP-YOLO model according to an embodiment of this application is shown. As shown, the network architecture of the SLP-YOLO model mainly consists of four parts: Input, Backbone, Neck, and Output. The Input part is responsible for receiving real-time monitoring data of the target to be detected, such as an image of a strawberry, providing input data for the model. The Backbone part is used to perform preliminary feature extraction and multi-scale feature generation on the input real-time monitoring data. Specifically, the Backbone part first uses a convolutional layer (Conv) to perform preliminary feature extraction on the input real-time monitoring data, such as an image; then it sequentially passes through multiple Spatial Depth Decomposition Convolutional (SPDConv) modules and a feature fusion C2f module. The SPDConv module performs interval partitioning and channel concatenation operations on the input feature map to achieve pooling-free downsampling, thereby effectively preserving the detailed information of the target. The C2f module then performs feature fusion processing. The backbone network is connected to the SPPF-LSKA module at its end. This module integrates a large separable kernel attention mechanism with fast spatial pyramid pooling, which can enhance the multi-scale feature fusion capability and fine-grained feature perception capability. The neck network part concatenates and fuses the features of different scales output by the backbone network through the Concat layer, and then uses the Upsample layer to perform upsampling operation to achieve cross-scale feature fusion. At the same time, it combines the C2f and SPDConv modules to further process and transform the features. The output part outputs the target maturity detection results through the Detect layer, which includes the category of the detected target maturity and the target's location information, thereby realizing the identification and localization of target maturity.

[0056] In some specific embodiments, the output shows the detection results of strawberry maturity, which includes the detected maturity category and the location information of the strawberries, thereby realizing the identification and location of strawberry maturity. Specifically, the strawberry maturity categories include ripe and immature, or color-changing stage, flowering stage, green and red categories.

[0057] For details, please refer to Figure 3 , Figure 3The diagram illustrates the structure of the Spatial Depth Decomposition Convolutional (SPDConv) module according to an embodiment of this application. As shown, for an input feature map of size H×W×C, the SPDConv module first performs a point-separated partitioning operation, dividing the input feature map into four sub-maps of size H / 2×W / 2×C. These four sub-maps are then concatenated according to their corresponding channel dimensions to obtain a feature map of size H / 2×W / 2×4C, thereby achieving pooling-free downsampling and preserving more local details. To address the problem of lost fine-grained information caused by convolution stride and pooling layers, this application improves the network structure of the YOLOv8 model as follows: In the Backbone part, except for the first convolutional layer, all other 3×3 convolutional layers are replaced with SPDConv layers; in the Head part of the output part, all standard convolutional layers are also replaced with SPDConv layers. Figure 3 In this diagram, S represents the spatial size of the input feature map, C1 and C2 represent the number of channels in the input and output feature maps, respectively, (ij) represents the pixel position of the feature map, x and y represent the horizontal and vertical coordinate axes, respectively, and ⊕ and ⊂ represent the input and output feature map channels, respectively. These represent matrix addition and convolution operations, respectively.

[0058] Further, refer to Figure 4 , Figure 4 A structural diagram of a fast spatial pyramid pooling (SPPF-LSKA) module incorporating a large separable kernel attention mechanism, according to an embodiment of this application, is shown. Figure 4 As shown, the multi-level max pooling operation in the traditional SPPF structure easily leads to the dilution and loss of weak detail features of small targets, and its excessively large receptive field is detrimental to detection tasks that require focusing on detailed features. Therefore, this application improves SPPF by employing a large separable kernel attention (LSKA) mechanism. The improved module structure is shown below. Figure 4The SPPF-LSKA module is shown in the diagram. This module first performs convolution on the input to generate an intermediate feature map, then performs downsampling through three max pooling operations to reduce the feature map size while retaining key information. Next, the intermediate feature map and the pooled feature map are concatenated and fused along the channel dimension to form a fused feature map, which is then introduced into the LSKA module. LSKA decomposes the two-dimensional weight kernels of depthwise separable convolutions and depthwise separable dilated convolutions into cascaded one-dimensional horizontal and vertical convolutions. It first receives the fused feature map as input, and then processes it sequentially through different convolutional layers. Specifically, the model first uses a 1×(2d-1) depthwise separable convolutional layer (DW-Conv) to capture horizontal spatial features. Next, a (2d-1)×1 depthwise separable convolutional layer extracts vertical spatial features. The extracted horizontal and vertical spatial feature maps are then element-wise summed to form a orientation-aware feature map. This map is then passed through a 1×|k / d| depthwise dilated convolutional layer (DW-D-Conv) to expand the receptive field. Finally, a 1×1 convolutional layer is used to adjust the number of channels, yielding the output feature. Here, d and k represent hyperparameters. d controls the spatial size of the convolutional kernel and the receptive field. Adjusting d allows for flexible control of the model's ability to capture horizontal and vertical features. A larger d corresponds to a larger convolutional kernel, capturing features over a wider area; a smaller d focuses on local details. The value 'k' controls the receptive field size of the dilated convolution. A larger 'k' results in a wider receptive field, allowing the model to capture feature dependencies over longer distances. The LSKA module decomposes the two-dimensional weight kernels of depthwise separable convolutions and depthwise separable dilated convolutions into cascaded one-dimensional horizontal and vertical convolutions, performing a weighted transformation on the original features. This captures more spatial and channel-related information, improving the ability to perceive fine-grained dimensional information and thus enhancing the accuracy of micro-target detection. The efficiency of LSKA not only enhances the cross-scale feature fusion capability of the SPPF module but also provides efficient and robust algorithmic support for handling real-world strawberry maturity detection scenarios.

[0059] Continue to refer to Figure 5 , Figure 5FIG. 0 shows a schematic diagram of the layer-adaptive magnitude (LAMP) pruning strategy according to an embodiment of the present application. Considering the limitations of factors such as computing resources and storage space, it is crucial to achieve the lightweight of the model. Since the SPDConv module and the SPPF-LSKA module are adopted in the present application, the computing cost has increased. Therefore, it is necessary to reasonably compress the computing amount and the number of parameters of the model to achieve the lightweight goal while ensuring the overall accuracy of the model. There are usually certain redundant parameters between the layers of the neural network model, and these redundant parameters will affect the volume and accuracy of the model. In view of this, the present application uses the layer-adaptive sparsity for the magnitude-base pruning (LAMP) method to prune the model to remove the redundant parts in the network structure. The entire pruning process can be referred to Figure 5 , LAMP performs global pruning by minimizing the strategy of model output distortion, adaptively determines the pruning ratio according to the absolute value of the parameter weights of each layer, and automatically assigns appropriate pruning ratios to each layer. Specifically, the weights of all layers are globally sorted according to the LAMP score, and then pruning operations are performed on the weights with the lowest scores. Finally, the pruned model is fine-tuned to restore or even improve the network performance before pruning. When u < v, |W[u]| ≤ |W[v]| holds, and its calculation formula is as follows: In the above formula, u and v represent weight indices, and W[u] and W[v] respectively represent the mapped weights corresponding to u and v. The numerator (W[u]) 2 is the square value of the current weight, reflecting its absolute importance; the denominator ∑ v≥u (W[v]) 2 is the sum of the squares of all subsequent weights starting from the current weight, which is used to normalize the importance of the current weight. That is, the denominator represents the sum of the squares of all remaining connection weights in the same-layer connection, and the connections corresponding to the indices less than u will be pruned until the threshold corresponding to the global sparsity constraint is reached. It can be seen from this formula that the larger the square of the weight, the higher the LAMP score, and this score can effectively reflect the importance of each connection.

[0060] S103. Train the SLP-YOLO model.

[0061] S104. Input the real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target to be detected.

[0062] The present application applies a lightweight method for detecting the maturity of a target to the cultivation and picking of strawberries. ]

[0063] In some specific embodiments, reference is made to Figure 6 and Figure 7 , Figure 6 and Figure 7 The figures show partial strawberry ripeness images and partial data-enhanced images from S0C-V14, according to embodiments of this application. As shown, both the training and validation datasets used in this application are sourced from the Roboflow platform (https: / / universe.roboflow.com), a large-scale computer vision and large model training data resource platform. Specifically, Strawberry Object Detection-v14 (SOD-v14) serves as the model training dataset, containing 4405 strawberry images labeled with two categories: "ripe" and "unripe." StrawberryClassification-v4 (SC-v4) serves as the dataset for validating the model's generalization ability, containing 1499 strawberry images labeled with four categories: turning stage, flowering stage, green, and red. These two datasets are specifically designed for strawberry picking and recognition scenarios and are open datasets. Their core objective is to improve the automated and efficient recognition capability of small targets in complex natural environments. To address the challenges of identifying strawberry ripeness due to variations in fruit growth posture and the imbalanced distribution of the dataset, this application selects pre-processed data versions from the Roboflow platform to enhance model generalization ability by increasing data diversity. All dataset images are uniformly resized to 640×640 pixels, and a multi-dimensional data augmentation strategy is employed, generating three augmented versions for each original sample. Augmentation operations for SOD-v14 include horizontal / vertical flipping, 90° and 180° rotation, grayscale conversion, and saturation and brightness adjustments; augmentation operations for SC-v4 include horizontal / vertical flipping and saturation adjustments. Targeted data augmentation highlights the detailed features of strawberry ripeness images and strengthens the network model's learning ability for different categories of data features (especially small object features). After augmentation, the SOD-v14 and SC-v4 datasets contain 8347 and 4296 images respectively, both divided into training, validation, and test sets in an 8:1:1 ratio, providing a standardized data foundation for model training and performance evaluation.

[0064] In some specific embodiments, the training parameters for training the SLP-YOLO model are set as follows: input image size is 640 pixels × 640 pixels, training batch size is 100, batch size is 16, initial learning rate is 0.01, initial momentum is 0.937, confidence threshold and IoU threshold for maturity detection results are set to 0.5 and 0.7 respectively; the pruning rate of the layer adaptive amplitude pruning strategy is 50%.

[0065] In some specific embodiments, to comprehensively evaluate the performance of the improved SLP-YOLO model, this application employs multi-dimensional verification methods, including ablation experiments, pruning experiments, comparative experiments, comparison experiments of recognition effects before and after improvement, and generalization experiments, to analyze the model from the two core dimensions of detection accuracy and detection efficiency. Key evaluation metrics selected during the evaluation process include: precision (P), mean average precision (mAP) calculated within the Intersection over Union (IoU) threshold range of 0.5 to 0.95, frame rate per second (FPS), and giga-floating-point operations (GFLOPS), etc., to quantitatively characterize the model's overall performance in terms of target recognition accuracy, real-time processing capability, and computational complexity.

[0066] Table 1 Comparison of Fusion Performance of Different Modules

[0067]

[0068] In the ablation experiments, as shown in Table 1, replacing most of the original convolutions with SPDConv resulted in an increase in precision by 1.8 percentage points and average precision by 1.2 percentage points, despite the increased floating-point computation and parameter count due to the increased number and size of convolutions. Combining the SPPF module with the LSKA module in the backbone network to form the SPPF-LSKA module resulted in a slight increase in GFLOPS and parameter count due to the pyramid structure splicing multi-scale features, but improved precision by 2.1 percentage points and average precision by 1.6 percentage points. When replacing SPDConv and SPPF-LSKA simultaneously in the corresponding positions of YOLOv8n, with GFLOPS increasing to 11.8 and parameter count increasing to 5.01×106, precision improved by 2.5 percentage points and average precision by 1.2 percentage points. In summary, after the fusion of these two improvements, the improved YOLOv8n model exhibits better detection performance compared to the unimproved original model, but at the cost of increased model size.

[0069] Table 2 Results of the pruning experiment

[0070]

[0071] In the pruning experiment, SLP-YOLO was the model obtained by pruning and optimizing the unpruned SL-YOLO model. The 100% pruning percentage was calculated based on the ratio of computational cost before and after pruning, and the pruning parameter values ​​were also set according to this ratio. This pruning strategy aims to reduce complexity and computational resource requirements while maintaining high detection performance. The results of the fine-tuning experiment after pruning are shown in Table 2. As the pruning rate gradually increases, the number of parameters, computational cost, and size of the SLP-YOLO model decrease significantly. When the pruning rate is 50%, the accuracy is the highest, with an average accuracy improvement of 0.9 percentage points, an FPS improvement of 67.05 percentage points, and a reduction in floating-point calculations, parameter count, and model memory usage, resulting in a significant decrease in model computational cost. When the pruning rate is 90%, although the model compression effect is more significant, its detection performance is relatively poor compared to the model with a 50% pruning rate. The principle of model pruning is to achieve maximum model volume compression while minimizing the loss of detection accuracy. Based on the results of channel pruning, a pruning rate of 50% was ultimately selected as the optimal choice for the SL-YOLO model. After reasonable pruning, the number of channels in the backbone and neck networks decreased significantly, demonstrating the significant effect of the LAMP pruning strategy on channel reduction. The pruning process prioritizes compressing layers with redundant intermediate feature representation capabilities, effectively eliminating inefficient channels. In this process, the model becomes more efficient and lightweight, with almost no loss in detection performance, fully demonstrating the efficiency of the pruning strategy in optimizing the SL-YOLO network structure and validating the practicality of the proposed method.

[0072] Table 3. Performance Comparison Before and After Model Improvement

[0073]

[0074]

[0075] As shown in Table 3, SLP-YOLO, compared to the baseline model (YOLOv8n), not only possesses fast and efficient detection capabilities but also reduces complex computational costs. After channel pruning, the difference in channel structure compared to the baseline led to a decline in detection performance during the optimization process. After fine-tuning, the model's detection performance almost recovered to the level before pruning; that is, while achieving model lightweighting, the overall high-efficiency detection performance of the model was maintained. Ultimately, SLP-YOLO provides a solution that balances high performance and lightweight design for strawberry growth maturity detection. Overall, compared to the baseline model, our proposed scheme achieves a reduction of 2.61x, 1.37x, and 2.44x in parameter count, computational cost, and model size, respectively, verifying the practicality of this lightweight scheme (channel pruning).

[0076] Table 4 Comparison of Detection Performance of Different Networks

[0077]

[0078] In the comparative experiments, to verify the superiority of the proposed SLP-YOLO model, under the same experimental environment, SLP-YOLO was compared with several widely used models in the field of object detection, including Fast-RCNN, SSD, YOLOv5n, YOLOv8n, YOLOv9t, YOLOv11n, and YOLOv12n. The experimental results are shown in Table 4. The proposed SLP-YOLO algorithm achieves the best accuracy and detection frame rate, while having the lowest number of floating-point calculations, parameters, and model memory usage, indicating that the improved model achieves lightweight and efficient detection of strawberry maturity. Compared with the YOLO series algorithms in Table 4, Faster R-CNN and SSD, due to their higher model complexity and slower inference speed, are not suitable for the real-time detection of strawberry maturity. The YOLO-based detection algorithms all achieve an average accuracy of 85% or higher, and their computational cost and real-time detection advantages are significantly superior. Compared to the baseline model YOLOv8n, the improved model in this application achieves an accuracy, average precision, and detection frame rate improvement of 2.5, 1, and 24.1 percentage points, respectively. Furthermore, compared to Faster R-CNN, SSD, YOLOv5n, YOLOv9t, YOLOv11n, and YOLOv12n, it outperforms the comparison models in terms of accuracy, average precision, and detection frame rate. Moreover, it has the lowest parameter count, floating-point computation cost, and model memory usage, demonstrating significant comprehensive performance and lightweight advantages.

[0079] In the comparison experiment of recognition performance before and after the improvement, in order to comprehensively evaluate the validation results of the SPL-YOLO model, the detection performance of YOLOv8n, YOLOv11n and the SPL-YOLO model of this application in the detection of strawberry maturity was compared. Tests were conducted under different light intensities, shooting angles, and multiple target conditions. The confidence threshold was set to 0.5 and the IOU threshold to 0.7 to better reflect real-world conditions. The results are as follows: Figure 8As shown in the diagram. In experimental scenario A, the confidence level of the SLP-YOLO model was generally higher than that of the other three models. In experimental scenario B, YOLOv11n misclassified some strawberry fruits. In scenario C, YOLOv5n, YOLOv8n, and YOLOv11n all had one or two missed strawberry targets. In scenario D, YOLOv5n, YOLOv8n, and YOLOv11n failed to detect and label a single strawberry target in the image. In the four experimental scenarios (A, B, C, and D), the improved model of this application had a higher overall confidence level and detected more small targets, indicating that the improved model performed better in detecting small targets, further enhancing its reliability in practical applications and providing more technical references for subsequent research in the field of intelligent strawberry applications.

[0080] Table 5 shows the performance comparison of the model before and after improvement on the SC-v4 dataset.

[0081]

[0082] In the generalization experiment, the natural environment presents certain complexities, such as lighting conditions, fruit growth variations, and occlusion issues. The training data used may not cover all possibilities. To verify the generalization ability of the SLP-YOLO model, this application uses the SC-v4 dataset for generalization experiments. Under the premise of ensuring the algorithm's operating environment is consistent with the above experiments, the experimental results are shown in Table 5. As shown in Table 5, compared to the base model YOLOv8n, the proposed SLP-YOLO model improves recall, average precision, and detection frame rate by 0.9%, 1.7%, and 12.5%, respectively; simultaneously, the computational cost is significantly reduced, such as a 27.16% reduction in floating-point computation and a 63% reduction in the number of model parameters. This indicates that the SLP-YOLO algorithm has superior generalization ability in strawberry maturity detection.

[0083] Continue to refer to Figure 9 As an implementation of the above method, in a second aspect, this application provides an embodiment of a lightweight target maturity detection system framework diagram 900, which is similar to... Figure 1 Corresponding to the illustrated method embodiment, this system can be specifically applied to various electronic devices. The system 900 includes a data acquisition module 901, a model building module 902, a model training module 903, and a maturity detection module 904, which are interconnected, wherein:

[0084] The data acquisition module 901 is configured to collect real-time monitoring data of the target to be detected;

[0085] The model building module 902 is configured to improve the backbone and neck networks of the YOLOv8 model based on the YOLOv8 network structure by using a spatial depth decomposition convolution module and a fast spatial pyramid pooling module that combines a large separable kernel attention mechanism to form an SL-YOLO model. The SLP-YOLO model is constructed by using a layer adaptive amplitude pruning strategy.

[0086] Model training module 903 is configured for training the SLP-YOLO model;

[0087] The maturity detection module 904 is configured to input the real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target to be detected.

[0088] In some specific embodiments, the maturity detection module includes a visualization detection unit built using PySide6 to achieve real-time display of the target maturity level.

[0089] Specifically, to simplify the strawberry ripeness detection process and allow inspectors to intuitively judge fruit ripeness, this application builds a visual detection interface based on PySide6 and integrates the SLP-YOLO algorithm to construct a strawberry ripeness visualization detection system. Compared to traditional PyQt, PySide has a wider licensing scope. (Reference) Figure 10 and Figure 11 , Figure 10 and Figure 11 The framework diagram and visualization effect diagram of the strawberry maturity visualization detection system according to embodiments of this application are shown respectively. Figure 10As shown, the strawberry maturity visualization detection system includes a detection algorithm module, an input source module, a parameter setting module, and a visualization module. The detection algorithm uses the SLP-YOLO algorithm, which is the core algorithm for strawberry maturity detection. It is responsible for analyzing and processing input strawberry images or videos to determine the maturity of the strawberries. The input source provides multiple data input methods, including images, videos, and real-time camera acquisition, meeting the needs for strawberry image data acquisition in different scenarios. The parameter setting module includes confidence threshold and IOU (Intersection over Union) parameter setting functions. Detection personnel can flexibly adjust these parameters according to actual detection needs and scenario characteristics to optimize the performance of the detection algorithm and the accuracy of the detection results. The visualization module presents the results processed by the detection algorithm to the detection personnel in an intuitive and clear way by displaying detection effects, heatmaps, and related information, helping them quickly and accurately determine the maturity of the strawberries. This detection system is designed by leveraging the advantages of sub-modules, improving the reusability and development efficiency of the overall architecture, reducing maintenance costs, and simplifying program design. The first step is the detection algorithm module, which imports weight files to select the corresponding detection algorithm. Secondly, there's the input source module, which acquires local data, including images, videos, and real-time camera detection data. Thirdly, there's the parameter setting module, which allows filtering of detection results by adjusting confidence and IOU thresholds. Finally, there's the visualization module, which includes detection result images and heatmaps. The strawberry maturity visualization detection effect is shown below. Figure 11 As shown.

[0090] This application firstly improves the ability to capture detailed information of the target during feature extraction by introducing SPDconv to replace most of the basic convolutional layers on the YOLOv8n architecture, retaining only one original convolutional layer in the backbone network. Secondly, it combines Large Separable Kernel Attention (LSKA) with Fast Spatial Pyramid Pooling (SPPF) to obtain the SPPF-LSKA module, replacing the original SPPF module in the backbone structure, thereby enhancing the receptive field and improving the multi-scale information fusion capability. Finally, for model lightweighting, the LAMP pruning strategy is used to compress the optimized model, reducing the number of parameters and the model size. Based on this, a visualization detection system is built using PySide6, integrating the model detection interface to achieve real-time display of strawberry maturity detection. Experimental results show that the lightweight SLP-YOLO improves precision (P), mean accuracy (mAP), and detection frame rate (FPS) by 2.5%, 0.9%, and 24.1% on the SOD-v14 dataset, respectively, and improves recall, mean accuracy, and detection frame rate by 0.9%, 1.7%, and 12.5% ​​on the SC-v4 dataset, respectively. Simultaneously, the number of model parameters is reduced by approximately 60%, and GFLOPS are reduced by 27.16%. This achieves a balance between performance improvement and lightweight design.

[0091] Although the principles of the present invention have been described in detail above with reference to preferred embodiments, those skilled in the art should understand that the above embodiments are merely illustrative explanations of the implementation of the present invention and are not intended to limit the scope of the present invention. The details in the embodiments do not constitute a limitation on the scope of the present invention. Any obvious changes, such as equivalent transformations or simple substitutions, based on the technical solutions of the present invention without departing from the spirit and scope of the present invention fall within the protection scope of the present invention.

Claims

1. A lightweight target maturity detection method, characterized by, The method includes: S1. Collect real-time monitoring data of the target to be detected; S2. Based on the YOLOv8 network structure, use a spatial depthwise separable convolution module and a fast spatial pyramid pooling module combined with a large separable kernel attention mechanism to improve the backbone network and neck network of the YOLOv8 model to form the SL-YOLO model, and use a layer adaptive magnitude pruning strategy to construct the SLP-YOLO model; S3. Train the SLP-YOLO model; S4. Input the real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target to be detected.

2. The light-weight target maturity detection method according to claim 1, characterized in that, The backbone network of the SL-YOLO model includes 1 convolutional layer, several spatial depthwise separable convolution modules, several feature fusion modules, and the fast spatial pyramid pooling module combined with the large separable kernel attention mechanism located at the end of the backbone network.

3. The method of claim 1, wherein The neck network of the SL-YOLO model includes feature fusion modules, upsampling layers, spatial depthwise separable convolution modules and connection layers, and several spatial depthwise separable convolution modules are provided in the detection head of the SL-YOLO model.

4. The method of claim 1, wherein The spatial depthwise separable convolution module realizes unpooled downsampling by performing pointwise division and channel splicing operations on the input feature map, and retains the target detail information in the real-time monitoring data of the target to be detected.

5. The method of claim 1, wherein The construction method of the fast spatial pyramid pooling module combined with the large separable kernel attention mechanism includes: Step a. Perform an initial convolution operation on the input feature map to generate an intermediate feature map, perform a maximum pooling operation on the intermediate feature map, and splice the intermediate feature map and the pooled feature map in the channel dimension to form a fused feature map; Step b. Use depthwise separable convolutional layers to capture the spatial features in the horizontal and vertical directions of the fused feature map respectively to form a direction-aware feature map; and use a depthwise separable dilated convolutional kernel to perform convolution on the direction-aware feature map to expand the feature perception range; Step c. Adjust the number of channels by a convolutional layer and output the processed feature map.

6. The method of claim 1, wherein When the layer adaptive magnitude pruning strategy performs pruning on the SL-YOLO model, it adaptively allocates pruning ratios according to the absolute values of the weights of the parameters of each layer of the SL-YOLO model.

7. The lightweight target maturity detection method according to claim 6, characterized in that, The layer adaptive magnitude pruning strategy also includes fine-tuning the pruned model, and the layer adaptive magnitude pruning strategy calculates the pruning score according to the following formula and prunes the weight with the lowest score; when u < v is satisfied, |W[u]| ≤ |W[v]| holds, and its calculation formula is as follows: In the formula, y and v represent weight indexes; W[u] and W[v] represent the mapping weights corresponding to u and v respectively; the numerator (W[u]) 2 represents the square value of the current weight, reflecting its absolute importance; the denominator v≥u (W[v]) 2 represents the sum of squares of all subsequent weights starting from the current weight, used for normalizing the importance of the current weight.

8. An application of a lightweight target maturity detection method, characterized in that, Apply the lightweight target maturity detection method according to any one of claims 1-7 to the cultivation and picking of strawberries.

9. The application of the lightweight target maturity detection method according to claim 8, characterized in that, The training parameters in step S3 are set as follows: the input image size is \(640\) pixels \(\times 640\) pixels, the training batch is \(100\), the batch size is \(16\), the initial learning rate is \(0.01\), the initial momentum is \(0.937\), the confidence threshold and IoU threshold of the maturity detection result are set to \(0.5\) and \(0.7\) respectively; the pruning rate of the layer adaptive magnitude pruning strategy is \(50\%\).

10. A lightweight target maturity detection system, characterized in that, The system includes: The data acquisition module is configured to collect real-time monitoring data of the target to be detected; The model building module is configured based on the YOLOv8 network structure. It uses a spatial depth decomposition convolution module and a fast spatial pyramid pooling module combined with a large separable kernel attention mechanism to improve the backbone network and neck network of the YOLOv8 model to form an SL-YOLO model. It also uses a layer adaptive amplitude pruning strategy to build an SLP-YOLO model. The model training module is configured to train the SLP-YOLO model; The maturity detection module is configured to input the real-time monitoring data of the target to be detected into the trained SLP-YOLO model to obtain the maturity detection result of the target to be detected. The maturity detection module includes a visualization detection unit built using PySide6 to achieve real-time display of the target maturity level.

Citation Information

Patent Citations

  • Strawberry maturity detection method and system based on improved YOLOv8

    CN117315473A