Multi-layer model pomegranate growth cycle detection method

Through the multi-level model method, combined with the improved YOLOv8n backbone network, Triplet Attention and SlimLGNeck structure, and combined with knowledge distillation technology, the pomegranate growth cycle detection model has been solved, and efficient detection on low-computing equipment is achieved.

CN120564048APending Publication Date: 2025-08-29WUXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510706238.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing pomegranate growth cycle detection model has high computational complexity, making it difficult to deploy on resource-constrained devices, has limited detection accuracy, insufficient feature extraction capabilities, and lacks lightweight optimization and knowledge distillation strategies.

Method used

A multi-level model method is adopted, including improved YOLOv8n backbone network, Triplet Attention attention mechanism, SlimLGNeck structure and knowledge distillation technology, feature extraction is performed through the Adown module, combined with Triplet Attention and SlimLGNeck structure for feature fusion, and optimized detection heads using Feature-CWD distillation to achieve lightweight and high-precision detection.

Benefits of technology

While reducing the computational complexity, it improves detection accuracy and feature extraction capabilities, so that pomegranate growth cycle detection can be efficiently deployed on low-computing equipment, providing efficient, low-cost and easy-to-deploy solutions for intelligent agricultural monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564048A_ABST
    Figure CN120564048A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-layer model pomegranate growth cycle detection method, and relates to the technical field of computer vision and intelligent agricultural technologies. In order to solve the technical problem that existing pomegranate growth cycle detection is insufficient in feature extraction capability, the technical scheme provided by the invention comprises the following steps: acquiring an image data set containing different growth stages of pomegranates to obtain a training image; inputting the training image into a backbone network of a YOLOv8n model, and outputting a primary feature map with enhanced feature expression; inputting the primary feature map into an SPPF module to obtain an attention-enhanced intermediate feature map; inputting the intermediate feature map into a neck network, and performing cross-scale feature fusion to obtain an efficiently fused advanced feature map; and inputting the advanced feature map into a detection head optimized by knowledge distillation to complete target detection and identification of different growth cycles of pomegranate. The method can be widely applied to application scenes such as intelligent agriculture, orchard management, agricultural Internet of Things and unmanned aerial vehicle inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and intelligent agricultural technology, and specifically to a pomegranate growth cycle detection method based on a multi-level model. Background Art

[0002] In recent years, the rapid development of computer vision and deep learning technologies has made intelligent agricultural monitoring a research hotspot. For applications in agriculture, such as fruit detection, crop growth monitoring, and pest and disease identification, researchers have proposed a variety of deep learning-based object detection algorithms. Among them, the YOLO (You Only Look Once) family of algorithms has been widely used in fruit detection and classification tasks due to its high efficiency and accuracy.

[0003] Currently, many studies have explored the growth cycle detection of different crops. For example: Camellia oleifera fruit testing Zhu et al. proposed an improved lightweight model (YOLO-LM) based on YOLOv7-tiny to improve the efficiency of camellia fruit maturity detection. The model optimizes the feature extraction module and reduces computational complexity, enabling it to run efficiently in resource-constrained environments.

[0004] Dragon fruit testing Nan et al. proposed the WGB-YOLO network to address the challenge of detecting dragon fruit in dense orchard environments. This method, which utilizes a feature enhancement module and an improved feature fusion strategy, achieves a mean average precision (MAP) of 86.0%, outperforming other models while also offering faster detection speed.

[0005] Grape maturity testing Chen et al. developed an improved algorithm (ESP-YOLO) based on YOLOv5s specifically for ripeness detection of fresh table grapes. On embedded devices, the model achieved a mean average prediction accuracy (MAP) of 98.3% and performed well in real-time detection tasks.

[0006] Mango picking point positioning Li et al. proposed an improved YOLOv8 model that combines object detection and instance segmentation techniques for an automated mango picking system. This method achieved fruit and stem detection accuracies of 98.9% and 97.1%, respectively, and was able to accurately locate the optimal picking point (with a 92.01% success rate).

[0007] Winter jujube detection and positioning Yu et al. proposed the MLG-YOLO model, which combined lightweight design and three-dimensional positioning technology to achieve efficient winter jujube detection and positioning. Experimental results showed that the mAP was 92.70%, the detection accuracy reached 100% in the laboratory environment, and 92.82% in the orchard environment.

[0008] Tea production estimates Li et al. proposed a tea bud counting method based on target detection using YOLOv5, and combined it with Hungarian matching and Kalman filtering algorithms to achieve efficient tea yield estimation. The average accuracy on the test set reached 91.88%, and the counting results were highly correlated with manual counting (R²=0.98).

[0009] Although the above research has made some progress in fruit detection and crop monitoring, there is still limited research on automated detection of the pomegranate growth cycle. For example: TP-YOLO (Du et al.): Through lightweight design (ShuffleNetV2 + SE attention mechanism), it reduces the amount of computation while maintaining detection accuracy (mAP 94.4%).

[0010] PG-YOLO (Wang et al.): Introduces depthwise separable convolution and multi-head self-attention (MHSA), achieving a mAP of 93.4% in small pomegranate fruit detection and an 8% increase in detection speed.

[0011] YOLO-Granada (Zhao et al.): A lightweight detection algorithm based on the optimization of YOLOv5, with a detection accuracy of 92.2% and a 54.7% reduction in model parameters.

[0012] However, existing pomegranate growth cycle detection still faces the following key problems: The model is complex and difficult to deploy on resource-constrained devices Existing models are mostly based on standard YOLO or its variants, which have high computational complexity and are difficult to meet the needs of real-time agricultural applications. For example, while YOLOv8 offers high accuracy, it has high computational overhead and cannot run efficiently on low-computing devices such as smart agricultural monitoring terminals, drones, and IoT edge devices.

[0013] Detection accuracy is limited, making it difficult to accurately identify the growth cycle The pomegranate growth cycle includes multiple key stages (bud, early fruiting, flowering, mid-growth, and maturity), but existing methods fail to fully optimize model structures to achieve high-precision detection in complex orchard environments. For example, existing research primarily focuses on fruit detection while ignoring the precise delineation of pomegranate growth stages, resulting in limited classification accuracy.

[0014] Insufficient feature extraction capabilities and low attention to key features Traditional YOLO models are susceptible to occlusion and lighting changes in complex backgrounds, leading to object detection errors. For example, models like YOLOv5 can sometimes misidentify leaves or branches as fruit, affecting the stability of detection results.

[0015] Lack of lightweight optimization and knowledge distillation strategies Existing research mostly uses standard convolutional neural network (CNN) architectures, but rarely incorporates lightweight optimization and knowledge distillation techniques to improve model efficiency. Even lightweight YOLO variants, such as YOLO-LM and YOLO-Granada, fail to fully utilize distillation strategies to optimize detection performance, resulting in a poor balance between detection accuracy and computational efficiency. Summary of the Invention

[0016] In order to solve the technical problems faced by the existing pomegranate growth cycle detection, such as high model complexity, limited detection accuracy, and insufficient feature extraction capability, the present invention provides the following technical solutions: The multi-level model pomegranate growth cycle detection method includes the following steps: Step 1: collect and preprocess an image dataset containing pomegranates at different growth stages to obtain standardized training images; Step 2: Input the training image obtained in step 1 into the backbone network of the improved YOLOv8n model, perform feature extraction and downsampling through the Adown module, and output a primary feature map with enhanced feature expression; Step 3: Input the primary feature map outputted in step 2 into the SPPF module containing the Triplet Attention mechanism to obtain an attention-enhanced intermediate feature map. Step 4: Input the intermediate feature map output in step 3 into the neck network containing the SlimLGNeck structure to perform cross-scale feature fusion to obtain an efficiently fused high-level feature map; In step 5, the high-level feature map output in step 4 is input into the detection head optimized by knowledge distillation to complete the target detection and recognition of pomegranates in different growth cycles.

[0017] Furthermore, a preferred embodiment is provided, in which the Adown module uses a combination of average pooling and maximum pooling to perform downsampling, thereby reducing the number of model parameters and improving feature extraction capabilities.

[0018] Furthermore, a preferred embodiment is provided, in which the Triplet Attention mechanism includes three parallel branches: channel attention, spatial attention, and sequence attention, which are used to enhance the model's attention to key features.

[0019] Furthermore, a preferred embodiment is provided, in which the SlimLGNeck structure includes a GSConv module and a VoVGSCSP module, the GSConv module is used to reduce the amount of calculation, and the VoVGSCSP module is used to enhance the cross-level feature fusion effect.

[0020] Furthermore, a preferred embodiment is provided, in which the GSConv module performs calculations by combining group convolution with spatial convolution to reduce computational complexity and hardware resource requirements.

[0021] Furthermore, a preferred implementation is provided, in which the knowledge distillation adopts Feature-CWD distillation to transfer feature information from the teacher model to the student model to optimize the performance of the student model.

[0022] Based on the same inventive concept, the present invention also provides a pomegranate growth cycle detection device based on a multi-level model integrating Triplet Attention and knowledge distillation, including the following modules: Module 1: Collect and preprocess a dataset of images of pomegranates at different growth stages to obtain standardized training images; Module 2: Input the training image obtained in module 1 into the backbone network of the improved YOLOv8n model, perform feature extraction and downsampling through the Adown module, and output a primary feature map with enhanced feature expression; Module three: Input the primary feature map output by module two into the SPPF module containing the Triplet Attention mechanism to obtain the intermediate feature map enhanced by attention; In module 4, the intermediate feature map output by module 3 is input into the neck network containing the SlimLGNeck structure to perform cross-scale feature fusion to obtain an efficiently fused high-level feature map; In module five, the high-level feature map output by module four is input into the detection head optimized by knowledge distillation to complete target detection and recognition of pomegranates in different growth periods.

[0023] Based on the same inventive concept, the present invention also provides a computer storage medium for storing a computer program. When the computer program is read by a computer, the computer executes the method described.

[0024] Based on the same inventive concept, the present invention also provides a computer, comprising a processor and a storage medium. When the processor reads the computer program stored in the storage medium, the computer executes the method described above.

[0025] Based on the same inventive concept, the present invention also provides a computer program product, which is a computer program. When the computer program is executed, the method described above is implemented.

[0026] Compared with the prior art, the technical solution provided by the present invention is beneficial in that: The present invention optimizes the backbone network of the YOLOv8n model by introducing the Adown module, while reducing the computational complexity of the model, maintaining or even improving the detection accuracy. Compared with the traditional stride convolution, the Adown module combines the average pooling and maximum pooling strategies, so that key information can be retained during the downsampling process, thereby enhancing the feature expression capability. Compared with existing research, such as YOLO-LM, which mainly reduces the amount of calculation through pruning and channel reduction, the present invention adopts the Adown module in a more stable way, avoiding the problem of accuracy reduction that may be caused by pruning, so that the lightweight model still has high detection accuracy in a low computing power environment.

[0027] Introducing the Triplet Attention mechanism before the SPPF module effectively improves the model's focus on key features and enhances its ability to recognize small objects and objects in complex backgrounds. Compared to the SE attention mechanism or CBAM, Triplet Attention calculates attention weights based on three dimensions: channel, width, and height. This enables the model to accurately identify pomegranates at all growth stages from different viewing angles. Compared to TP-YOLO's SE attention mechanism, this solution's Triplet Attention mechanism has a stronger ability to perceive objects of different scales, avoids the problem of local feature neglect that can occur with the SE module, and improves overall detection accuracy.

[0028] The SlimLGNeck structure is used to optimize the neck network of YOLOv8n. Combining the GSConv and VoVGSCSP modules not only reduces the computational complexity but also optimizes the feature fusion capability. GSConv effectively reduces the amount of computation by introducing grouped convolution and spatial convolution, while the VoVGSCSP module enhances cross-level feature fusion, allowing features of different scales to be efficiently transferred. Compared to the PG-YOLO optimization method for the backbone network in existing research, the SlimLGNeck structure of the present invention focuses more on improving computational efficiency. While reducing the amount of model computation, it does not affect the recognition effect of different growth stages, thereby increasing the detection speed while maintaining a high mAP.

[0029] The fusion of knowledge distillation technology (Feature-CWD) optimizes the feature transfer between the teacher model and the student model, achieving a lightweight model with high detection capabilities even in low-computing resource environments. Traditional knowledge distillation methods are mostly based on soft labels (Soft Targets) for optimization. However, this solution adopts the Feature-CWD distillation method, which enables the student model to efficiently learn the feature information of the teacher model while maintaining a low number of parameters and computational complexity. In experiments comparing distillation methods, Distill-YOLO after Feature-CWD distillation improved the mAP50-95 indicator by 1.86% compared to the original YOLOv8n, while reducing the computational complexity by approximately 14.6%. Compared with other distillation methods (such as BCKD and Logical-L2), it achieves a better balance between model compression and performance optimization.

[0030] Overall, the present invention is superior to existing technologies in terms of computational complexity, detection accuracy, lightweight optimization, and target feature extraction, enabling the pomegranate growth cycle detection system to be efficiently deployed on low-computing power devices, providing an efficient, low-cost, and easy-to-deploy solution for smart agricultural monitoring.

[0031] It can be widely used in application scenarios such as smart agriculture, orchard management, agricultural Internet of Things and drone inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a schematic diagram of the SlimLGNeck structure; Figure 2 This is the Triplet Attention mechanism diagram. DETAILED DESCRIPTION

[0033] In order to make the advantages and benefits of the technical solution provided by the present invention more clearly reflected, the technical solution provided by the present invention is now further described in detail with reference to the accompanying drawings, specifically: Embodiment 1: This embodiment provides a multi-level model pomegranate growth cycle detection method, comprising the following steps: Step 1: collect and preprocess an image dataset containing pomegranates at different growth stages to obtain standardized training images; Step 2: Input the training image obtained in step 1 into the backbone network of the improved YOLOv8n model, perform feature extraction and downsampling through the Adown module, and output a primary feature map with enhanced feature expression; Step 3: Input the primary feature map outputted in step 2 into the SPPF module containing the Triplet Attention mechanism to obtain an attention-enhanced intermediate feature map. Step 4: Input the intermediate feature map output in step 3 into the neck network containing the SlimLGNeck structure to perform cross-scale feature fusion to obtain an efficiently fused high-level feature map; In step 5, the high-level feature map output in step 4 is input into the detection head optimized by knowledge distillation to complete the target detection and recognition of pomegranates in different growth cycles.

[0034] The Adown module uses a combination of average pooling and maximum pooling to perform downsampling, reducing the number of model parameters and improving feature extraction capabilities.

[0035] The Triplet Attention mechanism includes three parallel branches: channel attention, spatial attention, and sequence attention, which are used to enhance the model's attention to key features.

[0036] The SlimLGNeck structure includes a GSConv module and a VoVGSCSP module. The GSConv module is used to reduce the amount of calculation, and the VoVGSCSP module is used to enhance the cross-level feature fusion effect.

[0037] The GSConv module uses a combination of group convolution and spatial convolution to perform calculations to reduce computational complexity and hardware resource requirements.

[0038] The knowledge distillation adopts the Feature-CWD distillation method to transfer feature information from the teacher model to the student model to optimize the performance of the student model.

[0039] Implementation Method 2: This implementation method further explains the technical solution provided in Implementation Method 1 in detail. Specifically: 1. Data Collection and Preprocessing In this solution, the first step is to build a high-quality dataset containing pomegranates at different growth stages to ensure that the subsequent object detection model has good generalization ability.

[0040] (1) Data collection: In a natural orchard environment, high-resolution cameras, drones, and fixed monitoring equipment were used to obtain images of pomegranates at different growth stages, including buds, early fruiting, flowers, mid-growth, and mature fruits. To improve the representativeness of the data, data collection should cover different lighting conditions (sunny, cloudy, and night), different background complexities (fruit tree leaves, branches, and weeds), and different shooting angles (looking down, looking straight, and looking up).

[0041] (2) Data cleaning: Screen the collected images and remove blurred, overexposed, or noisy images to ensure the high quality of the input data and lay the foundation for subsequent model training.

[0042] (3) Data enhancement: The dataset is standardized, including normalized size (640×640), color balancing, contrast adjustment, random cropping, rotation, flipping, etc., to enhance the robustness of the model so that it can adapt to the target detection task in the complex orchard environment.

[0043] (4) Data labeling: Use YOLO or COCO format to label the data, and use manual labeling tools to select pomegranate fruits at different growth stages to ensure that the training data has clear labels and improve the accuracy of target detection.

[0044] After data collection and preprocessing, a structured, high-quality dataset is obtained, which will be used as input for model training to improve the target detection model's ability to recognize the pomegranate growth cycle.

[0045] 2. YOLOv8n model optimization and lightweighting After obtaining the dataset, the second step is to optimize the YOLOv8n model to reduce computational complexity and improve detection performance, making it suitable for low-computing-power devices.

[0046] (1) Introducing the Adown module to optimize the backbone network: YOLOv8n's traditional downsampling method uses strided convolution, but there is a certain loss in information retention. To this end, this solution introduces the Adown module, which combines maximum pooling with average pooling to enhance the expression of key features. The Adown module uses a multi-path processing method to fully retain feature information, reduce the number of model parameters, and improve detection accuracy.

[0047] (2) Optimizing the Attention Mechanism: The Triplet Attention mechanism is introduced before the SPPF module. Attention weights are calculated using the channel, width, and height dimensions, enabling the model to better focus on key feature areas in complex backgrounds, reducing false detections and missed detections. Compared to the SE attention mechanism, Triplet Attention provides richer feature information, which helps improve detection accuracy.

[0048] (3) Reduce the computational complexity of the model: Optimize the convolution calculation method of YOLOv8n to reduce the amount of computation while maintaining detection capabilities, making it suitable for resource-constrained embedded devices or IoT terminals.

[0049] The optimized YOLOv8n model became the core of subsequent target detection tasks. It has a lighter structure and can perform real-time detection in complex orchard environments, providing input for subsequent neck network optimization.

[0050] 3. Neck Network Optimization To further enhance the feature representation capabilities of detection models in resource-constrained environments, this implementation design utilizes a lightweight gated convolutional block (LGCB)—a lightweight gated convolutional block (LGCB). Compared to existing modules (such as C2f, SE, and CBAM), LGCB no longer relies on multi-branch stacking or heavily parameterized structures. Instead, it achieves information enhancement through the fusion of local perception and dual attention. Architecturally, LGCB first utilizes GSConv for local feature extraction to capture fine-grained spatial information. It then introduces a gated channel attention mechanism (Gated ChannelAttention) to achieve full-channel semantic modeling through lightweight 1D convolutions, which is more efficient than traditional SE. Furthermore, a spatial attention path based on depthwise separable convolutions is added to guide the model's focus on salient regions. Finally, the outputs of the two attention mechanisms are fused and compressed with the original features, achieving a balanced balance between feature enhancement and computational overhead. This module can seamlessly replace the C2f module in YOLOv8, offering significant advantages in lightweight deployment and generalization performance.

[0051] The module formula is:

[0052] It is the output feature map after the input feature X passes through GSConv: It extracts local spatial features with low computational complexity and serves as the basis for the subsequent attention mechanism. is the channel attention weight, Indicates that channel attention is applied to the local feature map, and each channel is scaled by the corresponding attention coefficient. is the spatial attention weight, This is a residual connection, designed to preserve the original feature information and facilitate gradient propagation. The overall logic of the formula: local features are weighted by both channel-wise and spatial-wise attention, and residually fused with the original features. Finally, a 1×1 convolution is used to further compress or transform the channels to output the final feature map.

[0053] 4. Knowledge Distillation Optimization Distill-YOLO After optimizing the backbone network and Neck structure, the fourth step uses knowledge distillation technology to enable the student model to maintain high-precision detection capabilities while reducing computational complexity.

[0054] (1) Constructing a teacher model: YOLOv8m is selected as the teacher model and trained on the complete dataset to enable it to have high-precision target detection capabilities.

[0055] (2) Training the student model: Using knowledge distillation technology, the lightweight student model (Distill-YOLO) can learn the high-quality features of the teacher model while maintaining a low number of parameters and computational complexity.

[0056] (3) Adopting the Feature-CWD distillation method: This scheme adopts channel width distillation (Feature-CWD) to ensure that the student model's feature extraction ability at different scales is not affected, thereby improving the model's generalization ability and detection accuracy.

[0057] (4) Optimizing computing resource requirements: Through knowledge distillation, the student model can achieve efficient reasoning on low-computing power devices, which is suitable for application scenarios such as edge computing and smart agricultural terminals.

[0058] After knowledge distillation optimization, a lightweight Distill-YOLO model is obtained. This model significantly reduces the computing resource requirements while maintaining high detection accuracy, providing a feasible technical solution for practical applications.

[0059] 5. Training and Experimental Verification After the Distill-YOLO model is built, the fifth step is to train and test it to evaluate the optimization effect and compare it with the standard YOLOv8n model.

[0060] (1) Dataset partitioning: The preprocessed dataset is divided into training set, validation set, and test set to ensure that the model can generalize under different data distributions.

[0061] (2) Model training: Distill-YOLO is trained using the stochastic gradient descent (SGD) optimizer to ensure that the model can converge and have high detection capabilities.

[0062] (3) Performance evaluation: The model’s key indicators such as mAP50, mAP50-95, F1 score, and recall rate were evaluated on the test set, and compared with the original YOLOv8n to analyze the improvements in detection accuracy and computational efficiency of the optimized model.

[0063] (4) Experimental comparison: Test the impact of different distillation methods on the student model and verify the advantages of Feature-CWD to ensure that the proposed distillation method can reduce computational complexity while maintaining accuracy.

[0064] After training is completed, the final optimized Distill-YOLO model is obtained, providing an efficient target detection solution for the deployment phase.

[0065] 6. Results Analysis and Deployment After model training and evaluation are complete, the sixth step is to analyze the results and deploy the optimized Distill-YOLO model.

[0066] (1) Analyze the detection effect: By visualizing the detection results, analyze the detection effect of pomegranates at different growth stages, including indicators such as detection accuracy, false detection rate, and missed detection rate.

[0067] (2) Comparison with original YOLOv8n: Compare the advantages and disadvantages of the optimized Distill-YOLO and standard YOLOv8n in terms of computational complexity, detection accuracy, and inference speed to prove the effectiveness of this solution.

[0068] (3) Model deployment: The optimized Distill-YOLO model is deployed to smart agricultural terminal devices, such as smart cameras, drones, and IoT edge computing devices, to achieve real-time detection of the pomegranate growth cycle and improve the level of automation in agricultural management.

[0069] Ultimately, the optimized Distill-YOLO model was successfully applied in the field of smart agriculture, providing an efficient, low-cost, and deployable solution for orchard management.

[0070] Implementation Method 3: Combination Figure 1-2 This embodiment describes a lightweight pomegranate growth period detection method based on knowledge distillation. Combining the improved YOLOv8 algorithm, the Distill-YOLO model was developed for real-time growth period detection of pomegranates in a natural orchard environment. Through knowledge distillation technology, this embodiment successfully optimized a large YOLOv8 model into a small and efficient model, effectively solving the high cost and labor-intensive problems of traditional manual inspections. The research in this embodiment introduces a new lightweight paradigm, providing a potential solution for pomegranate growth period detection. The specific conclusions are as follows: (1) Model design and optimization: This study introduced the ADown and SlimLGNeck modules to effectively reduce the number of model parameters and computational complexity while ensuring detection accuracy. In addition, the combination of the TripletAttention module further improved the model's detection capabilities in natural environments, enabling it to more accurately and efficiently identify pomegranates in complex orchard environments.

[0071] (2) Proposal and Optimization of the Distill-YOLO Model: By combining the improved YOLOv8 model with the aforementioned optimization modules, the Distill-YOLO model was successfully trained. Based on the optimized student and teacher models, the overall performance of the model was further improved. Specifically, the Feature-CWD (channel width distillation) method was used to effectively retain the high accuracy of the teacher model while significantly reducing the number of model parameters and computational complexity.

[0072] (3) Validation Experiments: First, we conducted ablation experiments in the same training environment to analyze the individual impact of each improvement on model performance. Second, we selected the optimal model architecture by comparing different attention mechanisms. After determining the teacher model, we compared different distillation methods and ultimately selected the Feature-CWD distillation method to further optimize the performance of the Distill-YOLO model.

[0073] (4) Experimental results: Experimental results show that Distill-YOLO achieves 92.31%, 79.45%, and 88.72% in mAP50, mAP50-95, and F1 scores, respectively. Compared with the original YOLOv8n model, both accuracy and performance are improved. In addition, the number of parameters and FLOPs of the improved model are reduced to 2.56M and 7.0 GFLOPs, respectively, and the weight size is only 5.2MB. Compared with other lightweight models, Distill-YOLO shows better performance in terms of accuracy and speed, showing its potential for deployment on resource-constrained devices.

[0074] Specifically, YOLOv8 and its improved YOLOv8 are the latest object detection models in the YOLO series, setting a new benchmark in object detection with exceptional speed, accuracy, and lightweight design. By improving the backbone network (Darknet-53, C2f, and SPPF modules), the neck module (C2f replaces C3), and the head structure (decoupling the head and anchor-free approach), YOLOv8 achieves efficient integration of multi-scale features and lightweight optimization. Compared to YOLOv9 and YOLOv10, YOLOv8 performs better in detecting small objects and maintaining a balance between speed and accuracy, making it suitable for a variety of scenarios. The five versions provided (n, s, m, l, x) further enhance its flexibility, making it a reliable baseline model for real-world applications.

[0075] This implementation proposes an improved YOLOv8n model, named Distill-YOLO. This implementation optimizes the backbone and neck networks of the original model. In layers 3, 5, and 7, this implementation introduces the ADown module, replacing traditional convolution operations. This module not only downsamples the feature maps but also enhances the representation of key features through a built-in attention mechanism, effectively reducing the number of model parameters and improving computational efficiency and overall performance. This implementation introduces the TripletAttention mechanism before the SPPF module in layer 10. This mechanism applies attention calculations across the channel, width, and height dimensions to further enhance focus on important features, ensuring accurate capture of key information in complex backgrounds. This improves the model's feature extraction capabilities, particularly for detecting small objects and details. In the neck network, this implementation introduces a Slim-neck architecture, combining the GSConv and VoVGSCSP modules. This design effectively reduces computational complexity, ensures efficient feature fusion, and maintains high detection accuracy while reducing computational burden. GSConv reduces computational overhead by optimizing convolution operations, while VoVGSCSP improves the efficiency of cross-level feature fusion, further enhancing the model's performance in multi-scale object detection. Furthermore, to further improve the model's accuracy and robustness, this implementation utilizes a CWD distillation approach. This method transfers model knowledge from a large model (the teacher model) to a smaller student model, optimizing the student model's training process and significantly improving the model's inference performance and generalization capabilities.

[0076] The ADown module, introduced in YOLOv9, is a lightweight subsampling convolutional module specifically designed for object detection. It aims to preserve image feature information as much as possible while reducing computational complexity and parameter count to ensure detection accuracy. This module combines the advantages of average pooling and max pooling, extracting background information and key features through a block-wise operation. The ADown processing pipeline first downsamples the input feature map using average pooling, halving its size. The downsampled feature map is then split into ×1 and ×2 parts along the channel dimension. A 3 × 3 convolution is applied to the ×1 part for feature extraction and dimensionality reduction; while a max pooling and 1 × 1 point-wise convolution are applied to the ×2 part to enhance nonlinear feature representation and further reduce dimensionality. Finally, the two convolved feature maps are concatenated to form the output of the ADown module. Unlike traditional convolutional downsampling, YOLOv8 uses a 3 × 3 convolution kernel with a stride of 2, batch normalization, and the SiLU activation function for downsampling. However, in real-world applications, devices often have limited computing power, which can impact model performance. The ADown module, through its multi-branch architecture and multi-path processing strategy, achieves comprehensive feature extraction, reduces information loss, and enhances network flexibility and feature representation capabilities. During downsampling, the ADown module combines max pooling and average pooling to more comprehensively extract feature information. The ADown module replaces the convolutional modules in the P3, P4, and P5 layers of the YOLOv8 backbone network, significantly improving detection accuracy while reducing computational overhead and hardware requirements, providing a highly efficient solution for real-time object detection tasks.

[0077] In practical orchard monitoring, given the limited resources and deployment requirements of edge devices, this paper introduces the SlimLGNeck architecture. This module significantly improves detection accuracy through its lightweight design, making it more adaptable to the actual needs of orchard environments. To accelerate prediction computation, the input image in a convolutional neural network (CNN) is typically transformed in the backbone, gradually converting spatial information into channel information. As feature maps are passed layer by layer through the network, the spatial dimension is compressed and the channel dimension is gradually increased. However, some semantic information in the feature maps may be lost during each transformation. Dense convolution can maximize the preservation of this information flow, while sparse convolution may cut off this information. This implementation replaces traditional convolution with a lightweight convolution operation, GSConv. The computational cost of GSConv is approximately 60% to 70% of that of a regular convolution, but its contribution to model learning is similar to that of traditional convolution. By combining grouped and spatial convolutions, GSConv reduces computation while preserving the diversity of feature maps. This design not only improves computational efficiency but also reduces model complexity while maintaining accuracy. However, if GSConv is used in all stages of the model, the number of network layers will increase and the inference time will increase significantly. Therefore, it is more appropriate to use GSConv in the Neck stage, when the channel dimension has been maximized, the width and height dimensions have been minimized, and further conversions are no longer necessary. The processing logic of GSConv includes: first, downsampling the input feature map through ordinary convolution, then applying depthwise convolution (DWConv), and finally connecting the results of the two. Next, a shuffle operation is performed to recombine the compact semantic information from the previous two convolution operations. The goal of GSConv is to make the output of depthwise separable convolution (DSC) as close as possible to the effect of ordinary convolution through this method, while reducing the computational complexity of the model. The shuffle operation enables the information generated by ordinary convolution to effectively penetrate into every part generated by DSC convolution, thereby achieving a better balance between accuracy and computational efficiency.

[0078] Building on GSConv, this paper introduces a new module, GSBottleneck, and uses it to reconstruct the C2f module in the YOLOv8 model. To further enhance the detection model's feature representation capabilities in resource-constrained environments, this implementation design utilizes a lightweight gated convolutional block, LGCB (Lightweight Gated Convolutional Block). Compared to existing modules (such as C2f, SE, and CBAM), LGCB eliminates the need for multi-branch stacking or heavily parameterized architectures. Instead, it achieves information enhancement through the fusion of local perception and dual attention. Architecturally, LGCB first utilizes GSConv for local feature extraction to capture fine-grained spatial information. It then introduces a gated channel attention mechanism (Gated Channel Attention), which achieves full-channel semantic modeling through lightweight 1D convolutions, achieving higher efficiency than traditional SE. Furthermore, a spatial attention path based on depthwise separable convolutions is added to guide the model's focus on salient regions. Finally, the outputs of the two attention mechanisms are fused and compressed with the original features, achieving a balanced approach of feature enhancement and computational overhead. This module can seamlessly replace the C2f module in the YOLOv8 neck, and has obvious advantages in lightweight deployment and generalization performance.

[0079] The module formula is:

[0080] It is the output feature map after the input feature X passes through GSConv: It extracts local spatial features with low computational complexity and serves as the basis for the subsequent attention mechanism. is the channel attention weight, Indicates that channel attention is applied to the local feature map, and each channel is scaled by the corresponding attention coefficient. is the spatial attention weight, This is a residual connection, designed to preserve the original feature information and facilitate gradient propagation. The overall logic of the formula: local features are weighted by both channel-wise and spatial-wise attention, and residually fused with the original features. Finally, a 1×1 convolution is used to further compress or transform the channels to output the final feature map.

[0081] The Triplet Attention mechanism uses an innovative three-branch structure to calculate attention weights. This provides a cost-effective yet effective attention mechanism. Traditional methods for calculating channel attention involve calculating a weight and then uniformly scaling the feature map using that weight.

[0082] Triple Attention consists of three parallel branches, each responsible for capturing different interactions. Some branches track interactions between channels (C) and spatial dimensions (H or W), while the last branch functions similarly to a Convolutional Block Attention Module (CBAM) to construct spatial attention. The outputs of these three branches are then averaged to produce a composite attention score. The structure of Triple Attention is shown in Figure 2. Figure 2 shown.

[0083] Specifically divided into: Spatial Attention, Channel Attention, Sequence Attention The purpose of spatial attention is to assign weights to each spatial position in the input feature map, determining which positions should be emphasized. Spatial attention is usually calculated based on the spatial dimensions of the input (i.e., the features of each position). The specific formula is as follows:

[0084] in, is the input feature map; and is the convolution operation, Responsible for generating a low-dimensional feature map, and The low-dimensional feature map is mapped back to the same spatial size as the input. The spatial attention map is then combined with the input feature map Multiply to achieve weighted operation: . Channel attention is used to assign different importance weights to each channel of the feature map. The calculation of channel attention can be achieved by aggregating the global features of each channel. The specific formula is as follows:

[0085] in is a global pooling operation that compresses the information of each channel into a scalar; Is a weight matrix used to learn the importance of each channel; Used to generate channel attention maps . Channel attention map and input feature map Multiply by channel dimension:

[0086] Sequence attention is mainly used to capture the temporal relationship or spatial-temporal dependency in the input data. Its calculation is usually based on the self-attention calculation of each position of the input feature map. The formula is as follows:

[0087] In the formula, Q and K are respectively obtained by inputting feature maps Perform linear transformation on the query and key matrices; is the scaling factor, and finally, the sequence attention map is combined with the input feature map Multiply them together to get the weighted feature map:

[0088] Finally, the outputs of the spatial, channel, and sequence attention modules are fused, which is the weighted fusion result of the three attention modules and is finally passed to the subsequent layers of the network for processing. The final result is:

[0089] Knowledge distillation is a model compression technique that is mainly used to transfer the knowledge in a large and complex deep learning model (called the teacher model) to a smaller and lightweight model (called the student model) in order to reduce the computing resource requirements while maintaining the performance of the original model as much as possible.

[0090] The core idea of ​​knowledge distillation is to use "soft targets" from the teacher model to guide the student model's learning, in addition to using standard supervised learning objectives (such as cross-entropy loss) when training the student model. The teacher model's output typically includes a probability distribution over categories, rather than just hard labels (i.e., categories). This probability distribution contains richer information, such as similarity or ambiguity between categories. By learning these soft labels, the student model can better understand the teacher model's behavior, thereby achieving higher generalization capabilities. During the knowledge distillation process, the teacher model generates "soft targets" that contain more information about similarities between categories, rather than just the correct category labels. These probability distributions, as the teacher model's output, help the student model understand the deeper characteristics of the data, enabling better training. In knowledge distillation, in addition to conveying category information through the teacher model's soft labels, the teacher model's logical reasoning and feature extraction capabilities are also crucial for the student model.

[0091] In knowledge distillation, the student model is not only trained through traditional supervised learning objectives (such as cross entropy loss), but also learns more similarities and ambiguities between categories by comparing with the "soft labels" output by the teacher model. To achieve this goal, knowledge distillation usually trains the student model through the following loss function:

[0092] in: is the student model output and the true label The cross entropy loss between is used to ensure the performance of the student model on traditional supervised tasks; is the student model output Output from the teacher model Kullback-Leibler ( ) divergence, which measures the difference in their probability distributions; is the weight coefficient used to balance the cross entropy loss and The contribution of divergence loss, usually adjusted between 0 and 1; is a temperature factor used to smooth the output of the teacher model.

[0093] In this way, knowledge distillation effectively transfers the knowledge of the teacher model to the student model, so that the student model can not only learn from the true labels, but also obtain more potential information between categories from the soft labels of the teacher model, thereby improving the generalization ability of the student model.

[0094] Effect verification Table 1 shows the performance of the YOLOv8m teacher model, including three key metrics: F1 score, mAP50, and mAP95, reaching 88.1, 92.45, and 79.89, respectively. These metrics demonstrate that YOLOv8m, as a teacher model, delivers high-quality results in object detection tasks, providing a reliable baseline for the student model distillation process. By transferring knowledge from the teacher model to the more lightweight student model, it is expected that while maintaining high detection accuracy, the model's inference speed and parameter size will be further optimized to meet the application requirements of real-time and resource-constrained scenarios.

[0095]

[0096] This implementation compares and evaluates the YOLOv8n and Distill-YOLO models for detecting plant growth stages in five different stages: bud, early fruit, flower, mid-growth, and Ripr. The performance of the two models was evaluated using three key metrics: F1 score, mean average precision at IoU = 0.50 (mAP50), and mean average precision at IoU = 0.50-0.95 (mAP50-95). The experimental results are shown in Table 2. Distill-YOLO outperforms YOLOv8n in most stages. Specifically, Distill-YOLO achieved superior F1 scores, mAP50, and mAP50-95 in the early fruiting stage (F1: 90.44 vs. 87.91, mAP50: 92.83 vs. 90.35, mAP50-95: 83.00 vs. 79.84) and flowering stage (F1: 86.88 vs. 85.43, mAP50: 95.11 vs. 93.89, mAP50-95: 78.69 vs. 77.35). Although both models showed high performance in the Ripr stage, Distill-YOLO still showed slight improvements in all metrics (F1: 95.28 vs. 95.03, mAP50: 98.73 vs. 96.27, mAP50-95: 91.14 vs. 88.60). In the aggregated "All" category, which combines all growth stages, Distill-YOLO again demonstrates superior performance, achieving higher scores in F1 (88.72 vs. 86.84), mAP50 (92.31 vs. 90.87), and mAP50-95 (79.45 vs. 77.59). These results demonstrate that Distill-YOLO provides a more robust and accurate object detection framework for plant growth stage detection, particularly in terms of the mAP metric.

[0097]

[0098] Experimental results, shown in Table 3, show that by integrating optimization techniques such as Adown, TripletAttention, SlimLGNeck, and Distill, the enhanced YOLOv8n model significantly outperforms the baseline YOLOv8n model across multiple key performance metrics. Specifically, after applying all optimization methods, Model 8 achieved an F1 score of 88.72, mAP50 and mAP50-95 scores of 92.31 and 79.45, respectively, significantly outperforming the baselines of 86.84, 90.87, and 77.59. Furthermore, Model 8's parameter count was reduced to 2.56M, FLOPs to 7.0G, and the model weight to just 5.3MB, demonstrating its ability to effectively reduce computational complexity while improving detection performance. Other enhanced models, such as Model 5 (integrating Adown and SlimLGNeck) and Model 6 (integrating TripletAttention and SlimLGNeck), also achieved F1 scores of 87.26 and 87.68, mAP50 scores of 90.92 and 91.39, and mAP50-95 scores of 77.75 and 77.86, respectively, in different combinations, while maintaining low parameter count and computational complexity. This demonstrates that SlimLGNeck plays a key role in optimizing the model architecture to improve accuracy and efficiency. Furthermore, Model 7, by integrating Adown, TripletAttention, and SlimLGNeck, further improves performance, reaching an F1 score of 87.56, mAP50 and mAP50-95 scores of 91.88 and 78.32, respectively. Notably, Model 8, which incorporates Distill technology, further improves various metrics compared to Model 7, while maintaining the same parameter count and computational complexity. This demonstrates the effectiveness of knowledge distillation in refining the model's predictive capabilities without adding additional computational burden. Overall, the proposed multi-optimization strategy significantly improves the detection accuracy and operational efficiency of the YOLOv8n model through the synergistic effect of different techniques. Model 8, in particular, demonstrates the best balance of performance and efficiency, making it suitable for resource-constrained real-world applications. These results demonstrate the potential of optimization techniques to enhance the performance of object detection models and provide strong support for their future application in real-world deployments.

[0099]

[0100] Ablation studies highlight the significant contribution of the proposed modules (ADown, TripletAttention, and SlimLGNeck) to the detection model's performance. Individually, each module provides measurable improvements in detection accuracy, as reflected in higher F1 scores, mAP50, and mAP50-95 metrics. Notably, combining all three components into the "Distill" model yields the best results, achieving an F1 score of 88.7, a peak mAP50 of approximately 0.9, and an mAP50-95 of approximately 0.8, outperforming the baseline YOLOv8n in all metrics. The "Distill" model also exhibits faster convergence during training and greater stability in later epochs.

[0101] Different distillation methods were applied to optimize the performance of the student model using the YOLOv8m teacher model (F1 = 88.1, mAP50 = 92.45, mAP95 = 79.89). These methods are primarily categorized as feature-based and logical-based. Among feature distillation methods, Feature-cwd (channel width distillation) performed the best, achieving F1 = 88.72, mAP50 = 92.31, and mAP95 = 79.45, approaching or even slightly exceeding the teacher model, demonstrating that this method effectively captures and transfers the high-quality features of the teacher model. However, Feature-mgd (guided mask distillation) performed poorly, with F1 = 27.74, mAP50 = 17.58, and mAP95 = 6.26. This indicates that it fails to effectively adapt the teacher model's information in terms of feature guidance and mask generation, resulting in a significant decline in distillation performance. In logical distillation, different loss functions were used for supervised optimization of the student model. Logical-l1 and Logical-l2 use L1 and L2 losses, respectively, to enforce feature consistency. Logical-l2 slightly outperformed, achieving F1 = 87.92, mAP50 = 91.7, and mAP95 = 78.09, demonstrating good stability and accuracy. Logical-BCKD (KL-divergence-based knowledge distillation) also achieved relatively stable performance, with F1 = 87.37, mAP50 = 91.16, and mAP95 = 77.89, but its overall performance was slightly lower than that of distillation with L2 loss, as shown in Table 4.

[0102]

[0103] In the mAP50 and mAP95 training results of different distillation methods, Feature-cwd demonstrated outstanding performance advantages, significantly outperforming other distillation methods. Its core advantage lies in its ability to efficiently transfer the feature information of the teacher model, especially under the constraints of detailed alignment and channel width at the feature level, ensuring that the student model converges quickly and achieves stable high performance. In addition, Feature-cwd exhibits faster convergence speed in the early stages of training and maintains stable performance in the final stage, demonstrating its accuracy and robustness in feature information transfer, far exceeding other methods such as Logical-l1, Logical-l2, and BCKD, and avoiding the performance bottlenecks encountered by Feature-mgd. This shows that Feature-cwd can fully explore and utilize the knowledge of the teacher model during the distillation process and is an efficient solution for optimizing feature learning of student models.

[0104] The above further describes the technical solution provided by the present invention in detail through several specific embodiments in order to highlight the advantages and benefits of the technical solution provided by the present invention. However, the several specific embodiments described above are not intended to limit the present invention. Any reasonable modification and improvement of the present invention, combination of embodiments and equivalent replacement based on the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pomegranate growth cycle detection method based on a multi-level model, characterized in that: The following steps are involved: Step 1: collect and preprocess an image dataset containing pomegranates at different growth stages to obtain standardized training images; Step 2: Input the training image obtained in step 1 into the backbone network of the improved YOLOv8n model, perform feature extraction and downsampling through the Adown module, and output a primary feature map with enhanced feature expression; Step 3: Input the primary feature map outputted in step 2 into the SPPF module containing the Triplet Attention mechanism to obtain an attention-enhanced intermediate feature map. Step 4: Input the intermediate feature map output in step 3 into the neck network containing the SlimLGNeck structure to perform cross-scale feature fusion to obtain an efficiently fused high-level feature map; In step 5, the high-level feature map output in step 4 is input into the detection head optimized by knowledge distillation to complete the target detection and recognition of pomegranates in different growth cycles.

2. The pomegranate growth cycle detection method of the multi-level model according to claim 1, wherein The Adown module uses a combination of average pooling and maximum pooling to perform downsampling, reducing the number of model parameters and improving feature extraction capabilities.

3. The pomegranate growth cycle detection method of the multi-level model according to claim 1, wherein The Triplet Attention mechanism includes three parallel branches: channel attention, spatial attention, and sequence attention, which are used to enhance the model's attention to key features.

4. The pomegranate growth cycle detection method of the multi-level model according to claim 1, wherein The SlimLGNeck structure includes a GSConv module and a VoVGSCSP module. The GSConv module is used to reduce the amount of calculation, and the VoVGSCSP module is used to enhance the cross-level feature fusion effect.

5. The pomegranate growth cycle detection method of the multi-level model according to claim 1, wherein The GSConv module uses a combination of group convolution and spatial convolution to perform calculations to reduce computational complexity and hardware resource requirements.

6. The pomegranate growth cycle detection method of the multi-level model according to claim 1, wherein The knowledge distillation adopts the Feature-CWD distillation method to transfer feature information from the teacher model to the student model to optimize the performance of the student model.

7. A pomegranate growth cycle detection device based on a multi-level model integrating Triplet Attention and knowledge distillation, characterized by: Includes the following modules: Module 1: Collect and preprocess a dataset of images of pomegranates at different growth stages to obtain standardized training images; Module 2: Input the training image obtained in module 1 into the backbone network of the improved YOLOv8n model, perform feature extraction and downsampling through the Adown module, and output a primary feature map with enhanced feature expression; Module three: Input the primary feature map output by module two into the SPPF module containing the Triplet Attention mechanism to obtain the intermediate feature map enhanced by attention; In module 4, the intermediate feature map output by module 3 is input into the neck network containing the SlimLGNeck structure to perform cross-scale feature fusion to obtain an efficiently fused high-level feature map; In module five, the high-level feature map output by module four is input into the detection head optimized by knowledge distillation to complete target detection and recognition of pomegranates in different growth periods.

8. A computer storage medium for storing a computer program, characterized in that: When the computer program is read by a computer, the computer executes the method according to claim 1 .

9. A computer comprising a processor and a storage medium, characterized in that When the processor reads the computer program stored in the storage medium, the computer executes the method according to claim 1 .

10. A computer program product, being a computer program, characterized in that When the computer program is executed, the method according to claim 1 is implemented.

Citation Information

Cited By

  • Improved YOLOv8n-based cloudy orchard multi-class sheltered pear detection method

    CN121661511A